# AI in Education Knowledge Base — Concepts and FAQs > Full text of 216 concept pages and 32 FAQ pages from the AI in Education Knowledge Base: the syntheses of what the research shows, and the questions those syntheses answer. Small enough to attach to a chat that will not take the complete file. Each page below is linked at its address on the site, so an assistant that can browse may prefer to follow the link; the text is included so an assistant that cannot browse can still answer from it. The individual research papers behind these syntheses are linked from each page and are not reproduced here. # Concepts ## [AI in Education](https://edtechdev.github.io/aied/concepts/ai-education/) > **AI in Education (AIED)** — the broad, interdisciplinary field that applies artificial intelligence to teaching and learning, and studies its design, use, evaluation, and consequences. As the knowledge base's umbrella concept, AI in education encompasses **AI for education** (using AI to improve instruction and assessment) and **education about AI** (developing AI literacy and critical understanding). It sits at the intersection of instructional technology, the [[learning-sciences|learning sciences]] — the empirical research field that asks whether a learner changed rather than only whether a tool performed — computer science, [[educational-policy-ai|educational policy]], [[ethics]], and [[equity-in-ai-education|equity]]. This page is an introduction to the field and a map to every concept the knowledge base covers. ## Questions to Consider - 'AI in education' spans two directions: AI for education (using AI to improve teaching) and education about AI (building literacy and critical understanding). Which is more familiar to you, and which do you tend to overlook? - The field's history is framed as a recurring tension between control and learner agency. As you watch AI tools being adopted, where do you notice that same tension playing out today? - AI in education sits at the intersection of technology, learning science, policy, ethics, and equity. Which of those lenses do you naturally apply when evaluating an AI tool — and which are you likely to forget? - This knowledge base organizes the field into strands: pedagogy, learning theories, technologies, disciplines, assessment, feedback, stakeholders, and governance. If you were mapping your own use of AI, which strand would you find yourself in? - AIED includes teaching students to use AI critically as a goal in itself. In your context, is AI treated more as a subject to be taught or a tool to be used — and does that balance reflect what learners actually need? - Students may learn AI literacy by using AI critically, not just by learning about AI. How might hands-on, critical use build understanding that passive instruction cannot? ## Introduction AI in education is the umbrella that all other concept pages collectively define. The knowledge base organizes the field into the major strands below, each linking to the relevant concept pages. A landmark [[history-of-aied|historical]] perspective, **[[mishra-control-vs-agency-history-2025|Mishra et al.]]** trace AIED from cybernetics and the 1956 Dartmouth conference through [[intelligent-tutoring|cognitive tutors]] and Papert's [[constructivist|constructionism]], arguing that today's [[generative-ai|GenAI]] debates re-enact the field's foundational control-vs-[[agency]] tension. ## How the knowledge base is organized: the umbrella pages The knowledge base's concept coverage is anchored by several **umbrella pages** that group related concepts into navigable strands. These are good entry points for exploring the field: - **AI in Education** — this page, the field's overview and map to all coverage. - **[[ai-literacy|AI literacy]]** — the umbrella for understanding, using, and critically evaluating AI, spanning [[prompt-engineering|prompt engineering]], [[critical-thinking|critical thinking]], [[ethics|AI ethics]], and [[reducing-ai-misuse|responsible use]]. Alongside it, [[human-ai-collaboration|human–AI collaboration]] and [[agentic-ai|agentic AI]] frame how people and AI work together. - **[[pedagogy|Pedagogies and teaching strategies]]** — the umbrella for how teaching happens: the teaching methods and strategies AI operates within ([[active-learning|active]], [[collaborative-learning|collaborative]], [[project-based-learning|project-based]], [[problem-based-learning|problem-based]], [[experiential-learning|experiential]], [[game-based-learning|game-based]], [[socratic-method|Socratic]], [[scaffolding]], and more), including the distinct context of [[online-teaching-and-learning|online teaching and learning]]. - **[[learning-theories|Learning theories]]** — the umbrella for how learning happens: the theoretical frameworks ([[behaviorism]], [[cognitive-psychology|cognitivism]], [[constructivist|constructivism]], [[sociocultural-learning|sociocultural]], cognitive, [[motivation|motivational]]) that shape AI design and evaluation. - **[[ai-technologies|Technologies]]** — the umbrella for the technical layer: the AI systems ([[llm|LLMs]], [[generative-ai|generative AI]], [[multimodal]], [[educational-robotics|robotics]]) and techniques ([[rag]], [[prompt-engineering|prompt engineering]], [[reinforcement-learning|reinforcement learning]], [[pedagogical-llm-training|model training]], [[agentic-ai|agentic orchestration]]) that power AIED. The learner-modeling family — [[knowledge-tracing|knowledge tracing]], [[cognitive-diagnosis|cognitive diagnosis]], [[simulating-students|simulating students]], and the systems that consume them ([[intelligent-tutoring|intelligent tutoring]], [[adaptive-learning|adaptive learning]], [[personalized-learning|personalized learning]]) — is grouped under the [[student-modeling|Learner Modeling and Adaptive Instruction]] umbrella within this technical strand. - **[[discipline-specific-aied|AI in the disciplines]]** — the umbrella for how AI is applied across subject areas ([[math-education|mathematics]], [[physics-education|physics]], [[language-learning|language learning]], [[cs-education|computer science]], [[writing-education|writing]], [[stem-education|STEM]], [[engineering-education|engineering]], [[business-education|business]], [[teacher-education|teacher education]], [[medical-education|health professions]], and more) — and, past the academic subjects, the professional and applied strand that assesses demonstrated practice rather than correctness ([[nursing-education|nursing]], [[information-technology|information technology]], [[vocational-education|vocational education and training]], [[design-education|design education]]) — and educational levels ([[k-12]], [[higher-ed|higher education]], [[adult-learning|adult learning]]). - **[[assessment]]** (with [[formative-assessment|formative]], [[summative-assessment|summative]], [[authentic-assessment|authentic]], [[oral-assessment|oral]], and [[automated-assessment|automated]] strands) — the umbrella for how AI both assesses learners and reshapes assessment validity and integrity. - **[[feedback]]** — the umbrella for how feedback is generated, delivered, and used: the feedback loop, [[ai-feedback-quality|feedback quality]], [[feedback-literacy|feedback literacy]], and its assessment contexts ([[formative-assessment|formative]], [[peer-assessment|peer]], [[automated-assessment|automated]]). - **[[stakeholders|Stakeholders in AI education]]** — the umbrella for who the actors are: learners, [[teacher-role|teachers]], [[learning-design|learning designers]], [[administrator|administrators]], and [[educational-policy-ai|policymakers]]. - **[[ai-ed-evaluation|AI ed evaluation]]** and **[[research-methods-aied|research methods]]** — the umbrellas for how we know whether AI works: efficacy studies, [[benchmark|benchmarks]], [[rct|randomized controlled trials]], [[meta-analysis-systematic-review|meta-analysis]], and [[learning-gains|learning gains]] as the core outcome measure. Readers should also weigh the [[limitations-in-aied-research|cross-cutting limitations of this evidence]]. - **[[governance|AI governance]]**, **[[educational-policy-ai|educational AI policy]]**, and **[[equity-in-ai-education|equity]]** — the umbrellas for the institutional, regulatory, and fairness layer (see also [[regulation]] and [[privacy]]). These umbrella pages are linked throughout the sections below; each strand below names both its umbrella and its constituent concepts. ## Two dimensions of AI in education AI in education research spans two interconnected directions: - **AI for education** — using AI systems to enhance teaching, learning, assessment, and administration. This includes [[intelligent-tutoring|AI tutoring]], [[adaptive-learning|adaptive learning]], [[personalized-learning|personalized learning]], [[automated-essay-scoring|automated essay scoring]], [[automated-question-generation|automated question generation]], [[automated-assessment|automated assessment]], [[formative-assessment|formative assessment]], [[learning-analytics|learning analytics]], and [[feedback|feedback loops]]. - **Education about AI** — teaching learners and educators to understand, use, and critically evaluate AI. The core is [[ai-literacy|AI literacy]], supported by [[prompt-engineering|prompt engineering]], [[critical-thinking|critical thinking]], [[ethics|AI ethics]], [[governance|governance education]], digital literacy, and [[reducing-ai-misuse|responsible use]]. These two dimensions are not separate: [[ai-literacy|using AI well]] requires understanding it, and teaching about AI is enriched by using it. This [[human-ai-collaboration|human-AI collaboration]] is a central theme. ## Foundations of AI in education The field's cross-cutting and foundational concepts anchor the knowledge base's coverage and appear first in the sidebar. They open with an **Essentials** group — the concepts every reader should start with: the umbrella itself, [[misconceptions|misconceptions about AI]], [[ai-literacy|AI literacy]], [[agentic-ai|agentic AI]], [[cognitive-offloading|cognitive offloading]], [[framing-ai-use-for-students|how AI use is framed for students]], [[reducing-ai-misuse|reducing AI misuse]], [[academic-integrity|academic integrity]], [[teacher-role|teaching]], [[learning-design|learning design]], and [[educational-development|educational development]]. The **field** strand then covers [[history-of-aied|the field's history]], the [[limitations-in-aied-research|cross-cutting limitations of the evidence base]], [[philosophy-of-ai-in-education|its philosophy]], the [[theories-and-frameworks|theories and frameworks]] map, and [[theory-development-aied|theory development]]. The cross-cutting themes — [[human-ai-collaboration|human–AI collaboration]], [[agency|learner agency]], [[learner-identity|learner identity]], [[design-thinking|design thinking]], [[curriculum-design|curriculum design]], [[critical-thinking|critical thinking]], and [[computational-thinking|computational thinking]] — cut across every strand, because the inaccurate mental models people hold about AI are upstream of [[ai-misuse-learning-harm|misuse]] and under-calibrated [[trust-calibration|trust]]. ## Learning and instruction How AI supports teaching and learning is the heart of the field. Key concepts include: - **Core pedagogies:** [[pedagogy|pedagogies and teaching strategies]] — the umbrella for the knowledge base's teaching-methods coverage — along with [[active-learning|active learning]], [[collaborative-learning|collaborative learning]], [[group-work|group work]], [[project-based-learning|project-based learning]], [[problem-based-learning|problem-based learning]], [[productive-failure|productive failure]], [[inquiry-based-learning|inquiry-based learning]], [[experiential-learning|experiential learning]], [[game-based-learning|game-based learning]], [[learning-by-teaching|learning by teaching]], [[scaffolding]], [[socratic-method|the Socratic method]], [[critical-pedagogy|critical pedagogy]], [[pedagogical-partnerships|pedagogical partnerships]], [[storytelling-in-education|storytelling]], [[learning-design|learning design]], [[online-teaching-and-learning|online teaching and learning]], and [[video-education|video in education]]. - **Learning theories and processes:** the [[learning-theories|learning theories]] umbrella ([[behaviorism]], [[cognitive-psychology|cognitivism]], [[constructivist|constructivism]], [[sociocultural-learning|sociocultural]], [[distributed-cognition|distributed cognition]], [[situated-learning|situated learning]], [[embodied-learning|embodied learning]], [[community-of-inquiry|community of inquiry]]) sits alongside learner-facing processes like [[self-regulated-learning|self-regulated learning]], [[self-determination-theory|self-determination theory]], [[motivation]], [[self-efficacy]], [[self-directed-learning|self-directed learning]], [[metacognition]], [[desirable-difficulties|desirable difficulties]], [[transfer-of-learning|transfer of learning]], [[prior-knowledge|prior knowledge]], [[icap-framework|ICAP cognitive engagement]], [[refutation-text|refutation text]], [[retrieval-spacing-interleaving|retrieval, spacing and interleaving]], and [[activity-theory-aied|activity theory]]. - **Learner engagement and experience:** [[student-engagement|student engagement]], [[help-seeking]], [[social-emotional-learning|social-emotional learning]], [[well-being]], [[creativity]], [[problem-solving|problem solving]], [[mastery-learning|mastery learning]], and [[student-ai-interaction|student–AI interaction]] shape how learners actually encounter and are affected by AI, while the [[social-norms-ai-use|social norms that settle around AI use]] decide how openly any of it can be discussed. ## Technologies and techniques The [[ai-technologies|Technologies]] page is the umbrella for the technical layer: - **Models and techniques:** [[generative-ai|generative AI]], [[llm|large language models]], [[rag|retrieval-augmented generation]], [[multimodal|multimodal models]], [[educational-nlp|educational NLP]], [[reinforcement-learning|reinforcement learning]], [[knowledge-graph|knowledge graphs]], [[educational-robotics|robots in education]], [[conversational-ai|conversational AI]], [[simulation]], and [[pedagogical-llm-training|training pedagogical LLMs]]. The underlying methods matter too: [[machine-learning|machine learning]] is where these systems are built, [[speech-and-voice-technologies|speech and voice technologies]] carry spoken tutoring and language practice, [[visualization]] covers dashboards and visual analytics for learners and instructors, and [[virtual-and-augmented-reality|virtual and augmented reality]] hosts immersive practice whose visual layer the model can now generate. Newer interaction styles belong here too — most prominently [[vibe-coding|vibe coding]], the natural-language-driven workflow in which the user specifies a program by prompting an LLM and judges the resulting behavior rather than reading or editing source, which reframes [[cs-education|programming]] as an act of expression and verification and lowers the barrier to [[teacher-role|end users]] building their own tools. Integration-depth frameworks such as [[samr-model|SAMR]] and adoption theories such as [[technology-acceptance-model|TAM]] classify how deeply AI is taken up and how much it transforms the task. - **Learner modeling and adaptive systems:** the technical systems that represent and adapt to the learner are grouped under the [[student-modeling|Learner Modeling and Adaptive Instruction]] umbrella — [[knowledge-tracing|knowledge tracing]], [[cognitive-diagnosis|cognitive diagnosis]], [[simulating-students|simulating students]], [[intelligent-tutoring|intelligent tutoring]], [[adaptive-learning|adaptive learning]], [[personalized-learning|personalized learning]], [[recommender-systems-and-learning-paths|recommender systems and learning paths]], [[pedagogical-agent|pedagogical agents]], [[affective-tutoring|affective tutoring]], [[affective-computing|affective computing]], [[human-in-the-loop-ai|human-in-the-loop AI]], and [[learning-analytics|learning analytics]]. These sit at the technical layer because they are the AI systems themselves, distinct from the pedagogies they enact. ## AI in the disciplines AI is applied across disciplines and educational levels. The knowledge base's [[discipline-specific-aied|overview of AIEd in the disciplines]] maps subject-area coverage — alongside [[learning-sciences|learning sciences]], which is not a taught subject but the cross-cutting research field that studies learning itself and takes subject matter as one variable among others: - **Subject areas:** [[math-education|mathematics]], [[physics-education|physics]], [[chemistry-education|chemistry]], [[biology-education|biology]], [[cs-education|computer science]], [[engineering-education|engineering]], [[stem-education|STEM]], [[writing-education|writing]], [[language-learning|language learning]], [[english-education|English education (EAP/EFL/ESL)]], [[science-education|science education]], [[business-education|business, economics, and management]], [[humanities-education|humanities and social sciences]], [[arts-design-and-media-education|arts, design and media education]], [[medical-education|medical and health professions]], [[legal-education|legal education]], and the professional and applied strands — [[nursing-education|nursing]], [[information-technology|information technology]], [[vocational-education|vocational education and training]], and [[design-education|design education]]. ## Levels and contexts The same AI tool meets very different settings, and the knowledge base separates the educational level from the pedagogy so that findings do not silently transfer across them: [[k-12|K-12 schools]], [[early-childhood-elementary-ai-education|early childhood and elementary education]], [[higher-ed|higher education]], [[adult-learning|adult learning]], [[vocational-education|vocational education and training]], [[special-education|special education]], and [[teacher-education|teacher education]]. Domain-adjacent concepts that cut across levels include [[universal-design-for-learning|universal design for learning]], [[neurodiversity]], [[multilingual-learning|multilingual learning]], and [[social-emotional-learning|social-emotional learning]]. ## Assessment and measurement AI transforms both how we assess learners and how we evaluate AI systems themselves: - **Assessment and feedback:** [[assessment]], [[formative-assessment|formative assessment]], [[summative-assessment|summative assessment]], [[authentic-assessment|authentic assessment]], [[eportfolio|e-portfolio]], [[feedback]] and [[feedback-literacy|feedback literacy]], [[ai-feedback-quality|AI feedback quality]], [[peer-assessment|peer assessment]], [[oral-assessment|oral assessment]], [[automated-assessment|automated assessment]], [[automated-essay-scoring|automated essay scoring]], and [[automated-question-generation|automated question generation]]. Because a model can now produce plausible finished work on demand, the knowledge base foregrounds the capability that remains the learner's own: [[evaluative-judgment|evaluative judgment]], the capacity to appraise the quality of one's own work, peers' work, and AI output against reasoned criteria. It is the construct several feedback and authentic-assessment studies converge on — the hybrid feedback condition outperforming direct AI feedback in a multisite experiment, the sustainability gap in AI formative feedback, and the practical move of assessing the decisions students make rather than only the artifact — and it is a core reason AI-era redesign shifts from [[ai-detection|detection]] toward tasks whose integrity survives inspection. [[group-work|Group work]] is likewise assessed along both process and product, where teams must negotiate whose and what kind of AI engagement counts as acceptable. - **Measurement and validity:** [[assessment-validity|assessment validity]], [[psychometrically-aware-ai|psychometrically aware AI]], [[educational-measurement|educational measurement]], [[item-response-theory|item response theory]], [[self-report-measures|self-report measures]] (the instrument behind a large share of this evidence, and a recurring limitation), [[ai-detection|AI detection]], [[remote-proctoring|remote proctoring]], and [[academic-integrity|academic integrity]]. ## Research methods and evaluation How we know whether AI works is its own strand, and the knowledge base treats it as one: - **Research methods:** [[research-methods-aied|research methods in AIED]] as the umbrella, with [[qualitative-research|qualitative]], [[quantitative-research|quantitative]], [[mixed-methods-research|mixed-methods]], [[design-based-research|design-based]], and [[usability-research|usability]] approaches, plus [[rct|randomized controlled trials]], [[meta-analysis-systematic-review|meta-analysis and systematic review]], and [[network-analysis|network analysis]]. - **Evaluation of AI systems:** [[ai-ed-evaluation|AI ed evaluation]] and [[benchmark|benchmarks]] for judging a system's capability, with [[learning-gains|learning gains]] as the outcome that matters, and the [[limitations-in-aied-research|cross-cutting limitations]] of this evidence and [[interpreting-and-applying-aied-research|how to read a single study]] as the cautionary counterweight. ## People AI in education changes the roles of every stakeholder. The knowledge base's [[stakeholders|Stakeholders in AI education]] page is the umbrella covering all of them: - **Learners:** [[student-experience|student experience]], [[career-development-and-readiness|career development and readiness]], and [[anxiety-and-stress|AI anxiety and stress]] shape how students encounter AI. - **Families and communities:** [[parents-and-families|parents and families]] are the audience schools address most often about AI and the one with the least research behind the guidance, so their concerns belong in the stakeholder picture rather than outside it. - **Instructors and teaching frameworks:** [[teacher-ai-competency|teacher AI competency]], [[tpack|technological pedagogical content knowledge (TPACK)]], [[samr-model|SAMR]], and [[educational-development|educational development]] address educator preparation and support. - **Builders:** [[educational-technology-developers]] — the product designers, software developers, learning engineers and analytics designers who turn a model capability into something an institution can procure. They are a distinct audience from the practitioners and administrators above, and they sit outside the institutions that adopt their tools, which is why defaults, co-design and post-funding maintenance appear in this knowledge base as pedagogical questions rather than commercial ones. ## Institutions and policy The institutional layer is where AI decisions are actually made and defended: [[administrator|administrators]] and institutional leaders, [[educational-policy-ai|educational AI policy]], [[governance|AI governance]], [[change-management|change management]] as the work of making an adoption stick, [[regulation|AI regulation]], and the procurement and platform questions that follow from [[technology-acceptance-model|technology adoption]], [[open-source|open source]], and [[edtech-platform|edtech platforms]], alongside [[lifelong-learning|professional and lifelong learning]] and [[professional-training|professional training]]. ## Equity, ethics, and responsible use Fairness, access, and responsibility are central to AI in education: - **Equity and access:** [[equity-in-ai-education|Equity]], [[differential-effects-across-learner-groups|differential effects across learner groups]] (the question of who a finding holds for), [[digital-divide|digital divide]], [[bias-mitigation|bias mitigation]], [[culturally-relevant-pedagogy|culturally relevant pedagogy]], [[multilingual-learning|multilingual learning]], [[inclusive-learning|inclusive learning]], [[accessibility]], [[assistive-technology|assistive technology]], [[neurodiversity]], [[universal-design-for-learning|universal design for learning]], and [[global-south|Global South]] studies. - **Ethics and responsibility:** [[ethics|AI ethics]], [[ai-misuse-learning-harm|AI misuse and learning harm]], [[legal-issues-and-risks|legal issues and risks]], [[ai-use-disclosure|AI use disclosure]], [[guardrails]], [[privacy]], [[hallucination-risk|hallucination risk]], [[ai-sycophancy|AI sycophancy]], [[trust]], [[trust-calibration|trust calibration]], [[reducing-ai-misuse|reducing AI misuse]], [[framing-ai-use-for-students|how AI use is framed for students]], [[pedagogical-safety|pedagogical safety]], [[sustainability]], and [[cognitive-offloading|cognitive offloading]]. A [[meta-analysis-systematic-review|systematic review]] of the field's ethics literature ([[agarwal-ethical-values-norms-aied-2026|Agarwal et al. 2026]], 25 articles) consolidates AIED ethics into six main ethical values — non-discrimination, data stewardship, human oversight, goodwill, explicability, and educational aptness — and maps the ethical norms onto a stakeholder-by-value matrix. It finds end users largely passive in the ethical literature (student voices essentially absent) and calls for integrating ethics into AIED design and a greater focus on the educational (pedagogical) dimension of AIED ethics. ## Emergent and cross-cutting themes Several themes cut across the field: - **Trust and critical use:** [[trust]], [[trust-calibration|trust calibration]], [[ai-sycophancy|AI sycophancy]], [[critical-thinking|critical thinking]], [[cognitive-offloading|cognitive offloading]], [[critical-pedagogy|critical pedagogy]], and [[reducing-ai-misuse|reducing AI misuse]] (see also [[framing-ai-use-for-students|how AI use is framed for students]]). How learners and teachers decide to adopt and rely on AI is modeled by [[technology-acceptance-model|technology acceptance]] research, while [[global-south|Global South]] studies foreground equity and cultural context in adoption. Empirically, how AI explains itself shapes this trust: [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor et al. (2025)]] showed that [[explainable-ai|explainable AI]] builds teachers' trust in AI recommendations through understandability, with domain-driven (curricular-language) explanations trusted and accepted more than data-driven feature-importance output. - **The evolution of the field:** the knowledge base traces AI in education from early [[intelligent-tutoring|intelligent tutoring systems]] and [[knowledge-tracing|knowledge tracing]] to LLM-driven [[intelligent-tutoring|tutoring]], [[pedagogical-agent|agents]], and [[agentic-ai|agentic AI]] — a rapid shift from tool-centric studies to sociotechnical frameworks ([[design-thinking|design thinking]], [[curriculum-design|curriculum design]], [[institutional-change-framework-ai|institutional change]]), and from hand-authored systems to user-driven workflows in which a learner or [[teacher-role|non-programmer]] specifies behavior in natural language ([[vibe-coding|vibe coding]]). [[rismanchian-ai-education-four-decades-aixed-2026|Rismanchian & Doroudi]] formalize this trajectory with their AI×Ed framework, tracing papers across four decades of proceedings to show that the field moved from a diverse mix — including substantial research treating AI as an analogy to human intelligence and learning — toward a near-exclusive focus on applied, data-driven, researcher-facing uses, a turn that the rise of [[llm|LLMs]] now appears to partly reverse (three of the four "AI-as-analogy" papers at AIED 2024 were LLM-based). - **Emotion, anxiety, and career futures:** AI induces and shapes emotional responses — [[anxiety-and-stress|AI anxiety and stress]] spanning proctoring surveillance, integrity fears, and career displacement — while [[career-development-and-readiness|career development and readiness]] addresses how education prepares learners for an AI-disrupted labor market (see also [[well-being]]). ## Field maturity The knowledge base reflects a field in rapid evolution — from early intelligent tutoring systems to LLM-driven tutoring and agentic AI; from detection-focused academic-integrity tools to assessment redesign; from tool-centric studies to sociotechnical and equity-focused frameworks. The evidence base increasingly emphasizes rigorous [[research-methods-aied|research methods]], [[ai-ed-evaluation|evaluation]], [[rct|experimental designs]], and long-term outcomes. ## Connections AI in education connects to every concept in the knowledge base — it is the field that all other concept pages collectively define. Use this page as a starting point to navigate the full knowledge base. ## Connected Concepts - [[ai-literacy]] — umbrella: understanding, using, and evaluating AI - [[human-ai-collaboration]] — umbrella: how people and AI work together - [[pedagogy]] — umbrella: teaching methods and strategies - [[learning-theories]] — umbrella: how learning happens - [[ai-technologies]] — umbrella: models, techniques, and systems - [[discipline-specific-aied]] — umbrella: AI across subject areas and levels - [[assessment]] — umbrella: how AI assesses learners and reshapes validity - [[feedback]] — umbrella: how feedback is generated, delivered, and used - [[stakeholders]] — umbrella: who the actors are - [[learners]] — umbrella: the learner-side concepts (experience, identity, agency, interaction, learner models) - [[ai-ed-evaluation]] — umbrella: how we know whether AI works - [[research-methods-aied]] — umbrella: efficacy research methods - [[governance]] — umbrella: the institutional and regulatory layer - [[educational-policy-ai]] — umbrella: policy, guidance, and implementation - [[equity-in-ai-education]] — umbrella: fairness, access, and inclusion - [[learning-sciences]] — the empirical field behind AIED - [[ethics]] — the ethical dimensions of AI in education - [[misconceptions]] — the mental models people bring to AI - [[interpreting-and-applying-aied-research]] — reading a study, and carrying a finding into practice - [[limitations-in-aied-research]] — cross-cutting limits of the evidence base - [[meta-analysis-systematic-review]] — what the reviews and meta-analyses establish - [[history-of-aied]] — how the field evolved - [[philosophy-of-ai-in-education]] — the philosophical foundations - [[theories-and-frameworks]] — the map of theory and framework nodes - [[theory-development-aied]] — building and revising theory ## Connected Articles Field-wide reviews of AI in education — the studies that survey the whole field or a whole educational level rather than one topic: - [[typology-generative-ai-tools-education-2026]] — Typology of Generative AI Tools for Education - [[raza-farooq-aied-review-2020-2025]] — Review of Artificial Intelligence in Education from 2020 to 2025 - [[rismanchian-ai-education-four-decades-aixed-2026]] — The evolution of AI-and-education research across four decades (AIxEd framework) - [[mishra-control-vs-agency-history-2025]] — Control vs. agency: a history of AI in education - [[liang-genai-systematic-review-human-ai-2026]] — Generative AI in education: systematic review of 56 empirical studies - [[genai-higher-education-systematic-review-2026]] — Generative AI in higher education: systematic review of 125 studies - [[stanford-evidence-base-ai-k12-2026]] — The evidence base on AI in K-12: a review of 818 papers - [[caruana-pre-university-ai-education-slr-2026]] — Pre-university AI education: systematic literature review of 42 studies - [[genai-educational-outcomes-meta-analysis]] — Generative AI and educational outcomes: comprehensive meta-analysis - [[caeai-ai-companions-learning-over-performance-2026]] — a research agenda for AI companions built around learning rather than performance --- ## [AI Literacy](https://edtechdev.github.io/aied/concepts/ai-literacy/) > **AI literacy** — the knowledge, skills, and critical dispositions needed to understand, evaluate, and effectively use AI [[ai-technologies]] in educational contexts. AI literacy spans foundational understanding of how AI works, practical competence in using AI tools, critical [[ai-ed-evaluation|evaluation of AI]] outputs, and ethical awareness of AI's societal implications. ## Questions to Consider - What does it mean to be 'AI literate'? Is it knowing how to use the tools, understanding how they work, or being able to critically evaluate their output — and which matters most for your own goals? - Self-reported AI literacy diverges sharply from measured performance — teachers overestimate their skills by about 40%. How confident are you in your own AI skills, and what evidence would you trust to actually test that confidence? - AI literacy is described as a metacognitive social practice, not a checklist of skills. Because LLMs are probabilistic and opaque, what does it mean to 'critically evaluate' output when you can't see inside the model? - There's a tension between critical-use literacy for academic settings and workflow-integration literacy for employment — educators and employers value different skills. Which kind of AI literacy is your context actually teaching? - Technical AI-literacy training alone, without self-efficacy and self-regulation support, can actually increase dependency on AI. How could learning to use AI make you more reliant on it rather than more capable? - AI literacy is also a recognition skill: spotting when an AI is agreeing with you because it's being sycophantic versus because you're right. Have you ever caught an AI just telling you what you wanted to hear? ## Introduction AI literacy has rapidly emerged as a core competency for learners, educators, and institutions as [[generative-ai]] becomes embedded in education. Unlike general digital literacy, AI literacy requires understanding probabilistic systems that can hallucinate, exhibit [[bias-mitigation|bias]], and shift [[agency]] from human to machine — making [[critical-thinking]] and [[cognitive-offloading|Over-Reliance]] central to the construct. ### Dimensions of AI literacy AI literacy [[research-methods-aied|research]] in this knowledge base spans four interconnected dimensions: Frameworks increasingly trace how these dimensions are enacted in practice: [[dai-chan-responsible-genai-research-ai-literacy-2026|Dai & Chan (2026)]] show how postgraduate researchers enact all four dimensions when using [[generative-ai]] across the research workflow and scaffold research-specific responsible-use guidelines from them, while [[san-orhan-karsak-ai-cognition-micro-credentials-2026|Şan & Orhan Karsak (2026)]] use Word-Association mapping to show that Turkish undergraduates' AI cognition is instrumentally "black-box" dominated, with ethical-governance concepts structurally isolated (τ=−0.819) — evidence that credential design must deliberately bridge experiential tool use and ethical frameworks. **Foundational knowledge:** Understanding what [[llm|LLMs]] are, how they differ from rule-based systems, and their fundamental limitations. This includes awareness of model capabilities, training data biases, and the distinction between task-specific AI and general-purpose models. Research in [[prompt-engineering]] examines how understanding prompt mechanisms affects effective AI use. **Practical competence:** The ability to use AI tools effectively — from [[prompt-engineering]] to interpreting outputs. Studies of [[genai-usage-design-students-survey|student GenAI usage patterns]] reveal that tool access alone doesn't produce competence; structured practice and [[scaffolding]] are essential. The [[gaide-vibe-coding-k12-teachers|vibe coding framework]] shows how [[k-12|K-12 teachers]] can develop practical AI literacy through guided tool creation. [[miles-prompt-literacy-human-centered-genai-framework-2026|Miles, Haber-Curran and Arar (2026)]] argue that this dimension has a part optimization training misses: prompt literacy, as distinct from prompt engineering, treats prompting as a rhetorical and ethical act, so instruction that only tunes outputs for better performance leaves the ethical and epistemological dimensions of LLM use untouched. **Critical evaluation:** The capacity to assess AI outputs for accuracy, bias, and appropriateness. [[ai-literacy-assessment-misalignment|Research on literacy assessment]] shows a 40% gap between self-reported and performance-based AI literacy — people consistently overestimate their evaluation skills. That divergence is the clearest case in the knowledge base of a [[self-report-measures|self-report measure]] failing to track the competence it names. This connects to [[cognitive-offloading|Over-Reliance]] research showing that students who [[trust]] AI uncritically learn less. Recent conceptual work pushes evaluation toward *verification*: the [[pearls-epistemic-verification-2026|PEARLS framework]] (Wang) treats AI output as a provisional knowledge claim whose warrant must be assembled and examined across six dimensions (Process, Evidence, Access, Reproducibility, Legitimacy, Source), and advances **verification-driven learning** as the mechanism by which learners build expertise while checking AI claims. Complementing this, the [[student-centered-genai-responsible-framework-2026|student-centered responsible-use framework]] (Alsammani) offers ten student-facing guidelines across Learning and Growth, Ethics and Integrity, and Awareness and Safety pillars that externalize [[metacognition|metacognitive]] prompts at the point of decision. For [[k-12|secondary]] learners, the [[aarc-ai-research-competency-2026|AI-Assisted Research Competency (AARC) framework]] (Beau, Flaquière & Lazar 2026) grounds this in virtue epistemology and AI intuition, defining research literacy as conducting inquiry with AI without surrendering [[agency|authorship]], judgment, verification, or responsibility — operationalized as an analytic rubric and a verify–cite–reflect commitment routine. At the [[assessment]] end, [[human-capability-test-learning-outcomes-ai-2026|Saleh (2026)]] proposes a *human capability test* that turns evaluation into a design principle: ask what a student must demonstrate independently, what can be strengthened through AI augmentation, and what the student must verify, defend, and take responsibility for. Disciplinary verification raises the same bar further. [[vega-baudrit-genai-university-chemistry-education-review-2026|Vega-Baudrit and Rivera Alvarez (2026)]] argue that in university chemistry the core AI literacy demand is representational translation across Johnstone's macroscopic, submicroscopic and symbolic domains, because an output can be locally persuasive in one domain while contradicting charge balance, a mechanism or a safety limit in another. Their remedy is verification-centered integration: make verification an assessed activity, keep prompt logs and revision histories as reasoning traces, and treat prompting as an epistemic act that specifies constraints, assumptions and adequacy criteria, with students identifying a false assumption or justifying rejection of a generated answer. **Ethical and institutional awareness:** Understanding AI's broader implications — from [[academic-integrity]] to [[equity-in-ai-education]] to [[privacy]]. AI literacy at the institutional level involves policy development, [[educational-development]], and [[governance|governance frameworks]] — institutional AI literacy is a matter of [[educational-policy-ai|policy]] as much as pedagogy. The [[sangwa-epiq-ai-faculty-readiness-2026|EPIQ-AI framework]] frames institutional AI literacy as a sociotechnical alignment challenge, not just individual training. - **AI literacy as a governance capacity for sustainable development.** [[ai-literacy-sdg-governance-framework-2026|Islam, Morshed, and Islam (2026)]] reconceptualize AI literacy as a governance-oriented capacity rather than a purely educational or technical skill, linking it to all seventeen UN Sustainable Development Goals. Their six-level **AIRE Taxonomy** (Recognize → Comprehend → Apply → Analyze → Integrate → Govern) extends Bloom's hierarchy by adding ethical synthesis and strategic foresight, positioning advanced competencies (Analyze–Govern) as the pathway from foundational literacy to institutional and policy-level governance — an "18th SDG" heuristic that treats literacy as a cross-cutting cognitive and ethical bridge. A survey of 300 professionals in a national context found strong technical awareness but limited ethical and governance readiness, with **ethical reasoning and reflective thinking the strongest predictors of sustainable, trustworthy AI use** and governance literacy the strongest predictor of AI–SDG nexus awareness (β = 0.64). This empirically grounds the knowledge base's emphasis on critical-use literacy and ties AI literacy directly to [[sustainability]] and [[educational-policy-ai|policy]] integration. ### How AI literacy is developed Research points to [[collaborative-learning|collaborative]] and [[active-learning|active]] approaches as most effective. [[rismanchian-ai-education-four-decades-aixed-2026|Rismanchian & Doroudi]] position AI literacy as more than applied skill: in their AI×Ed framework the learner is a distinct end user of AI (accessed through AI literacy and AI education), and they argue for renewed AI-literacy work that encourages reflection on learning itself — treating literacy as a route back into "AI as an analogy to human intelligence" research on how people learn, a strand the field largely abandoned. The [[icap-framework]] (Interactive → Constructive → Active → Passive) provides a useful taxonomy of cognitive engagement: students learn AI literacy best when they co-construct knowledge rather than passively receive information. Designers should select the mode that fits the learning goal and favor the deeper (constructive and interactive) modes where possible. Practical activities — designing prompts, evaluating outputs in groups, debating AI [[ethics]] — outperform lectures. **[[writing-education|Composition]] starts from lived experience — but concept depth needs scaffolding.** [[burriss-multimodal-composition-critical-ai-literacy-2026|Burriss et al. (2026)]] found that youth anchored abstract AI [[ethics]] principles in systems they encounter daily — electronic "hall passes," school-regulated laptops, plagiarism detectors — with one group explaining that AI in schools "was personal to all group members," and 15 of 18 end-of-unit survey responses reported that making the video changed their understanding of AI ethics. The authors are explicit about the limit: composing did not by itself guarantee conceptual grounding — "informed consent" was translated as "We gotta approve," and some terms (e.g., "failsafe") were never explained, prompting them to plan deeper conceptual work in later iterations. [[creativity|Creative]] production develops AI literacy best when explicit conceptual [[scaffolding]] accompanies the making rather than being assumed to emerge from it. **AI intuition as the experiential complement to literacy.** A recurring gap is that K-12 frameworks assume learners approach AI through declarative, rule-based knowledge, when in practice they first develop a practical "feel" for how AI responds — experimenting with prompts, observing behavior, adapting strategies — before they can articulate formal principles. [[ai-intuition-ai-literacy-k12-2026|Beau & Lazar (2026)]] formalize this as **AI intuition**: an experiential, inductive, often tacit understanding that develops through iterative interaction with AI systems and supports context-sensitive judgment under uncertainty. Their **dual framework** situates AI literacy (structured, largely static competencies) alongside AI intuition (a dynamic learning process), mapping both across the standard dimensions — understand concepts, use tools, evaluate critically, apply ethically, reflect — so learners combine conceptual clarity and [[guardrails]] (literacy) with the practical wisdom to trust-but-verify and decide when to disengage (intuition). Because intuition cultivates "expert observation" of AI (e.g. stress-testing prompts to find where a model fails), the authors argue it is a safeguard against [[cognitive-offloading|over-reliance]] and critical-thinking erosion; it is distinct from [[prompt-engineering]] (which optimizes outputs) in foregrounding *epistemic judgment*. This connects literacy to [[experiential-learning]], [[constructivist]], and developmentally grounded classroom practice, and points teacher preparation toward facilitating inductive exploration rather than only delivering concepts. **Motivation is a precondition, not just an outcome.** [[liang-ai-learning-motivation-sdt-2026|Liang et al. (2026)]] found, across 2,086 secondary students in a year-long AI curriculum, that students who transitioned into or remained in a *Self-Determined* [[motivation|motivational]] profile (high autonomy, competence, and relatedness need satisfaction) showed the **greatest AI-literacy gains**. AI literacy develops through sustained engagement, and that engagement is itself shaped by motivation and psychological-need support — so effective AI-literacy instruction should attend to learners' motivation, not only their skills. **Instructional emphasis and task openness reach performance through different needs.** [[yu-designing-ai-literacy-self-determination-2026|Yu, Lin and Chen (2026)]] randomized 320 undergraduates to four groups in a 2 x 2 experiment on an AIGC image-generation task scored with an objective rubric (ICC = 0.89) and found that thinking-based instruction, covering model limitations, critical evaluation and ethics, outperformed skills-based prompting instruction (M = 9.21 vs 7.69, F(1, 316) = 50.79, p < 0.001, partial eta squared = 0.138), with the wider margin on open-ended tasks (10.10 vs 8.00). The structural model shows the two design choices travel by different routes: instruction worked indirectly through autonomy (beta = 0.035) and competence (beta = 0.079), while task openness operated mainly through autonomy (beta = 0.029), competence was the strongest predictor of performance (beta = 0.358) and relatedness had no independent effect. What the instruction contains therefore matters more than whether learners get prompting practice, and it matters most when the task is open-ended. The variables surrounding AI literacy have since been mapped more broadly. [[ai-literacy-correlates-affective-behavioral-cognitive-2025|A 2025 systematic review of AI literacy's correlates]] synthesized 31 studies across 14 countries (N = 12,071) and found the most consistent associations in the affective and behavioral band: AI self-efficacy, positive AI attitudes, motivation and digital competence all moved with AI literacy, while AI anxiety and negative attitudes moved against it. Demographic variables barely registered, with age and socio-economic status correlating weakly or not at all. The review's caution is about instruments: the same studies that produced strong correlations from self-assessment also showed far weaker ones when AI literacy was tested rather than self-rated. [[hu-psychological-predictors-continued-chatgpt-use-2026|Hu (2026)]] adds an ordered account of what AI literacy predicts rather than what predicts it. Surveying 450 university students in mainland China who already use ChatGPT, the study found AI literacy related directly to continued use (beta = 0.16, p = 0.002) and indirectly along a serial path through [[trust]] and academic [[self-efficacy]] (indirect effect = 0.07, 95% CI [0.04, 0.10]), with the literacy to trust link the largest association in the model (beta = 0.50). [[anxiety-and-stress|AI anxiety]] moderated that first link (interaction beta = -0.25, simple slopes falling from 0.76 to 0.25 across the anxiety range), so the chain thinned for more anxious students and literacy instruction alone should not be assumed to reach them. The design is cross-sectional and the sample is restricted to existing users, so the ordering is a modeling assumption rather than an observed sequence. A core applied aim of AI literacy is [[reducing-ai-misuse]]: teaching students to use AI ethically and productively rather than substituting it for their own [[cognitive-offloading|cognitive work]]. Where AI literacy builds the *capacity* to evaluate and use AI critically, [[reducing-ai-misuse|reducing misuse]] is the behavioral and structural payoff — combining guardrailed tool design, [[assessment|assessment redesign]], and educative levers such as scaffolded think-first/AI-second sequences and prompting practice with deliberate [[feedback]]. The two concepts are mutually reinforcing: AI literacy supplies the critical dispositions that make misuse-reduction interventions durable, while misuse-reduction evidence (e.g. the [[ai-misuse-learning-harm|performance–learning gap]]) motivates why literacy must go beyond operational skill to critical judgment. AI literacy also needs developmentally appropriate forms for the youngest learners. [[ai-play-framework-early-childhood-2026|AI-Play]] translates AI literacy competencies into play-based, **unplugged** activities for Pre-K–K2 learners — organized around *AI Body* (AI as a system built from parts), *AI Food* (AI learns from examples), *AI Brain* (AI improves through patterns and feedback), and a Pre/Post-AI ethical lens — addressing a persistent lack of developmentally grounded AI literacy guidance for [[early-childhood-elementary-ai-education|early childhood]] and making AI literacy accessible to non-technical [[teacher-role|educators]] and [[parents-and-families|families]]. Complementing this, [[vahedian-children-attitudes-ai-chatbot-2026|Vahedian Movahed & Martin (2025)]] found children aged 6–14 broadly trusted an age-tailored chatbot as an information source and most modeled it as "a smart computer program that learns" (nascent [[machine-learning]] understanding), yet showed gaps in critical engagement and digital-safety awareness — evidence that young learners' trust can outpace their critical evaluation skills, making explicit [[trust-calibration]] and safety instruction a necessary part of AI literacy for children. Measurement is keeping pace with that developmental turn at the younger end. [[ai-literacy-self-assessment-questionnaire-primary-2025|Thianwan and Srikoon's 2025 validation of an AI literacy self-assessment questionnaire for upper primary students]] developed a 15-item instrument for Grades 4 to 6, organized around Learning About AI, Learning About How AI Works, and Learning for Life with AI, and confirmed a three-factor structure in two samples (n = 335 and n = 579) with an overall Cronbach's alpha of .934. The authors frame it as a formative diagnostic of perceived rather than demonstrated literacy, and note that children's self-assessment accuracy is constrained by metacognitive ability that is still developing. **Project-based learning as a delivery mechanism.** [[ai-literacy-course-satisfaction-pbl-scale-2026|Zhu & Kong (2026)]] developed and validated an AI project-based learning (AI-PBLS) scale and, in a Hong Kong sample of 1,027 secondary and university students (446 with complete data), used structural equation modeling to show that [[self-efficacy|empowerment]] in using AI for problem solving and AI [[ethics|ethical awareness]] jointly **mediate** the relationship between perceived [[project-based-learning]] and satisfaction with an AI literacy course. This positions PBL not merely as a delivery format but as a mechanism that builds learner confidence and ethical reasoning alongside competence, reinforcing the link between [[project-based-learning|PBL]] and meaningful AI literacy development. It also supplies a validated measurement instrument for future [[educational-measurement|measurement]] of AI literacy course experiences. **AI literacy as a core gap in [[conversational-ai]] frameworks.** The [[conversational-ai-agents-umbrella-review-2026|umbrella review of conversational AI agents]] (Ganguly et al. 2025, 34 reviews) identifies **limited AI literacy support** as a major gap in CAI frameworks, and its ethical-use roadmap makes foundational assessment (including strengthening AI literacy) the first pillar alongside participatory design, ethical-use guidelines, and continuous evaluation of cognitive impact. It further finds that AI-literacy, training, and awareness rank among the most-emphasized ethical directions in the CAI literature.([[conversational-ai-agents-umbrella-review-2026]]) **Effectiveness of AI literacy interventions.** A [[liu-ai-literacy-interventions-meta-analysis-2026|three-level meta-analysis of 59 studies]] (172 effect sizes, 7,211 participants) estimates a large overall effect of AI literacy interventions (g = 0.837, p < .001) — but with a wide prediction interval [−0.292, 1.966], so effectiveness varies considerably across settings. Interventions in East Asia and Europe outperformed those in North America, and knowledge-focused interventions outperformed those targeting skills, attitudes, or ethics. The authors argue AI literacy education should therefore move beyond knowledge toward skills, practices, ethics, and attitudes, supported by integrated and reflective [[pedagogy|pedagogies]] (project- and problem-based, inquiry-based, experiential) and GenAI-supported tools — a shift that aligns with the participatory, producer-oriented forms of [[computational-thinking]] described elsewhere in this knowledge base. Classroom evidence for deliberately *brief* designs is thinner, but it runs in the same direction and adds a behavior-level outcome that surveys cannot supply. [[clerc-ai-literacy-workshop-llm-regulation-2026|Clerc et al. (2026)]] gave 116 French students in grades 8–9 a two-hour workshop on how LLMs work and fail, paired with practice in predicting whether a prompt would elicit a usable answer, evaluating the response, and revising or asking again. Two days later, during six LLM-supported science problems with a control group, trained students accepted underspecified prompts less often (51.5% vs. 66.7%, OR = 0.47), asked follow-up questions after a weak response far more often (59.2% vs. 27.9%, d = 0.80), and judged answer correctness more sensitively to prompt quality (interaction OR = 2.52), with a modest score difference (11.38 vs. 10.29 of 20, p = .040). The workshop never rehearsed the test tasks, so what transferred was a regulatory stance rather than task familiarity — and the authors are explicit that two days is not durability. **AI-supported critical media literacy in elementary school.** Demir and Akar (2026) provide a concrete elementary-school demonstration of GenAI-supported literacy instruction: an 18-hour, 5E-model program for fourth-grade Turkish students in which ChatGPT and Grammarly were embedded phase-by-phase as pedagogical agents (ChatGPT for reflective questions and Q&A, Grammarly and Canva AI for content refinement, Padlet for [[peer-assessment|peer feedback]]), aligned to the Turkish Language and Social Studies curricula. The AI-supported group showed large gains in media reading (+3.50), writing (+1.67), and total media literacy (+5.17, all p < .01) with between-group effect sizes of Cohen's *d* = 1.12–1.31, while the control group advanced only modestly. [[qualitative-research|Qualitative]] analysis surfaced six domains of critical media literacy growth — digital self-protection and data privacy, purposeful and responsible media use, safe communication and boundary awareness, critical evaluation and misinformation awareness, online risk awareness, and media ethics/digital citizenship — evidence that developmentally appropriate, discipline-embedded GenAI use can build the critical-evaluation and ethical dimensions of AI literacy, not just operational skill. **Frameworks for structuring AI literacy.** Several recent contributions offer structured progressions for building AI literacy. **[[ukraine-ai-literacy-secondary-framework-2026|Marienko, Markova, and Semerikov (2026)]]** propose a five-level framework (Awareness, Application, Evaluation, Creation, Ethics) integrated with three paradigms of AI in education (AI-directed, AI-supported, AI-empowered), developed through a [[mixed-methods-research|mixed-methods]] study of Ukrainian secondary educators (national survey n = 2018; PD evaluation n = 1130). They found 84% of educators use AI but only 11% can identify specialized services beyond ChatGPT, and a professional-development intervention produced a 24% improvement in AI competence — evidence that targeted PD advances literacy beyond surface-level tool familiarity. The framework's grounding in [[constructivist]], connectivism, and TPACK connects to [[tpack]] and [[teacher-ai-competency]]. Complementing this, **[[science-integrated-ai-literacy-curriculum-dbr-2026|Moore et al. (2026)]]** used a two-year [[design-based-research|DBR]] process with a youth and AI-expert advisory board to design a science-integrated ML curriculum for high school youth, finding ML-knowledge gains in both cohorts (Cohort 2 M2−M1 = 0.175 vs Cohort 1 0.076) and greater gains among female and non-White participants — evidence that participatory, discipline-integrated design can advance both AI literacy and [[equity-in-ai-education]]. The construct is contested precisely where educators are concerned. A review of 34 studies on AI literacy in teacher education found that the field lacks consensus on what AI literacy means for professional educators as distinct from general digital literacy, and UNESCO's AI Competency Framework for Teachers specifies 15 competencies across five dimensions while functioning as a global policy reference rather than implementation guidance. The faculty in the [[ai-integration-instructional-design-collaboratory-2026|cross-institutional teacher preparation collaboratory]] treated AI literacy as an integrated dimension of [[professional-training|professional preparation]] rather than a separate technology skill, so candidates were expected to audit, compare and justify AI output within disciplinary coursework rather than to demonstrate tool familiarity. **Initial teacher preparation is where the structure is thinnest.** [[pinto-ai-initial-teacher-training-mathematics-review-2026|Pinto et al. (2026)]] reviewed 11 studies of AI in initial [[teacher-education|teacher training]] for pre-service primary mathematics teachers and found that nine were a single session or a handful of sessions embedded in existing courses, prompting was explicit training content in only some of them, and [[ethics]] appeared in just three, none of which addressed transparency, algorithmic bias or accountability. The risks they report are literacy-level failures: pre-service teachers failed to notice conceptual errors produced by ChatGPT, accepted generated content uncritically, and delegated more problem solving to the tool as the tasks became harder. Their recommendation is to build [[teacher-ai-competency|AI competency]] across the whole training sequence, from a preparatory phase into the practicum, rather than in an isolated module. **AI-interaction literacy: the interactional dimension.** [[brunnstrom-ai-interaction-literacy-srl-2026|Brunnström and Palmqvist (2026)]] propose a narrower, interactional competence — the ability to *steer, evaluate, and learn from* iterative interaction with GenAI — as a specific enactment of the applicational, evaluative, and integrational dimensions in broader frameworks such as [[ai-literacy-heptagon-2026|the AI Literacy Heptagon]]. Their reflective demonstration shows what it consists of in practice: recognizing that a fluent answer is pitched above one's own schema, requesting simplification, narrowing scope, and redirecting the system toward focused practice. Two design consequences follow — the skill is unevenly distributed, so unguided use may advantage already-confident students and widen gaps ([[equity-in-ai-education]]), and it has to be taught explicitly rather than assumed, with teachers covering how to formulate productive prompts and when to stop using the tool ([[self-regulated-learning]]). ### Critical AI literacy: beyond skills to power and resistance A distinct strand of the knowledge base treats AI literacy not merely as skills or critical evaluation but as a *critical and political practice* that interrogates power, authority, and whose knowledge counts. This connects AI literacy to [[critical-pedagogy]] and [[equity-in-ai-education]]: - **"Resisting AI" as a literacy stance.** Critical AI Literacy (CAIL) can encompass *resisting AI* — refusing the inevitability and techsolutionism of dominant discourse, and cultivating collective agency through dialogic, collaborative pedagogies.([[li-mroziak-reorienting-critical-ai-literacy]]) This positions education as a space where communities imagine and build alternative futures, rather than merely adapt to a given technological order. - **Community-based epistemologies.** Community-based AI learning grounds AI engagement in learners' lived and community-based ways of knowing, redistributing AI's epistemic authority through epistemic fine-tuning, redistribution of authority, and situated discernment.([[ojeda-ramirez-community-based-ai-learning]]) - **Critical and feminist frames.** Critical, feminist scholarship argues AI literacy should be framed within pedagogies of justice, resistance, and cultural [[sustainability]] — asking whose knowledge AI produces, who benefits, and who is accountable when systems fail.([[avraamidou-ai-colonization-science-education]]) - **Situated curriculum devices.** AI literacy can be built through [[situated-learning|situated]], active teaching instruments (e.g., Episodes of Situated Learning) that develop AI competencies across levels rather than through abstract instruction.([[panciroli-ai-literacy-episodes-situated-learning]]) - **Creative composition as critical AI literacy pedagogy.** Rather than treating AI literacy as technical knowledge or individual competencies, design opportunities for learners to compose messages about AI for real audiences — a [[multimodal]] product such as a public service announcement simultaneously demonstrates and communicates critical competence. In [[burriss-multimodal-composition-critical-ai-literacy-2026|Burriss et al.'s (2026)]] classroom study, 22 eleventh graders chose their own AI [[ethics]] issues and made video PSAs about school surveillance, [[privacy|informed consent]], and algorithmic accusation; across all seven films harm emerged from human–machine entanglement rather than a villainous tool, the plots still closed on hope and [[agency]], and the authors position [[storytelling-in-education|storytelling]] and advocacy — resistant to traditional assessment — as core to [[critical-pedagogy|critical AI literacy pedagogy]] that centers [[learner-identity|student identity]], choice, and voice.([[burriss-multimodal-composition-critical-ai-literacy-2026]]) - **Regulatory competence, not acceptance.** Kim (2026) reframes AI literacy in [[higher-ed]] as *[[regulation|regulatory competence]]* — an enacted practice of verifying, revising, selectively adopting, or rejecting AI output during academic work — showing that evaluative capacity and ethical awareness (not mere willingness to use AI) predict active, [[student-engagement|critical engagement]].([[ai-anxiety-strategic-regulation-writing-2026]]) - **Recognizing [[ai-sycophancy|sycophancy]] as a literacy skill.** A core evaluative competency is recognizing when an AI is *agreeing* with the learner versus being correct. [[contextual-sycophancy-ai-literacy|Contextual sycophancy]] shows AI literacy and prompting training reduce — but do not eliminate — sycophantic mirroring of user errors, and [[sycophantic-ai-social-interaction-2026|sycophantic AI]] is preferred by users precisely because it makes them feel understood. AI literacy must therefore teach learners to detect agreement-for-its-own-sake and to value corrective friction, connecting to [[trust-calibration]] and [[reducing-ai-misuse]]. These critical strands complement the operational and cognitive dimensions of AI literacy: where the latter ask "can the learner use and evaluate AI?", critical AI literacy asks "does the learner understand and challenge the power structures AI embodies?" ### Whose AI skills count? The instructor–employer framing divide A central open question in AI literacy is *which* skills matter and for whom. An Ithaka S+R study comparing how US instructors and employers prioritize the 26 skills of the HiBob AI Skills Framework found they agree on the importance of only one (setting realistic expectations for AI-augmented work). Instructors weight a **critical, responsible-use orientation** — recognizing AI's limits, [[explainable-ai|transparency]] and attribution, human accountability, proactive output review — which aligns with academic values of attribution, review, and information literacy. Employers weight **productivity-oriented skills** — workflow evaluation and redesign, automation, and [[human-ai-collaboration|human–AI teaming]] — that reflect team-based workplace efficiency. Because only three of 26 skills are taught by half or more instructors, and those taught skew toward the critical-use categories, the report identifies a concrete **AI skills gap**: whole categories employers value are neither prioritized nor taught in college curricula.([[ithaka-sr-ai-skills-college-graduates-2026]]) This divide frames AI literacy as a contested construct — critical-use literacy for academic settings versus workflow-integration literacy for employment — a tension relevant to [[framing-ai-use-for-students]], [[curriculum-design]], and [[professional-training]]. ### Designing AI literacy interventions The knowledge base's frameworks and empirical studies converge on a set of practical guidelines for educators and [[stakeholders|instructional designers]] building AI literacy interventions: **1. Use a structured competency framework as scaffolding, not a checklist.** Mature frameworks give designers a shared vocabulary and a developmentally sequenced target. The [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|SAIL framework]] organizes AI literacy into three domains (AI Concepts; Application and Technical Skills; AI Digital Citizenship) across four scaffolded levels — Understand and Explore → Apply and Integrate → Evaluate and Create → AI++. The [[ai-literacy-heptagon-2026|AI Literacy Heptagon]] cross-cuts seven dimensions (technical knowledge, application, critical thinking, ethics, social impact, integration, legal/regulatory) with four Bloom-aligned proficiency levels, and stresses that emphasis must be **adapted to disciplinary context** — technical programs weight application, [[humanities-education|humanities]] programs weight ethical and social-impact reasoning.([[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl]])([[ai-literacy-heptagon-2026]]) **2. Engage learners across ICAP modes.** Applying the [[icap-framework]], effective instruction gives learners opportunities to engage at multiple cognitive levels — passive exposure (AI concept lectures), active manipulation (hands-on tool use), constructive generation (creating AI artifacts, self-explaining), and interactive dialogue (collaborative [[problem-solving]] with peers and AI) — selecting the mode that fits the learning goal. A [[meta-analysis-systematic-review|systematic review]] found successful collaborative AI-literacy interventions spanned all four ICAP modes.([[hingle-collaborative-ai-literacy-2025]]) **3. Assess demonstrated competence, not self-perception.** Self-reported AI literacy diverges sharply from measured performance — teachers overestimate their AI skills by ~40%, and performance-based measures correlate with classroom AI integration far better than self-reports (r≈0.72 vs 0.31). Design interventions around performance-based assessment and calibration rather than confidence surveys, and use diagnostic profiles (overestimators vs. true novices) to target support.([[ai-literacy-assessment-misalignment]]) **4. Build metacognitive and critical dispositions, not just operational skill.** AI literacy is better understood as a [[metacognition|metacognitive social practice]] than a skills checklist: because LLMs are probabilistic and opaque, learners must monitor and adjust their strategies, cultivate scientific skepticism, and interrogate how algorithms shape knowledge production — not merely learn to operate tools. Participatory co-design and experimental, project-based spaces (not one-off tool training) are where this awareness grows.([[metacognitive-ai-literacy-beyond-skills-gap-2026]]) **5. Embed literacy in the discipline and make it sustained.** Movement toward higher stages of AI literacy (from uncritical use → informed use → critical evaluation → improvement) is most visible when experiences are **sustained and discipline-embedded** rather than delivered as standalone workshops. Design for repeated, contextual practice within authentic coursework.([[ai-literacy-continuum-higher-education]]) A concrete model of discipline-embedded critical literacy is the task-based taxonomy of [[dierickx-taxonomy-llm-tasks-critical-ai-literacy-journalism-2026|Dierickx et al. (2026)]] for journalism: it maps LLM-supported tasks across the four stages of the news workflow (newsgathering, sensemaking, editing, publication/distribution), each tied to a baseline prompt and a risk-and-mitigation strategy. By treating task definition and prompting as a [[situated-learning|situated]] form of professional judgment — not a neutral technical skill — it turns prompting itself into a vehicle for critical AI literacy (bias, [[hallucination-risk|hallucination]], overreliance, and the enduring value of [[human-in-the-loop-ai|human editorial oversight]]), and its underlying logic transfers to other knowledge-intensive professions (law, [[medical-education|medicine]], public policy). Complementary to such [[discipline-specific-aied|discipline-specific]] taxonomies, the [[dohn-boundary-object-classifying-genai-learning-activities-2026|Dohn et al. (2026) taxonomy]] classifies GenAI learning activities along six dimensions, of which **Epistemic Engagement** (understanding / using / critiquing / constructing GenAI) directly operationalizes AI literacy as the depth of learners' cognitive relationship with the technology — from passive exposure to active critique and construction. **6. Treat equity and the digital divide as design constraints.** AI literacy is a mechanism for addressing the three-level digital divide (access, skills, outcomes): closing the device gap is insufficient unless skills and critical use are built so benefits distribute fairly. Interventions should plan explicitly for learners who enter with less prior AI access, and incorporate cultural and governance perspectives rather than treating literacy as culture-neutral.([[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl]])([[digital-divide]]) **7. Pair literacy with misuse-reduction levers.** Because AI misuse actively harms durable learning, literacy instruction should be coupled with the structural and educative levers documented under [[reducing-ai-misuse]] — guardrailed "hint-not-answer" tool design, assessment redesign, and scaffolded think-first/AI-second sequences with deliberate prompting practice. **8. Start from learners' actual entry point.** Learners enter with distinct orientations — avoidance driven by fear, mistrust, or lack of access, versus uncritical reliance that masks misunderstanding. A diagnostic, stage-based approach (rather than a uniform curriculum) lets designers meet students where they are and move them toward critical, responsible engagement.([[ai-literacy-continuum-higher-education]]) **9. Extend provision beyond the audiences formal education already reaches.** The [[ai-literacies-young-adults-2025|AI Literacies framework for public service media]] argues from a landscape review of more than 40 frameworks and 35 expert interviews that provision has tilted toward technical and functional skills, and that AI literacies overlap and should be connected to digital, media and information literacy rather than taught as a standalone subject. Its structure — six competency areas, five values and three progression levels (*Understanding and Applying*, *Analysing and Evaluating*, *Synthesising and Specialising*) with an assessment guide — is a usable template for providers outside formal schooling, and its equity argument is a design instruction: reaching young people who are digitally or otherwise marginalized takes deliberate partnerships, not universal publication.([[ai-literacies-young-adults-2025]]) ### Measuring AI literacy A distinct research thread treats AI literacy not only as a target for instruction but as a construct to be measured. The knowledge base's assessment strand distinguishes **self-reported** from **performance-based** literacy: self-reports diverge sharply from demonstrated competence (teachers overestimate by ~40%), and performance-based measures predict classroom AI integration far better than confidence surveys (r≈0.72 vs 0.31). Intervention work reproduces that divergence at the level of behavior: in [[clerc-ai-literacy-workshop-llm-regulation-2026|Clerc et al. (2026)]], neither GenAI attitudes nor a general metacognitive-awareness scale predicted students' regulation of LLM interaction or their final task scores (r = .01 and r = .04, both non-significant), while the behaviors the workshop changed — rejecting underspecified prompts, judging answer correctness, asking a follow-up — did track answer quality. Validated instruments are emerging to close this gap — the [[jin-glat-genai-literacy-assessment|GLAT]] provides a psychometrically validated generative-AI literacy assessment, and diagnostic profiles (overestimators vs. true novices) let designers target support where it is needed. For design and research, this ties AI literacy to [[educational-measurement]] and to [[assessment]] broadly: a literacy framework is only as useful as the instruments used to track growth, and stage-based continua require reliable measurement to place learners along them. [[zhi-modeling-measuring-graduate-genai-literacy-2026|Zhi, Yang and Huang (2026)]] build a graduate-specific model on Marzano's taxonomy: grounded-theory analysis of interviews with 14 professors condensed 329 raw labels into 96 concepts, 15 categories and five dimensions (cognitive foundation, operational skills, higher-order thinking, metacognitive reflection, ethical responsibility), then operationalized them as a 15-item Likert scale whose five factors emerged in exploratory factor analysis (83.14% of cumulative variance) and held in confirmatory factor analysis on a second subsample (CFI = 0.934, RMSEA = 0.083), with reliability from 0.796 to 0.842 across 308 valid questionnaires. A related question is *what* the instruments can measure at all. [[burriss-multimodal-composition-critical-ai-literacy-2026|Burriss et al. (2026)]] note that existing AI literacy scales and competency frameworks assume individually measurable performance and so structurally exclude collaborative, creative expression — their unit's evidence was [[multimodal]] film artifacts, reflections, and civic discourse rather than a [[summative-assessment|summative]] score, and the authors argue such evidence can *complement* rather than replace conventional measures. Broadening the construct may therefore require broadening the admissible evidence, not merely adding modality-rich items to existing scales. A complementary strand measures how learners and teachers receive AI literacy *materials* rather than their literacy itself. [[age-tiered-ai-literacy-guidebooks-2026|Wang, Chuang and Wu (2026)]] had 794 students and 37 teachers rate two guidebook editions built to UNESCO's age threshold (ages 9-12 and 13-18) after roughly 30 minutes of guided classroom exposure. Acceptance held a four-factor structure (Performance Expectancy, Effort Expectancy, Perceived Playfulness, Behavioral Intention) with measurement invariance supported across the two student editions, younger learners scored higher on all four constructs, and perceived playfulness carried the largest association with intention in both cohorts. The authors are explicit about what this is not: perceived acceptance of an age-tiered resource is not AI literacy achievement, [[ethics|ethical]] reasoning, adoption or sustained use, and material-level acceptance should not be read as evidence that literacy improved. Measurement also extends to the educators who mediate learners' engagement with AI. Most AI-literacy assessments target students or general users, leaving a gap in [[teacher-education|teacher]] education — a gap the [[language-teachers-ai-literacy-edai-2026|Teachers' AI Literacy Scale (TAILS)]] addresses: grounded in the ED-AI framework with six dimensions (knowledge, evaluation, collaboration, contextualization, autonomy, and ethics), it was validated through exploratory and confirmatory factor analysis with [[language-learning|preservice language teachers]]. Such instruments support measuring and developing the AI literacy of the educators who mediate learners' engagement with AI. On the student side, [[genai-assessment-literacy-scale-2026|Nie et al. (2026)]] develop and validate a Generative AI Assessment Literacy Scale (GAA-LS) for higher-education students — an 18-item, five-factor instrument whose scores track [[feedback]] engagement, [[academic-integrity]], and responsible AI use, with an indirect path to integrity running through feedback — evidence that assessment-specific AI literacy is measurable and ties to integrity behavior. What that teacher-facing instrument corpus actually contains has since been audited: [[assessing-teachers-ai-literacy-measurement-tools-2026|Zainal, Mohd Matore and Maat's 2026 review of teacher AI literacy measurement tools]] appraised 33 instruments published between 2019 and 2025 and found that 31 (93.9%) were self-report scales of perceived confidence, only two tested knowledge objectively, and none used performance-based tasks, while fairness evidence was the weakest quality domain, with only five instruments reporting measurement invariance or differential item functioning. The 2026 update of that measurement literature reorganizes it into four domains — knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents — and reports that no validated individual-level instrument in the corpus covers the full combination of scope, permissions, recovery, state isolation, independent review and evidence-based closure that [[agentic-ai|agentic]] tool use demands. Its pooled subjective–objective correlation across three same-sample effects was r = .055, consistent with the divergence described above rather than with self-ratings as a usable proxy ([[competent-generative-ai-use-measures-review-2026|Verí (2026)]]). **The measurement landscape is itself disordered, and now partly mappable.** [[ai-literacy-measurement-conceptual-landscape-llm-2026|He, Zhang, Wang and Ji (2026)]] analyzed the AI literacy instrument corpus with an LLM-based coding pipeline and found the field measuring many constructs under one label — a jangle problem, with one construct carrying several names (Behavioral Commitment appears in eight candidate pairs, Intrinsic Motivation in five), alongside candidate jingle cases where instruments share a label but not their item content. Their semantic-similarity recovery of item and construct structure correlated only moderately with instruments' own reported reliabilities (r = 0.49 at item level against 0.45 and 0.35 at construct level), and they offer the method as screening rather than adjudication. The consequence for anyone selecting a measure is that convergent claims across studies are not safe to assume: two instruments with the same name may not operationalize the same construct. **What those instruments can support has now been audited, and the audit is sobering.** [[ai-literacy-instrument-development-systematic-review-2026|Jin, Gašević, Martinez-Maldonado and Yan (2026)]] screened 8,056 records down to 58 studies covering 47 unique instruments, appraised with COSMIN against a construct whose instrument-development window compressed into three years (two instruments before 2023, 28 in 2025 alone). The corpus remains 37-of-47 [[self-report-measures|self-report]], and the evidence clusters tightly around internal structure: structural validity was sufficient in 32 of 58 studies and internal consistency in 32 with none rated insufficient, but [[assessment-validity|construct validity]] was reported by only 15 studies (12 sufficient), criterion validity by one, measurement [[assessment|invariance]] by five (all sufficient), and just 2 of 58 earned a sufficient overall content-validity rating — comprehensiveness was insufficient in 53 despite 43 studies reporting cognitive interviewing. The pattern is that [[educational-measurement|measurement]] practice looks robust exactly where it is cheapest to demonstrate and thin where evidence must come from outside the instrument, so a clean factor structure and a high coefficient (which can reflect item redundancy as much as construct representation) should not be read as demonstrated capability. The authors' prescription is consolidation rather than proliferation: refine existing scales, test invariance whenever one crosses a language or an educational stage, and add criterion evidence by relating scores to performance tasks. **One of the exceptions is built to compare groups.** [[gails-generative-ai-literacy-scale-2026|Zhang et al. (2026)]] developed the Generative AI Literacy Scale (GAILS) — 43 items reduced to 34 through a five-expert Delphi review and a seven-person pilot, then validated in 341 North American adults — and tested scalar [[assessment|invariance]], finding that it held across female and male respondents and across student and workforce groups. That is what licenses the mean comparisons the instrument reports — men scoring modestly higher on Adaptive Operational Skills, students higher on Adaptive Operational Skills and Responsible GenAI Literacy — as competence differences rather than differential item functioning, and it is what makes the scale usable across [[higher-ed]] and workplace settings rather than student-only. Its own evidence is not unblemished: [[educational-measurement|fit]] was mixed (CFI = .967 and SRMR = .058 inside conventional cutoffs, RMSEA = .087 above them), total-scale α = .973 sits alongside one failed Fornell–Larcker discriminant comparison (Factor 1 √AVE = .828 against r = .873 with Factor 3), and, like almost everything in the corpus, it measures perceived competence with no behavioral criterion. A validated scale in this field is a starting point for the invariance and criterion work the review calls for, not a finished [[benchmark]]. ### Connections across the knowledge base AI literacy intersects with [[intelligent-tutoring|AI Tutoring]] (understanding when and how AI tutors are effective), [[teacher-ai-competency]] (educator preparedness), [[academic-integrity]] (knowing what constitutes appropriate AI use), and [[ai-education]] broadly. It is both a prerequisite for effective AI use and an outcome of well-designed AI integration — students learn AI literacy BY using AI critically, not just by learning ABOUT AI. AI literacy is **double-edged** for overreliance: [[student-dependency-on-ai-literacy-self-efficacy-2026|Maizel et al. (2026)]] found the skill-based dimensions of AI literacy (using/understanding, detecting) were *positively* associated with reported AI dependency, while AI [[self-efficacy]] and academic confidence were negatively associated — so technical AI-literacy training, absent self-efficacy and [[self-regulated-learning]] scaffolds, can increase dependency. AI literacy here becomes an enabling capacity whose *direction* depends on complementary motivational resources. - **Verification habits pay off only through regulation.** In [[davor-ai-supported-learning-higher-order-outcomes-2026|Davor, Larbi and Boateng (2026)]]'s survey of 533 university students in Ghana, AI verification literacy had no significant direct effect on [[critical-thinking|critical thinking]] (beta = .076, p = .090) or technical [[problem-solving]] (beta = .043, p = .385) and mattered only through [[metacognition|metacognitive self-regulation]] (beta = .167, p < .001), a full mediation the authors call metacognitive activation. AI task scaffolding predicted both outcomes (beta = .185 and .170), while [[cognitive-offloading|cognitive offloading]] tendency predicted them negatively (beta = -.240 and -.312) and depressed regulation as well (beta = -.294). Teaching students to check AI output is therefore not enough on its own: evaluative habits need planning, monitoring and reflection built into the task. - **Critique of AI output as a literacy practice:** [[pedagogy-ai-mistakes|Hosseini (2026)]] treats evaluating AI-generated errors as a core AI-literacy skill, using failure-mode analysis and iterative prompt refinement in a database design course. The study found students overestimated their AI abilities (self-reported literacy weakly, negatively correlated with objective competency), and that critique-based learning strengthened calibration. - **Socialist humanist AI literacy (2026):** A literature review critiques compliance-oriented AI literacy and proposes a socialist-humanist framing of asynchronous AI literacy and fair use in higher education, linking the historical digital divide to modern AI literacy and calling for approaches that serve human flourishing and equity rather than mechanical policy compliance ([[mechanical-compliance-human-flourishing-ai-literacy-2026]]). - **Cheap, light-touch warnings blunt AI persuasion — without costing general trust.** [[ai-literacy-warning-political-persuasion-2026|Orchinik and Rand (2026)]] preregistered two experiments (total N = 3,208 US adults) in which participants conversed with an [[llm]] instructed to shift their views on political topics. A single brief warning — that models can be prompted to persuade and may present information selectively — cut belief change by roughly one half (-48.1%, 95% CI [-59.5%, -36.8]) relative to control, and adding specific persuasion-technique warnings produced no further benefit. The property that matters for teaching is on the other side of the effect: general trust in [[generative-ai]] did not fall, so the intervention builds [[trust-calibration|calibrated trust]] rather than blanket skepticism. It is a one-paragraph, no-facilitation intervention, which is a rare cost profile for an AI literacy design. - **Learner governance of AI matters more than the design of the tool.** [[ai-literacy-tool-design-programming-education-2026|Azimi (2026)]] randomized 33 students in a master's data-analytics course between a scaffolded AI Study Coach embedded in the notebooks (n = 16) and unrestricted use of any [[generative-ai]] tools they chose (n = 17) for seven weeks. Assignment performance and concept-inventory gains were indistinguishable; the Coach condition reported higher [[self-efficacy|confidence]] instead. What separated students was AI literacy in practice: those who had formulated their own rules for when to use AI scored higher in both conditions, and the students with the deepest model understanding — every one of them self-taught — prompted most deliberately. The design implication runs against the control reflex: teach the [[self-regulated-learning|self-regulatory]] and model-understanding components of AI literacy rather than constrain tools. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[generative-ai]] — the technology AI literacy targets - [[llm]] — the systems at the heart of AI literacy - [[critical-thinking]] — core evaluative disposition - [[prompt-engineering]] — core practical competence - [[metacognition]] — literacy as metacognitive social practice - [[cognitive-offloading]] — the over-reliance risk literacy counters - [[reducing-ai-misuse]] — literacy's behavioral payoff - [[academic-integrity]] — knowing what constitutes appropriate AI use - [[self-assessment]] — judging your own work and skill, as technique and as measure - [[trust-calibration]] — calibrating appropriate trust - [[ai-sycophancy]] — literacy skill of detecting agreement - [[equity-in-ai-education]] — fair distribution of literacy - [[digital-divide]] — the access/skills/outcomes gap - [[ethics]] — ethical awareness dimension - [[teacher-ai-competency]] — educator preparedness - [[educational-development]] — building educator literacy - [[k-12]] — school-level literacy - [[higher-ed]] — university-level literacy - [[ai-education]] — the broader field ## Connected Articles - [[ai-literacies-young-adults-2025]] — Six competency areas, five values and three progression levels for public service media - [[powerful-learning-with-emerging-technology-2025]] — Evaluating emerging technology as a skill-building goal - [[typology-generative-ai-tools-education-2026]] — Typology of Generative AI Tools for Education - [[ai-literacy-heptagon-2026]] — The AI Literacy Heptagon - [[ai-literacy-continuum-higher-education]] — A Practical Five-Stage Continuum for AI Literacy - [[ai-literacy-measurement-conceptual-landscape-llm-2026]] — Mapping the AI literacy instrument corpus: 55 constructs, jangle and jingle pairs, LLM-based coding - [[ai-literacy-instrument-development-systematic-review-2026]] — COSMIN appraisal of 58 studies and 47 AI literacy instruments: validity evidence clusters around internal structure while criterion, content and invariance evidence stays largely missing (Jin et al. 2026) - [[gails-generative-ai-literacy-scale-2026]] — GAILS: 34-item GenAI literacy scale validated in 341 adults, scalar-invariant across sex and student/workforce groups (Zhang et al. 2026) - [[jin-glat-genai-literacy-assessment]] — GLAT: a validated generative AI literacy assessment test - [[ai-literacy-assessment-misalignment]] — AI Literacy Assessment: Self-Reported vs Performance Misalignment - [[genai-assessment-literacy-scale-2026]] — GAA-LS: validated Generative AI Assessment Literacy Scale for higher-ed students (Nie et al. 2026) - [[competent-generative-ai-use-measures-review-2026]] — Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use - [[ai-literacy-correlates-affective-behavioral-cognitive-2025]] — systematic review of what AI literacy correlates with - [[assessing-teachers-ai-literacy-measurement-tools-2026]] — systematic review of the tools that measure teacher AI literacy - [[ai-literacy-self-assessment-questionnaire-primary-2025]] — self-assessment questionnaire for upper primary students - [[metacognitive-ai-literacy-beyond-skills-gap-2026]] — AI literacy as a metacognitive social practice - [[liu-ai-literacy-interventions-meta-analysis-2026]] — Meta-analysis of AI literacy intervention effects - [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl]] — The Scaffolded AI Literacy (SAIL) Framework - [[age-tiered-ai-literacy-guidebooks-2026]] — Age-tiered AI literacy guidebooks evaluated with 794 students and 37 teachers: four-factor acceptance structure invariant across the 9-12 and 13-18 editions, with younger learners higher on every construct - [[li-mroziak-reorienting-critical-ai-literacy]] — Critical AI literacy: power, resistance, agency - [[miles-prompt-literacy-human-centered-genai-framework-2026]] — Prompt literacy as a foundational literacy distinct from prompt engineering: the five-phase human-centered GenAI engagement model (Miles, Haber-Curran & Arar 2026) - [[contextual-sycophancy-ai-literacy]] — Contextual sycophancy as an AI literacy intervention - [[student-dependency-on-ai-literacy-self-efficacy-2026]] — AI literacy, self-efficacy and dependency - [[clerc-ai-literacy-workshop-llm-regulation-2026]] — a two-hour workshop changed middle-school students' regulation of LLM interaction, while self-reports predicted nothing (Clerc et al. 2026) - [[ai-literacy-sdg-governance-framework-2026]] — AI literacy as a governance capacity for sustainable development: the AIRE Taxonomy and AI–SDG Nexus (Islam, Morshed & Islam 2026) - [[ukraine-ai-literacy-secondary-framework-2026]] — Five-level AI literacy framework for Ukrainian secondary educators (Marienko et al. 2026) - [[niri-steam-ai-literacy-review-2026]] — STEAM education for AI literacy: systematic review - [[hingle-collaborative-ai-literacy-2025]] — Collaborative AI Literacy Framework - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback - [[learning-analytics-genai-secondary-writing-2026]] — Using Learning Analytics to Support Secondary School Students' Writing with Generative AI - [[obyrne-co-constructing-ai-boundaries-agency-judgment-2026]] — reframes AI literacy as judging which interpretive work should not be delegated - [[davor-ai-supported-learning-higher-order-outcomes-2026]] — AI verification literacy mattered only through metacognitive self-regulation, while cognitive offloading predicted lower critical thinking and problem solving (Davor et al. 2026) - [[yu-designing-ai-literacy-self-determination-2026]] — Thinking-based instruction beat skills-based instruction on an objectively rated task, with autonomy and competence carrying the effect (Yu, Lin & Chen 2026) - [[hu-psychological-predictors-continued-chatgpt-use-2026]] — AI literacy to continued ChatGPT use runs through trust and academic self-efficacy, and AI anxiety thins the first link (Hu 2026) - [[zhi-modeling-measuring-graduate-genai-literacy-2026]] — Five-dimension Marzano-grounded model and 15-item self-report scale for graduate GenAI literacy (Zhi et al. 2026) - [[pinto-ai-initial-teacher-training-mathematics-review-2026]] — Eleven studies of AI in initial mathematics teacher training, mostly short-term, with ethics addressed in only three (Pinto et al. 2026) - [[vega-baudrit-genai-university-chemistry-education-review-2026]] — Verification-centered GenAI integration in university chemistry, with representational translation as the core AI literacy demand (Vega-Baudrit & Rivera Alvarez 2026) --- ## [Agentic AI](https://edtechdev.github.io/aied/concepts/agentic-ai/) > **Agentic [[ai-education]]** — AI systems that autonomously plan, execute, and adapt multi-step workflows to achieve learning goals, going beyond single-turn Q&A to act as persistent, goal-directed collaborators: [[intelligent-tutoring|AI tutors]] that scaffold over extended interactions, multi-agent systems that orchestrate instructional designs, and agents that co-regulate learning. This paradigm shift from a prompt-responding tool to an active collaborator carries both promise and risk: agentic AI can personalize and deepen learning, but it also threatens [[agency]], [[cognitive-offloading|cognitive effort]], and control. The knowledge base's [[agentic-ai-education-scoping-review|scoping review]], [[tool-invariant-framework-agentic-ai|tool-invariant framework]], and [[agentic-ai-pedagogical-best-practice-2026|pedagogical best-practice]] articles examine this tension. ## Questions to Consider - Agentic AI doesn't just answer questions — it plans, executes, and adapts multi-step workflows toward a goal, acting as a persistent collaborator. How is learning with a proactive agent different from learning with a tool you have to prompt? - The field draws a line between "conversational" and "agentic" AI (a system qualifies only if it meets several criteria such as planning, memory, autonomy, and goal-directed action). Where is that line in practice — and does calling every chatbot "agentic" hide more than it reveals? - The more an agent automates, the less cognitive work the learner does. Where is the line between an AI that scaffolds your learning and one that does your learning for you? - A [[meta-analysis-systematic-review|scoping review]] found only 29% of agentic AI studies grounded their systems in educational theory. If most systems aren't theory-based, what should make you skeptical when evaluating an 'intelligent' tutoring agent? - Multi-agent systems orchestrate specialized agents with distinct roles. When several agents work together in a classroom, who is accountable — and where should a human intervene? - The field's central tension is personalization versus learner agency and cognitive effort. If a tutor becomes so good at adapting that you never have to struggle, what learning are you actually getting? - Where on the Copilot-to-Autopilot spectrum should a given task sit — and does moving a task from "agent proposes, learner disposes" to "agent owns it" ever serve learning rather than just efficiency? - Hybrid agents grounded in established design theory outperformed pure prompting. Why might a theoretically-grounded system beat raw prompt-engineering — and what does that say about how an agent's 'smartness' is measured? ## Introduction Agentic AI refers to artificial intelligence systems that can autonomously plan, execute, and adapt multi-step workflows to achieve learning goals — going beyond single-turn question-answering to act as persistent, goal-directed collaborators in educational contexts. In education, agentic AI manifests as AI tutors that scaffold learning over extended interactions, multi-agent systems that orchestrate complex [[learning-design|instructional designs]], and autonomous agents that adapt their [[pedagogy|pedagogical]] strategies based on learner needs. This emerging paradigm shifts AI from a tool that responds to prompts to a collaborator that actively guides, adapts, and co-regulates learning processes. ## Defining and classifying agentic AI [[kostopoulos-agentic-ai-education-2025|Kostopoulos et al. (2025)]] supply an operational definition the field otherwise lacks. They propose a **six-criteria checklist** — a system counts as agentic if it meets at least four: autonomy (action independent of continuous human intervention), reasoning/planning, memory/context-awareness, goal-directed action toward [[learning-gains|learning outcomes]], adaptability, and dynamic collaboration/initiative. The ≥4 threshold deliberately **excludes reactive chatbots** (a static FAQ bot without planning or persistence does not qualify) while accommodating diverse architectures. They also organize the space along three axes: **pedagogical role** (tutor, learning coach/mentor, companion, instructor's assistant, [[curriculum-design|curriculum]] planner), **autonomy level** (reactive → adaptive → proactive → collaborative), and **embodiment** (text-based, avatar/graphical, [[embodied-learning|embodied]]/robotic). This taxonomy — particularly the autonomy spectrum and the checklist's exclusion of reactive tools — gives researchers and designers shared vocabulary for classifying agentic systems and distinguishing genuinely agentic from merely conversational AI. A worked discriminator makes the line concrete. A static FAQ chatbot that can only answer a fixed set of questions meets **zero** criteria (no planning, no persistence, no initiative) and is plainly not agentic. A [[conversational-ai|conversational]] tutor that remembers the current session but never acts unless prompted, holds no cross-session [[student-modeling|learner model]], and cannot set sub-goals may meet only one or two (memory, some reasoning) — conversational, not agentic. By contrast, a tutor that plans multi-turn lessons, stores learner progress in a persistent profile, **proactively** fires a hint when a learner stalls, and re-plans the next step based on that profile satisfies planning, memory, autonomy, and goal-directed interaction — at least four criteria, so it qualifies as agentic. The value of running this test is not pedantry: labeling every LLM chat interface "agentic" blurs the very design question — what the system initiates versus what the learner must initiate — that determines whether it scaffolds or supplants learning. **Bounded agency as an educational design stance.** [[ilieva-agentic-genai-higher-education-2026|Ilieva et al.'s (2026)]] AGAI-HE framework accepts the capability list that definitions such as the six-criteria checklist describe, then deliberately constrains it: goals, roles, data sources, tools, checkpoints, stopping conditions, and final decisions are defined or approved by educators, and each agentic function must trace to a learning requirement, assessment purpose, or governance control. The framework also draws a line between agentic orchestration and advanced [[prompt-engineering|prompting]] — an agentic workflow preserves task state, allocates functions, checks completion conditions, returns to earlier stages when evidence is insufficient, and records material decisions. Notably, its exploratory perception study with 130 higher-education students found no significant difference between agent-supported and chatbot-supported learning, which the authors read as evidence that agentic *capability* does not by itself produce a perceived pedagogical advantage ([[human-in-the-loop-ai|human oversight]]). ## The field: rapid expansion and current shape The knowledge base's [[agentic-ai-education-scoping-review|scoping review]] — the most comprehensive synthesis of the field to date, mapping **474 studies (2020–2026)** — documents a field that has grown **explosively since 2025**, but whose literature is still dominated by conference papers concentrated in [[higher-ed]], [[stem-education|STEM disciplines]], and text-based tutoring scenarios. The review analyzes publication characteristics, study designs, agent roles, AI models and architectures, six dimensions of agentic capability, and the extent of educational-theory integration, providing a roadmap for the field's frontiers and gaps. Notably, only **29% of the reviewed studies** (138 of 474) explicitly grounded their systems in educational theory, exposing a disciplinary divide between technically oriented and pedagogically oriented work. ## A role-based map of the field Where [[agentic-ai-education-scoping-review|the 474-study scoping review]] and [[kostopoulos-agentic-ai-education-2025|Kostopoulos et al.'s conceptual survey]] map research breadth and capability, [[baradziej-agentic-ai-higher-education-2026|Baradziej (2026)]] organizes agentic AI by the **role it plays** — an [[governance|institutional]] framing for deciding where to deploy and govern these systems. Across 48 higher-education studies, six roles emerge in order of evidential weight: personalized learning and adaptive tutoring (18/48), [[automated-assessment]] and feedback (12), teaching assistance and augmentation (11), administrative and student support (8), curriculum design and workforce alignment (5), and research support and academic operations (4). The role lens foregrounds a design choice that recurs in every deployment: how much moment-to-moment control the human retains (a "Copilot" relation) versus how much the agent owns (an "Autopilot" relation) — the autonomy axis that determines whether an agent scaffolds learning or supplants it. ## Design and evaluation of agentic systems [[research-methods-aied|Research]] in the knowledge base spans design and evaluation: - **Hybrid agents grounded in theory outperform pure [[prompt-engineering|prompting]]:** [[jeon-isd-agent-bench-2026|ISD-Agent-Bench]], a benchmark of **25,795 instructional-design scenarios**, finds the best-performing approach integrates classical ISD frameworks (ADDIE, Dick & Carey, Rapid Prototyping) with modern ReAct-style reasoning — hybrid (theory + technique) > pure theory > technique-only. Grounding [[llm]] agents in established educational-design theory provides a structural advantage raw prompting cannot replicate. - **Assessment frameworks for agentic tools:** [[tool-invariant-framework-agentic-ai|The tool-invariant framework]] proposes [[teacher-role]] and assessing computational methods in a way that does not depend on any specific AI tool, emphasizing [[computational-thinking]] fundamentals, [[authentic-assessment]] via oral defense, and verification — relevant to [[cognitive-offloading|Over-Reliance]] concerns. - **A reporting-and-governance scaler: the Autonomy–Oversight–Evidence (AOE) framework.** [[beyond-agent-label-agentic-ai-governance-2026|Dey (2026)]] argues the term *agentic AI* is applied so inconsistently that evidence and oversight cannot be compared across studies — systems that plan, remember, use tools, or coordinate multiple agents are lumped with static GenAI interfaces and conventional [[pedagogical-agent|pedagogical agents]]. His critical integrative review of fifteen peer-reviewed reviews finds evidence strongest for artifact-level outcomes (feedback accuracy, hallucination reduction) and weakest for durable learning, equity, workload, or institutional outcomes, with authentic deployments typically short, single-site, and weakly tied to oversight. The fix is a concrete reporting triplet — **autonomy level A0–A4** (reactive generation → bounded orchestration → delegated task autonomy → workflow autonomy → *consequential autonomy* over grading/admissions/progression), **oversight level O0–O4** (unspecified → retrospective audit → pre-use approval → checkpointed control → *continuous bounded supervision*), and **evidence-maturity stage M0–M5** (concept → prototype/benchmark → participant evaluation → authentic deployment → extended/multi-site → *replicated/institution scale*) — plus a **proportionality rule**: allowable autonomy should not outrun either evidence maturity or oversight strength (e.g. an A4 consequential system demands M4–M5 evidence, legal validation, and O4 with human final authority). Reporting *A2–O3–M3* turns vague "agentic deployment" claims into comparable, testable specifications and gives institutions a staged adoption, logging, and rollback grammar — the evidentiary complement to the [[baradziej-agentic-ai-higher-education-2026|Copilot-to-Autopilot]] spectrum above. - **Adversarial robustness testing:** [[adversarial-stress-testing-role-playing-agents|Multi-agent stress testing]] coordinates Interrogator, Target, and Judge agents to reveal failure modes invisible to single-strategy testing, reducing robustness scores by 0.17–0.20 points — critical for persona consistency and [[pedagogical-safety|safe deployment]] with learners. - **Domain applications:** agentic systems appear across domains, including [[learnmate2-llm-adaptive-learning|adaptive learning agents]], [[educlaw-bench-pedagogical-llm-agents-2026|pedagogical LLM agents]], [[guided-llm-scaffolding-independent-learning|guided LLM scaffolding]], [[cyberagents-gamified-cybersecurity-learning-2026|gamified cybersecurity learning agents]], and [[hdr-brachytherapy-agentic-ai-simulation-2026|clinical simulation agents]]. - **Web agents as learning-experience evaluators:** a single autonomous "describing" web agent that navigates an online [[learning-design|lesson]] like a student — [[ai-web-agents-lesson-design-2025|Wang, Mitchell & Piech (2025)]] — produces a description rich enough to predict student dropout and give the designer actionable feedback before real learners engage, outperforming a full simulated cohort and every baseline on a global CS1 course. This positions agentic evaluation (an agent as a stand-in critic of a learning experience) as a distinct, low-cost use of agentic AI alongside agents that teach or design. ### Multi-agent systems A growing and distinct strand of agentic AI involves **multi-agent systems** that orchestrate multiple specialized agents with distinct roles. The knowledge base documents several architectures: [[code-gen]] pairs a generator agent with a validator agent for human-in-the-loop [[automated-question-generation|question generation]]; [[adversarial-stress-testing-role-playing-agents|adversarial testing]] coordinates Interrogator/Target/Judge agents; multi-agent classrooms (e.g., [[human-in-the-loop-ai|MAIC]] with teacher, TA, and classmate archetypes) create varied peer-learning dynamics; and [[multi-agent-llm-social-learning|multi-agent social learning]] explores how interacting agents shape learning. Multi-agent design raises distinctive questions about [[human-in-the-loop-ai|human oversight]] (which agent is accountable, and where does a human intervene?), coordination costs, and how role differentiation supports or complicates [[scaffolding]]. Two 2026 systems sharpen what role specialization buys — and where it stops paying. MeduAI-SP ([[ai-standardized-patient-scaffolding-medical-2026|Yang et al., 2026]]) splits clinical-interview instruction across four LLM agents — a patient agent, a Socratic tutor agent, a turn-level evaluator, and a final evaluator — each carrying a distinct instructional function. The architecture withholds answers by design: the tutor was barred from disclosing the diagnosis or case answers, and the final evaluator's OSCE scores stayed hidden during the live encounter. Role specialization, rather than raw model capability, is what the trial credits for improved consultation behavior (71.8% vs. 55.6% on the final examination), offering a concrete template for pedagogy-aligned multi-agent design. A second study points to the limits of that pattern: in a 45-student, 15-group ethics discussion system ([[ethics-training-agents-group-ethics-discussion-2026|Seo et al., 2026]]), three LLM personas embodying care, deontological and pragmatic ethics produced no significant differences among themselves, and all were rated below human peers on contribution, diversity and influence (all Kruskal-Wallis p < .001). The authors attribute the gap to [[ai-sycophancy|sycophantic]] agreement — agents accepted nearly every contribution unless it was wholly wrong — and to a lack of visible reasoning, arguing that task-focused agent design must pair process transparency with safeguards (fallible-peer framing, a dedicated Questioning phase) to avoid [[cognitive-offloading|over-reliance]]. A step beyond orchestrating a few specialized agents is the **full agentic multi-agent ecosystem** proposed by [[sudarshan-agentic-ai-ecosystems-higher-education-2026|Sudarshan et al. (2026)]]: an institution-wide platform coordinating learning, teaching, and administrative agents through cross-functional [[feedback|feedback loops]] and distributed intelligence. Its distinctive move is treating [[inclusive-learning|inclusivity]] as a first-class architectural concern — coordinating [[accessibility]], cognitive-support, and [[well-being]] agents so that learners with [[special-education|special educational needs]] are supported across cognitive, sensory, and emotional dimensions in real time, rather than being served by an isolated assistive tool. The paper situates this inside a **human–AI co-evolution** loop (human behaviors and decisions shape AI adaptation, which in turn enhances human capability) that keeps humans in the loop. - **Participant-specific LLM agents for collaborative problem solving.** Fang (2026) fine-tunes individual LLM agents on real participants' dialogue data to represent each participant in collaborative problem solving simulations, with probabilistic speaker and thematic-code selection and sliding-window plus summarized memory. Validated with [[network-analysis|Epistemic Network Analysis]], the simulated dialogues are statistically indistinguishable from real ones (ENA distance 0.17, permutation p = 0.65) — a demonstration of agentic AI reproducing authentic collaborative discourse. - **Socially intelligent multi-agent tutoring.** Socially intelligent multi-agent tutoring prototypes such as ASTRA study how learners coordinate with AI in dyads, using differentiated Tutor and Facilitator agents to prompt coordination and balanced participation. The framework's trace-based evaluation enables reproducible analysis of interaction, participation balance, and verification in introductory programming. ## The central tension: automation vs. learning The [[agentic-ai-pedagogical-best-practice-2026|pedagogical best-practice]] work articulates the field's defining tension: as education AI shifts from passive [[conversational-ai|chatbots]] to **proactive agents** that initiate and pursue goals, personalization improves but **learner [[agency]] and cognitive effort** are at risk. The more an agent automates, the less [[cognitive-offloading|cognitive work]] the learner does. The design response — **intentional friction, dynamic [[scaffolding]], [[human-in-the-loop-ai]] oversight, and considered AI utilization** — acts as a principled guardrail. This connects to [[desirable-difficulties]], [[sociocultural-learning]], and the risk of [[cognitive-offloading|Over-Reliance]], and to the broader theme of preserving [[agency]] in AI-mediated learning. One empirical check comes from [[spec-driven-development-ai-agents-sdpbl-2026|a practical report on Spec-Driven Development in a software PBL course (Tanaka et al. 2026)]]: students using AI agents across development phases generated more code but showed code-comprehension dips that recovered only after instructor one-on-one interviews - concrete field evidence for the automation-vs-learning tension and for [[human-in-the-loop-ai]] monitoring as the mitigation. The survey literature turns these principles into **measurable design [[guardrails]]** rather than vague intentions. On scaffolding, [[kostopoulos-agentic-ai-education-2025|Kostopoulos et al. (2025)]] recommend **fading protocols** — gradually reduce hint frequency after each successful attempt — and targeting a **[[help-seeking]] ratio (AI-initiated hints ÷ total student actions) below 0.3**, so the agent is not the one driving most of the interaction. They pair this with **reflective checkpoints** (e.g., ask the learner to explain their reasoning before the agent offers the next cue) and adaptive fading curves whose intervention likelihood drops as proficiency rises. On [[explainable-ai|transparency]], agents should expose a "Why this suggestion?" rationale and keep **timestamped decision-traceability logs** (agent rationale, data sources, decisions) available for instructional auditing. On [[bias-mitigation|fairness]], they advise pre-deployment **disparate-impact testing across at least three demographic groups** (e.g., gender, language, geography) and involving diverse teachers in design. These metrics give an instructor or designer an audit lever: rather than asking "is the agent too helpful?", measure whether hints are fading, whether the learner is initiating, and whether the agent's reasoning is inspectable. A complementary vocabulary for the autonomy question comes from the aviation analogy surfaced in [[baradziej-agentic-ai-higher-education-2026|Baradziej's (2026)]] synthesis: how far a deployment sits on the **Copilot-to-Autopilot spectrum** — from an agent that assists a human retaining moment-to-moment control, to one that owns the whole task with the human only supervising. The same underlying system can be configured toward either end, and the choice is pedagogical before it is technical. Copilot-style configurations (agent proposes, learner disposes, human retains final say) tend to preserve [[agency]] and support the effortful processes that build learning; Autopilot-style configurations maximize task completion and efficiency but shift the cognitive load off the learner. Selecting a position on this spectrum — per task, not once globally — is a concrete way to operationalize the field's "intentional friction" principle. ## Positive implications of AI agents for education When designed well, agentic AI offers substantial benefits: - **Deeper, more adaptive personalization.** Persistent agents can sustain a learning conversation over many turns, tracking what a learner knows, adapting difficulty, and sequencing multi-step [[scaffolding]] — going beyond the one-shot responses of earlier chatbots. This supports [[adaptive-learning|adaptive]] and [[personalized-learning|personalized]] learning at scale. In [[baradziej-agentic-ai-higher-education-2026|Baradziej's (2026)]] synthesis the strongest evidence for this role reports academic gains of **15–25%** and engagement increases of up to **+40%**, with particular potential for learners historically ill-served by one-size-fits-all instruction (first-generation students, learning differences, second-language learners). - **Unburdening routine instructional work.** Agents can plan lessons, generate and validate questions, draft feedback, and orchestrate specialized sub-agents (e.g., generator + validator for question creation), freeing teachers for higher-value interaction. This is the promise of [[ai-tpack-teacher-multi-agent-workflow|teacher-facing multi-agent workflows]]. Assessment agents in particular report **90–95% agreement with human graders and 50–70% reductions in grading time**, though the same evidence flags [[bias-mitigation|bias]] and the [[metacognition|metacognitive]] cost of instant, unreflective feedback. - **Rich, varied interaction.** Multi-agent classrooms and simulated peers create diverse interaction dynamics (peer-like discourse, constructive disagreement, role-play) that single-agent systems cannot, supporting [[collaborative-learning]], [[socratic-method|Socratic-style probing]], and [[simulation]]. - **Productive friction.** Agents designed to challenge rather than agree can push learners toward deeper reconsideration. [[ai-agents-constructive-conflict-design-education-2026|Research on adversarial design agents]] shows that constructive-conflict agents prompted significantly more design iterations, broader exploration, and higher-rated final designs (N=48) — a form of [[desirable-difficulties|desirable difficulty]]. - **Scalable practice and simulation.** Agent-based simulations (simulated students, [[medical-education|clinical]] scenarios) let learners practice in low-risk environments before real-world application, as in [[hdr-brachytherapy-agentic-ai-simulation-2026|clinical simulation]] and [[simulating-students|simulated learners]]. - **Evidence-aware scaffolding.** Well-grounded agents can apply [[learning-theories|learning theory]] and known pedagogy in their interactions, and [[benchmark|benchmarks]] show theory-grounded agents outperform raw prompting. A narrower, instructor-built class of agents gets a distinct argument in [[ai-agents-joyful-assessment-third-space-2026|El Khoury and Ma's third-space proposal]]: custom GPTs, Gems and Copilot Studio agents designed by instructors for a specific pedagogical purpose and explicitly **not** autonomous systems. Their claimed value is as low-stakes rehearsal space — an oral-exam simulator, a clinical-communication simulation, an ESL pronunciation avatar — where students practice before judgment while the instructor's evaluative role stays intact, with the agent outside the social hierarchies students navigate with peers and instructors. Two boundaries are worth noting: the authors exclude agent-based *grading* from scope by design, and they argue assessment must remain relational — the instructor role cannot be substituted by a machine, though it can be extended and made more sustainable through careful design. ## Negative implications and risks of AI agents for education The same autonomy that enables these benefits also creates significant risks: - **Erosion of learner agency and cognitive effort.** The more an agent automates, the less cognitive work the learner does. Proactive agents that initiate, plan, and complete tasks can leave learners as passive consumers, hollowing out the effortful processes — drafting, recalling, revising — that build durable learning. This is the core [[cognitive-offloading|Over-Reliance]] and [[agency]] concern. The effect is measurable: in [[baradziej-agentic-ai-higher-education-2026|Baradziej's (2026)]] synthesis, passive learners in agentic-tutoring environments underperformed [[active-learning]] students by **8.7%** — evidence that the harm follows deployment design (letting the agent do the cognitive work) more than the technology itself. - **Over-automation of the learning process.** If an agent optimizes for task completion rather than learning, it can produce "answers" that bypass understanding — the very risk the [[tool-invariant-framework-agentic-ai|tool-invariant framework]] warns about, where the artifact no longer certifies the learner. - **Reduced metacognitive and self-regulated [[student-engagement|engagement]].** When agents handle planning and monitoring, learners may not develop the [[metacognition]] and [[self-regulated-learning|self-regulation]] that education aims to build. Agents must be designed to elicit, not replace, these processes. - **Misplaced trust and verification gaps.** Autonomous agents can produce plausible but unvalidated output; learners and teachers may [[trust-calibration|over-trust]] it. The need for robust verification and [[ai-literacy]] grows as agents take on more autonomy. - **Opacity, coordination, and accountability.** Multi-agent systems complicate [[human-in-the-loop-ai|human oversight]]: which agent is accountable for an error, and where does a human intervene? Coordination failures, persona drift, and emergent behaviors can undermine reliability and [[pedagogical-safety]]. - **Bias and equity.** Agents trained on data that encode bias can reproduce it at scale, and unequal access to capable agentic systems can widen [[equity-in-ai-education|educational inequity]]. Bias operates at multiple levels — training data, architecture, evaluation criteria, and test populations — so it needs institutional mitigation (bias audits, diverse datasets, transparent documentation, stakeholder involvement), not one-time checks. A subtler, culturally specific form is **epistemic hegemony**: because most agentic systems are trained on English, Western-produced data, they embed particular assumptions about knowledge, argumentation, and academic register. [[baradziej-agentic-ai-higher-education-2026|Synthesis evidence]] documents language-education tools marginalizing non-Western rhetorical traditions and penalizing linguistic features of non-English academic cultures, and [[global-south]] analyses show agentic pedagogies reproducing inequity when they ignore epistemological diversity and infrastructure constraints. - **Assessment integrity and skill decay.** When agents can generate work on demand, assessing genuine learning becomes harder, and over-reliance can erode foundational skills — the "comprehension debt" and certification problem the field flags. - **Ghost students and the verification gap.** [[bozkurt-ghost-students-agentic-ai-2026|Bozkurt, Crompton & Fell Kurban (2026)]] describe the **"ghost student"** — a digital surrogate created by coupling LLMs (the "mind") with agentic AI browsers (the "body") that can navigate Learning Management Systems, engage with content, and complete assessments with human-like mimicry, making the actual learner's presence optional. This creates a **verification gap** that traditional [[ai-detection|proctoring and detection]] tools are structurally unable to close, and it accumulates **cognitive debt** in the learner who is bypassed. As AI shifts from generative to agentic, this integrity and [[academic-integrity|verification]] threat grows — an agentic-specific risk beyond those of single-turn [[generative-ai|GenAI]]. ## AI agents and academic integrity Agentic AI poses distinctive integrity threats that go beyond the single-turn GenAI cases the field already struggles with. Because agents act autonomously over long horizons — and because "ghost students" (LLM "minds" coupled with agentic browser "bodies") can navigate [[online-teaching-and-learning|Learning Management Systems]], engage content, and complete assessments with human-like mimicry — they make the learner's genuine presence optional and create a **verification gap** that [[ai-detection|proctoring and detection]] cannot close. Several integrity implications follow: - **The artifact no longer certifies the learner.** When an agent can generate, plan, and execute an entire submission, the product's quality reflects the agent's capability, not the learner's. This is the [[tool-invariant-framework-agentic-ai|tool-invariant]] certification problem at its extreme — traditional "submit the work" assessment loses its evidential value. - **Verification, not detection, is the only viable response.** Detection-based policing is structurally unable to keep up with autonomous agents. The integrity question shifts from "can we catch AI agents?" to "can we verify what the learner can actually do?" — favoring [[authentic-assessment|process-based]], interactive, and [[human-in-the-loop-ai]] verification. - **Agentic completion of assessed coursework is now demonstrated, and it is a validity failure, not just an integrity one.** [[ai-agents-complete-lms-assessment-validity-2026|Hadjisolomou & El-Haddad (2026)]] documented agents (Claude for Chrome, Perplexity Comet, Claude Opus) logging into a live undergraduate LMS course and completing real assessed work from a single instruction — a 10-question quiz scored 10/10 in under 5 minutes, and a discussion-board post in which the agent fabricated a credible personal life story after mining peers' posts. Applying Kane's argument-based validity framework, they place a "human-production assumption" at the base of the scoring inference: agent completion removes its backing, so every unproctored asynchronous score — including honestly earned ones — loses interpretive support because authorship is unverifiable. Framing the problem as validity rather than integrity matters because an institution can punish misconduct and still lack grounds for the scores it reports, and because the same artifacts feed program-review and accreditation evidence chains. Their remedy, aligned with this page's "verification over detection," is assessment redesign for verified human presence (presence over product, integration over isolation, authenticity over genericity, low-stakes practice / high-stakes presence) with an equity-preserving menu of verified-moment options. - **Accountability is diffused.** In multi-agent systems, when an autonomous agent produces problematic output, it is unclear who is accountable — the learner, the system, or the institution. This blurs the attribution that academic-integrity processes assume. - **Cognitive debt accumulates silently.** Ghost students let learners bypass the effortful processes that build understanding, accruing [[cognitive-offloading|cognitive debt]] that surfaces only when independent performance is required. Integrity is thus tied to genuine learning, not just rule-compliance. - **It widens equity gaps.** Learners with access to more capable agentic systems gain an outsized advantage, and automated support may erode help for those who need it most — an [[equity-in-ai-education]] dimension of integrity. This connects the agentic-AI discussion to the knowledge base's [[academic-integrity]] coverage, which frames the response as assessment redesign and [[ai-literacy]] rather than detection alone. ## Productive friction and social interaction Not all agentic behavior need be smooth assistance. [[ai-agents-constructive-conflict-design-education-2026|Research on adversarial design agents]] shows that agents enacting **constructive conflict** prompted significantly more design iterations, broader exploration of alternatives, and higher-rated final designs among novice interaction designers (N=48) — a *productive friction* dynamic, where the conflict agent was frustrating but ultimately helpful. This connects to [[socratic-method|Socratic questioning]] and [[design-thinking]], and illustrates how agentic AI can support deep reconsideration rather than passive acceptance. A study of how agentic AI reaches learning outcomes through *psychological* rather than technological pathways supplies the missing measurement angle. [[pramod-agentic-ai-motivational-pathways-2026|Pramod and Patil (2026)]] surveyed 398 business students in India and modeled autonomy, competence and relatedness alongside interactivity, information sharing and perceived [[community-of-inquiry|social presence]]; autonomy was the strongest motivational driver (β = 0.504) and interactivity the strongest social one (0.468), with motivation and social presence feeding [[student-engagement|engagement]] at nearly the same strength (0.533 and 0.493) before engagement predicted perceived learning performance (0.671). The design lesson is that the social route is not automatic: a collaborative-environment construct moved perceived social presence less than plain responsiveness and information sharing did, so treating an agent as a chat interface rather than a participant leaves most of that pathway unused. ## Implications for instructors and instructional designers For teachers, faculty, and [[learning-design|instructional designers]], agentic AI changes both what is possible and what must be guarded: - **Reallocate effort to higher-value work.** Agents can take over lesson planning, question generation and validation, feedback triage, and resource retrieval. Instructors should treat these as automatable scaffolds that free time for what agents cannot do: relational teaching, contextual judgment, and the design of learning experiences. Teacher-facing [[ai-tpack-teacher-multi-agent-workflow|multi-agent workflows]] are a promising model. - **Keep the learner's cognitive work front and center.** The central design question is not "what can the agent do?" but "what must the *learner* do?" Instructional designers should configure agentic systems so they scaffold rather than replace learner planning, monitoring, and effort — using dynamic [[scaffolding]] and [[desirable-difficulties|intentional friction]] to protect [[agency]] and avoid [[cognitive-offloading|over-reliance]]. - **Design for verification and process, not just output.** When agents can generate work on demand, the artifact no longer certifies learning. Instructors should pair agentic tools with [[authentic-assessment|process-based assessment]] (oral defense, [[tool-invariant-framework-agentic-ai|tool-invariant]] tasks, verification checks) so that understanding — not just production — is measured. - **Curate and ground agents in pedagogy.** Benchmark evidence shows theory-grounded agents outperform raw prompting. Designers should ground agent behavior in established instructional frameworks (e.g., gradual release, Socratic questioning, [[learning-theories|learning theory]]) rather than defaulting to generic tool-chaining. - **Retain human oversight and judgment.** Multi-agent and autonomous systems make [[human-in-the-loop-ai]] design essential: decide where a human intervenes, who is accountable, and how failures are caught. Adversarial testing helps surface failure modes before deployment. - **Build instructor [[ai-literacy]].** Teachers and designers need accurate mental models of agentic AI to configure, monitor, and critique these systems — and to model responsible use for learners. This links to [[teacher-ai-competency]] and [[educational-development|faculty development]]. - **Watch for equity.** Agentic tools risk widening gaps if access is unequal or if automation erodes support for the learners who need it most; design with [[equity-in-ai-education]] in mind. - **Stand up institutional scaffolding before scaling.** The deployment question is not only design-level but institution-level. [[baradziej-agentic-ai-higher-education-2026|Baradziej's (2026)]] synthesis condenses the governance evidence into three pillars: develop [[ai-literacy]] among students *and* staff; build [[ethics|ethical]] infrastructure (data-protection policies, algorithmic-accountability and academic-integrity frameworks) *before* large-scale deployment; and deliver competence-based [[educational-development|educator training]] that goes beyond tool familiarization to pedagogical frameworks preserving human agency. Given only ~6.5% of faculty in some national contexts report direct AI use, the training gap is a binding constraint on responsible adoption. ### Techniques for ensuring academic integrity with agentic AI Because autonomous agents make detection futile, instructors should focus on techniques that **verify learning** and **make honest work visible**, rather than on policing: - **Prefer verification over detection.** Replace or supplement "submit the work" with interactions that require the learner to demonstrate understanding they cannot outsource: oral defenses, [[tool-invariant-framework-agentic-ai|tool-invariant]] tasks, live [[problem-solving]], and [[authentic-assessment|process-based]] assessment. The goal is to establish what the learner can do independently, not to catch an agent. - **Use interactive and staged assessment.** Require staged submissions (drafts, revisions, reflections) and follow-up [[conversational-ai|conversational]] checks that probe whether students understand their submitted work — the "AI Viva" and cognitive-stewardship approaches. Ghost students cannot sustain a live interrogation they did not perform. - **Set clear, purpose-driven expectations.** Ground integrity expectations in the course's purpose — what AI use is allowed, when, and why — rather than abstract rules. [[educational-policy-ai|Policy]] clarity that is aligned with pedagogy reduces the ambiguity students exploit and the misjudgments documented in integrity research. - **Make AI use visible and declared.** Structured, task-specific AI-use declarations (mapping use to cognitive stages) force reflection and normalize honest disclosure, shifting the culture from concealment to transparency. - **Build AI literacy as integrity education.** Teach students how to use agents responsibly and to judge output critically, framing integrity as genuine learning rather than rule-following. This includes [[ai-literacy]], understanding what agents can and cannot do, and the [[cognitive-offloading|learning cost]] of bypassing effort. - **Keep humans in the loop.** Maintain [[human-in-the-loop-ai|human oversight]] of assessment decisions, verify high-stakes submissions interactively, and design agentic tools so an instructor can always intervene. - **Close the verification gap with interaction.** For fully online or asynchronous contexts, use proctored or interactive components that require live presence, addressing the [[bozkurt-ghost-students-agentic-ai-2026|ghost-student]] threat directly rather than assuming detection will catch it. ## A balanced takeaway Agentic AI is neither a panacea nor an inevitable harm: its value depends on design. Used to scaffold learner agency, ground in pedagogy, and keep humans in the loop, agents can personalize and deepen learning; used to maximize automation and task completion, they can erode the very effort that produces learning. The recurring design principle is **intentionality** — deciding explicitly what the agent does and what it deliberately leaves for the learner. ## Connected Concepts - [[scaffolding]] - [[intelligent-tutoring]] - [[ai-literacy]] - [[prompt-engineering]] - [[curriculum-design]] - [[metacognition]] - [[adaptive-learning]] - [[educational-development]] - [[human-in-the-loop-ai]] - [[agency]] - [[cognitive-offloading]] - [[desirable-difficulties]] - [[sociocultural-learning]] - [[simulation]] - [[pedagogical-safety]] - [[ai-education]] - [[equity-in-ai-education]] - [[authentic-assessment]] - [[teacher-role]] - [[learning-design]] - [[teacher-ai-competency]] - [[academic-integrity]] - [[online-teaching-and-learning]] - [[educational-policy-ai]] ## Connected Articles - [[pramod-agentic-ai-motivational-pathways-2026]] — Autonomy, competence, relatedness and social presence as the pathways from agentic AI to engagement (Pramod & Patil 2026) - [[ai-agents-joyful-assessment-third-space-2026]] — AI agents, joyful assessment, and third space - [[ilieva-agentic-genai-higher-education-2026]] — The AGAI-HE framework: bounded, human-supervised agentic GAI in higher education (Ilieva et al. 2026) - [[beyond-agent-label-agentic-ai-governance-2026]] — critical integrative review introducing the AOE evidence/oversight framework - [[sudarshan-agentic-ai-ecosystems-higher-education-2026]] — Perspective on inclusive agentic multi-agent AI ecosystems in higher education - [[baradziej-agentic-ai-higher-education-2026]] — Systematic review of the roles of agentic AI in higher education (48 studies; six roles; tripartite responsible-integration framework) - [[kostopoulos-agentic-ai-education-2025]] — Agentic AI in education: state of the art and future directions (IEEE Access survey; operational definition + taxonomy) - [[agentic-ai-education-scoping-review]] — Scoping review of agentic AI in education (474 studies) - [[agentic-ai-pedagogical-best-practice-2026]] — The tension between automation and learning - [[tool-invariant-framework-agentic-ai]] — Teaching and assessing computational methods in the age of agentic AI - [[jeon-isd-agent-bench-2026]] — ISD-Agent-Bench: benchmarking instructional-design agents - [[adversarial-stress-testing-role-playing-agents]] — Adversarial stress testing of role-playing agents - [[ai-agents-constructive-conflict-design-education-2026]] — Constructive conflict AI agents in design education - [[ai-tpack-teacher-multi-agent-workflow]] — Teacher TPACK and multi-agent workflows - [[code-gen]] — Code generation agents - [[educlaw-bench-pedagogical-llm-agents-2026]] — Pedagogical LLM agent benchmark - [[guided-llm-scaffolding-independent-learning]] — Guided LLM scaffolding for independent learning - [[learnmate2-llm-adaptive-learning]] — LearnMate-2 adaptive learning agents - [[cyberagents-gamified-cybersecurity-learning-2026]] — Gamified cybersecurity learning agents - [[hdr-brachytherapy-agentic-ai-simulation-2026]] — Agentic AI in clinical simulation - [[bozkurt-ghost-students-agentic-ai-2026]] — Ghost students and the agentic-AI verification gap (Bozkurt et al. 2026) - [[ai-agents-complete-lms-assessment-validity-2026]] — AI agents completing LMS tasks; validity failure via the human-production assumption (Hadjisolomou & El-Haddad 2026) - [[llm-agents-collaborative-problem-solving-simulation-2026]] — Fine-tuned participant-specific LLM agents reproducing collaborative problem solving dialogues (Fang 2026) - [[astra-multi-agent-tutoring-benchmark-2026]] — ASTRA synthetic benchmark for multi-agent tutoring and participation-balanced collaboration - [[ai-web-agents-lesson-design-2025]] — AI Web Agents: a describing agent as a learning-experience evaluator (predicts dropout, gives design feedback before students engage) - [[spec-driven-development-ai-agents-sdpbl-2026]] - SDD with AI agents in software PBL; automation vs. comprehension - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[ethics-training-agents-group-ethics-discussion-2026]] — Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration --- ## [Cognitive Offloading](https://edtechdev.github.io/aied/concepts/cognitive-offloading/) > **Cognitive offloading** — the use of external tools (including AI) to reduce internal cognitive demand, shifting mental work from the learner to the system. In [[ai-education]], cognitive offloading is the central mechanism through which AI tools can either support or undermine learning: appropriate offloading frees cognitive resources for higher-order thinking, while excessive offloading bypasses the processing required for durable learning. **Over-reliance** is the harmful end of this spectrum — the unproductive pattern where offloading crosses from strategic support into learning displacement. ## Questions to Consider - Cognitive offloading is using external tools — including AI — to reduce internal cognitive demand. The page is explicit that offloading isn't inherently harmful: notebooks, calculators, and search engines all offload. What do you think makes AI-mediated offloading different and potentially more consequential than these familiar tools? - A common assumption is that using AI too much is the problem. But this page separates over-reliance from mere frequency — it's about a mode of use that substitutes for learning, not about how often AI is used. Can you describe a way of using AI that is frequent but healthy, and one that is rare but harmful? - The 'speedup illusion' is that AI-assisted work feels faster and easier, creating a misleading impression of productivity that masks reduced learning — students conflate task completion speed with learning. When have you felt productive doing something quickly and later realized you'd learned little from it? - Research finds the harm of offloading is conditional, not intrinsic: 'AI that coaches preserves or boosts skill; AI that substitutes risks decay.' What is the practical difference between an AI that scaffolds your thinking and one that replaces it — and how would you tell which you were getting? - One study shows that merely having access to AI advice nearly eliminated people's willingness to say 'I don't know' — even when the advice was wrong — while nearly doubling confidence and cutting accuracy to a third. What does this suggest about how AI changes our awareness of our own ignorance? - The page introduces a metacognitive equity gap — a 'Matthew Effect with AI': because productive AI use requires prior knowledge and metacognition, already-advantaged students benefit more while those who need practice most are most likely to delegate the learning. How should an educator or designer respond to the fact that the same tool can widen existing divides? ## Introduction Cognitive offloading is not inherently harmful — humans have always used external tools (notebooks, calculators, search engines) to reduce cognitive load. What makes AI-mediated offloading different is its comprehensiveness: [[llm|LLMs]] can generate complete solutions, explanations, and analyses, potentially eliminating the need for the very cognitive processes that produce learning. ### How cognitive offloading manifests in AIED research The knowledge base's articles document cognitive offloading across multiple dimensions: - **Instrumenting self-reflection rather than measuring a trait:** PAUSE (Patterns of AI Use: Self-Examination) turns the 2023–2026 offloading literature into a four-domain self-check — reasoning and critical thinking, creativity and originality, research and learning, social and communicative capacity — with reverse-scored behavioral items, citation anchors on every item, no composite score and no claim to validity ([[pause-ai-cognitive-offloading-self-reflection-2026|Alam, 2026]]). Its design position is that the useful intervention on offloading is prompting reflection rather than producing a diagnosis, since no instrument yet has the standing to justify a consequential decision about a person. - **Prompt patterns as offloading traces:** [[misiejuk-cognitive-offloading-prompting-2026|Misiejuk et al. (2026)]] use Co-Occurrence [[network-analysis]] to show that reactive prompts (disagreement without domain context) indicate higher offloading, while context-rich prompting with integrated instruction reflects engaged cognition. The *how* of AI use — not just whether it's used — determines the degree of offloading. - **Naturalistic message-level evidence at scale — offloading observed, not assumed:** [[student-cognitive-offloading-ai-higher-ed-2026|Piatnitckaia et al. (2026)]] coded 3,047 [[conversational-ai|ChatGPT]] messages from 46 undergraduates at one European university across a seven-week exam-preparation window, using GPT-4o-mini with enforced chain-of-thought rationales and validating the pipeline against a trained human rater on 200 messages (Cohen's κ = 0.76 for question type, 0.75 for Bloom level). Analyze topped that distribution at 27.28% (831 messages) ahead of Understand at 23.76%, the first fine-grained evidence that naturalistic [[student-ai-interaction|student–AI use]] routinely aims at higher-order work rather than memorization. A 16-student subsample with linked grades, whose 1,140 messages were hand-coded, supplies what prompt-level traces cannot: 49.6% of 125 dialogues showed no offloading, 34.4% light and 16.0% heavy, on a rubric that reserves *heavy* for cases where the AI produces the first draft and constructs the core intellectual product. Heavy offloading was overwhelmingly a Create phenomenon (80% of the 20 heavy dialogues, 42% of all Create dialogues), which the authors read as evidence that offloading degree and Bloom level are distinct dimensions worth measuring separately. The grade pattern was descriptive only — 17.6% heavy in the top tier, 17.5% in the middle and 5.9% in the bottom (χ² = 5.80, df = 4, p = 0.215) — and rests partly on one programming-focused student whose removal drops the top-tier rate to 8.9%, leading the authors to argue that discipline shapes delegation habits more than performance and to recommend [[metacognition|metacognitive]] feedback on actual usage patterns rather than prohibition. - **The speedup illusion:** [[cognitive-offloading-speedup-illusion|Research on the speedup illusion]] demonstrates that AI-assisted work *feels* faster and easier, creating a misleading impression of productivity that masks reduced learning. Students conflate task completion speed with learning, a metacognitive blind spot. - **Learning losses from unguided AI:** [[generative-ai-guardrails-harm-learning|High school math RCTs]] show that GenAI without [[guardrails]] produces worse learning outcomes than traditional instruction. [[generative-ai-reduced-study-time-math|Reduced study time]] correlates with reduced learning — students complete tasks faster but retain less. - **Offloading is not always harmful — the "coach" boundary condition:** [[coach-not-crutch-ai-writing|Lira et al. (2025)]] show that AI can reduce practice effort *and* improve the learning environment, yielding "work less, learn more." Adults who practiced writing with an AI tool wrote better no-AI letters than those who practiced alone — even beating personalized feedback from human editors — with no illusion-of-mastery inflation. The reconciliation with the harms above is the **form of offloading**: Lira et al.'s AI *scaffolded* (surfacing examples and feedback while keeping the learner in the loop) rather than *replacing* the cognitive act. [[ai-making-us-stupid|The skills-vs-basic-abilities perspective]] converges on the same boundary: **AI that coaches preserves or boosts skill; AI that substitutes risks decay.** So offloading's effect on learning is conditional, not intrinsic. - **Critical [[student-engagement|engagement]] vs. offloading:** [[favero-critical-ai-tutors-empower-enslave-2025|Favero et al.]] frame [[intelligent-tutoring|AI tutors]] as either empowering (supporting active cognition) or enslaving (enabling passive offloading), connecting to [[critical-thinking]] [[research-methods-aied|research]]. - **Metacognitive awareness:** [[metacognitive-awareness-experiential-vs-instructional|Studies on metacognitive awareness]] examine whether students recognize when they're offloading versus learning — and whether instructional interventions can improve this calibration. - **Embodied intelligence as the alternative to outsourcing:** [[zhu-e3-hot-embodied-intelligence-sustainable-learning|The E3-HOT framework]] argues that to counter AI-induced cognitive outsourcing and learning detached from authentic contexts, AI should be designed around *embodied intelligence* (situational embedding, embodied participation, cognitive creation) so learners sustain cognitive agency and higher-order thinking rather than offload it. This frames embodied, [[situated-learning|situated]] AI design as the positive counterpart to offloading risk, connecting to [[distributed-cognition]] and [[embodied-learning]]. - **The efficiency–[[regulation]] trade-off of distributed cognition:** [[hao-human-ai-collaborative-problem-solving-cognition|Hao et al.]] show that in human–AI collaboration, the mode that offloads most to AI (delegated reasoning) performs best on tasks but correlates with reduced self-regulation — empirical evidence that offloading's efficiency gain can come at the cost of the learner's regulatory engagement, converging with [[self-regulated-learning]] concerns. - **Fatigue and cognitive burden:** [[ai-fatigue-academic-contexts|AI fatigue research]] documents how constant [[student-ai-interaction|AI interaction]] creates its own cognitive burden, a paradox where offloading one task increases cognitive load from managing AI outputs. - **The metacognitive beliefs-vs-experiences framework:** [[cognitive-offloading-metacognitive-review-2026|Guo & Ye (2026)]] apply Nelson and Naren's dynamic metacognitive model to reconcile the field's contradictory intervention findings. They distinguish metacognitive *beliefs* (stable, self-referential self-conceptions that anchor offloading choices pre-task) from metacognitive *experiences* (dynamic, task-specific feelings that drive belief updating during-task), yielding the principle of **timing-component matching**: belief-targeting feedback is most effective before a task, while experience-targeting feedback (immediate correctness indicators) is most effective during it. They also formalize **substitutive offloading** (replacing internal processing with external aids) vs. **duplicative offloading** (supplementing it) — when external stores vanish, substitutive offloaders decline sharply while duplicative offloaders retain accuracy — and use **reminder bias** to quantify deviation from optimal offloading. This converges with the "coach vs. crutch" boundary: offloading that scaffolds preserves skill; offloading that substitutes risks decay. - **A formal problematic-use model for AI dependence in academic writing (I-PACE):** [[ai-dependence-academic-writing-ipace-2026|Liu, Zhuang & Wang (2026)]] extend the I-PACE model of addictive-technology use to generative AI dependence in college [[writing-education|academic writing]]. In a [[mixed-methods-research|mixed-methods]] Chinese sample, academic stress is the strongest predictor of AI dependence, [[ai-literacy]] is a protective factor (lower literacy → more psychological dependence), and perceived trust mediates the path from social influence to dependence — so dependence forms through a social-influence → trust → behavior pathway, not just individual tool use. Their [[qualitative-research|qualitative]] data add a policy dimension: students report strategic evasion of [[ai-detection|plagiarism detection]] and cite ambiguous rules about what counts as compliant AI use as an incentive to improvise, connecting offloading to [[academic-integrity]]. - **Cognitive debt and the episodic–habitual offloading distinction:** [[critical-thinking-paradox-genai-learning-2026|Lin & Al-Hada (2026)]] formalize the "critical-thinking paradox" — improved products alongside reduced cognitive engagement — through a differentiated three-level framework (surface/intermediate/deep AI roles) and the construct of *cognitive debt*: a potential cumulative decline in metacognitive calibration and unaided higher-order performance that persists beyond an AI-assisted episode. Their key conceptual advance is distinguishing **episodic offloading** (deliberate, task-specific delegation with retained awareness) from **habitual offloading** (routine, weakly monitored reliance), predicting that the latter on deep-processing tasks yields a product–process dissociation — higher-rated assignments but lower unaided delayed transfer. - **The outsourcing-to-reallocation spectrum: offloading can redistribute effort rather than reduce it.** [[yan-cognitive-outsourcing-genai-assessments-2026|Yan et al. (2026)]] applied Biggs' presage–process–product model to 38 undergraduates in Japan and China writing unsupervised argumentative essays, locating student–GenAI engagement on a spectrum bounded by **cognitive outsourcing** and **cognitive reallocation** — the GenAI-era analogue of surface versus deep approaches. The reallocation minority reported *unchanged total effort* with a shifted focus, moving resources from low-level retrieval to critical evaluation and reflective consolidation (writing reflection notes after each session to counter shallow retention), which qualifies the assumption that offloading necessarily subtracts cognition. Reallocation was the exception, however: 78.94% touched GenAI only before starting or after drafting rather than alternating it with independent work, and 76.32% relied on a single-turn ask–get-answer–stop pattern (23.68% sustained iterative dialogue), so the integrating behavior that produced reallocation had to be taught rather than assumed. - **Metacognitive training reduces reminder bias (direct empirical evidence):** [[metacognitive-training-optimal-cognitive-offloading-2026|Ngai & Gilbert (2026)]] provide the first clear demonstration that a brief intervention can make offloading measurably more optimal. Two preregistered experiments (N=164, N=416) found that **just five practice trials pairing a performance prediction with veridical, trial-by-trial feedback** improved metacognitive calibration and reduced reminder bias. The four-group additive design isolated the mechanism: **predictions alone were ineffective; adding performance feedback drove the improvement; explicitly labeling over-/under-confidence added nothing further**. The effect appeared on *absolute* (not signed) bias — training corrected individual miscalibration in both directions. This empirically validates the beliefs-vs-experiences framework above: it is *experience-targeting feedback*, not beliefs or prediction alone, that changes offloading behavior. The authors attribute success to financial incentive tied to offloading optimality plus immediate veridical feedback. - **An in-task reflection prompt makes reliance more discriminative rather than more defensive:** [[ren-metacognitive-awareness-genai-reliance-2026|Ren (2026)]] randomized 342 undergraduates across three conditions (no AI, open multi-turn [[conversational-ai|ChatGPT]] support, and the same support plus a brief reflection prompt on their own reasoning and on what would justify rejecting the AI explanation) and found that open support carried 62.4% acceptance of incorrect AI advice, which the reflection prompt cut to 39.7% (OR = 0.40), while recommendation accuracy stayed comparable across the two assisted conditions (66.3% vs. 67.0%) and alignment with correct advice remained high. Reflection also improved awareness calibration between perceived and behavioral reliance (0.59 vs. 0.41) and lowered an AI-specific attribution bias index from 0.42 to 0.21. This is the same *experience-targeting* mechanism as the training study above, applied inside the decision episode, and it supports the page's monitoring rather than frequency reading of over-reliance: knowing about model limits did not by itself stop students accepting plausible wrong advice. - **Metacognitive laziness and a new metacognitive [[equity-in-ai-education]] gap:** [[lodge-loble-cognitive-offloading-2026|Lodge & Loble (2026)]], a sector report for Australian schooling, argues the core risk of GenAI is cognitive offloading rather than [[academic-integrity|plagiarism]]. They adopt **metacognitive laziness** (Fan et al. 2024) — the convenience of AI undermining learners' engagement in essential self-regulatory processes, so learners abdicate metacognitive responsibility to the tool — and introduce a **metacognitive equity gap** (a "Matthew Effect with AI"): because leveraging AI productively requires [[prior-knowledge]] and metacognition, already-advantaged students benefit more while those who need the practice most are most likely to delegate the learning, widening existing divides. Their proposed remedy is **teacher augmentation** (giving the tool to expert teachers to scale their practice, supported by studies showing teacher-facing AI improves outcomes at far lower cost) rather than student-facing AI tutors, alongside Load Reduction Instruction and metacognitive prompts. - **Over-reliance is the leading [[ethics|ethical]] concern across all [[conversational-ai]] generations.** The [[conversational-ai-agents-umbrella-review-2026|umbrella review of conversational AI agents]] (Ganguly et al. 2025, 34 reviews) reports that human–AI relationship concerns — including over-reliance and the diminution of social interaction — are the **most frequently discussed ethical issue across all CAI generations**, predating GenAI. It also lists educational impact and cognitive concerns (including overreliance and degraded critical thinking) as the second most-discussed challenge category, underscoring that offloading's harm is a persistent, cross-generation theme rather than a GenAI-specific novelty.([[conversational-ai-agents-umbrella-review-2026]]) - **Teachers as reflective regulators of cognition:** [[teachers-reflective-regulators-cognition-offloading|Ho and Chen (2026)]] extend offloading theory to *professional* AI judgment by interviewing 18 in-service teachers. They identify a 'metacognitive ecology' in which teachers recognize, redistribute, and reflectively re-engage cognition with GenAI, [[framing-ai-use-for-students|framing AI]] as a cognitive partner rather than a thinking substitute — and flag 'professional drift' as a risk when offloading goes unreflective in AI-augmented [[teacher-role]] and [[administrator|administration]]. - **Offloading is value-based decision-making, and some students are more vulnerable than others.** [[seung-basham-cognitive-offloading-swld-2026|Seung & Basham (2026)]], a conceptual review in *Learning Disability Quarterly*, synthesize cognitive science, [[special-education]], and educational technology to model GenAI offloading as a cost–benefit decision shaped by performance goals, task difficulty, academic self-efficacy, and perceptions of the tool. They argue that **students with learning disabilities (SWLDs)** are especially vulnerable to suboptimal offloading — because heightened cognitive load, effort-avoidant performance goals, lower academic self-efficacy, and inflated expectations toward GenAI make premature or excessive delegation more likely. GenAI is framed as a **compensatory aid or shortcut depending on how offloading decisions interact with learner profiles and [[learning-design|instructional design]]**, with instructional [[guardrails]] the key moderating factor (teach metacognitive self-regulation, build [[ai-literacy]] to calibrate tool trust, sequence mastery experiences, and assess process not just product). This extends offloading's equity dimension: the same tool that reduces barriers to access can, if unguarded, substitute for the very practice SWLDs need most. - **Offloading risk is developmental.** [[niu-genai-children-creative-thinking-cognitive-development-review-2026|Niu et al. (2026)]] scoped 24 evidence sources on [[generative-ai|GenAI]] and children's [[creativity|creative thinking]] and found [[cognitive-offloading|over-reliance]] and prompt dependence among the recurring risks, strongest in the lower grades, alongside template-based thinking and children's difficulty judging whether an AI suggestion was original; one fMRI comparison recorded lower engagement of cognitive control and attention networks during child-ChatGPT interaction than in human conversation. Because younger children lacked the linguistic and metacognitive skills that text-based prompting demands, the authors recommend keeping [[ai-literacy|AI literacy]], authorship and unaided idea generation inside the activity itself and measuring independent creativity separately from AI-assisted output. This adds an age dimension to the vulnerability patterns above. - **The "thinking less vs. learning differently" question is conditional.** [[nesnin-cognitive-offloading-ai-students-2026|Nesnin et al. (2026)]] offer a broader analytical review concluding that AI is **not necessarily making students think less but transforming how they learn** — the outcome depends on use. AI that clarifies, verifies, and guides enhances learning; AI that replaces independent thinking yields passive dependence. This converges with the knowledge base's pervasive "scaffold vs. substitute" boundary and the conditional view of offloading's harm. - **Offloading is layer-sensitive: the depth of delegation matters, not just its occurrence.** [[layer-sensitive-cognitive-offloading-writing-2026|Chen (2026)]] introduces a layer-sensitive account for [[writing-education|academic writing]] — surface (grammar/vocabulary), structural (outline/sequencing), idea (claims/content), and reasoning (warrants/counterarguments/argumentative logic). In an eight-week quasi-experiment, open AI collaboration produced the highest *supported* writing but the lowest independent no-AI outcomes, with deeper offloading layers carrying the strongest negative association with independent [[higher-ed]] higher-order thinking (reasoning offloading indirect ab = −0.34 vs. surface −0.08). [[self-regulated-learning|Self-regulated writing]] attenuated but did not eliminate the harm. This refines the scaffold-vs-substitute boundary: some layers of delegation scaffold, while deeper layers substitute for the cognition that builds [[critical-thinking]]. - **A causal test of the offloading prediction, with the sign reversed: writing with ChatGPT produced *less* learning than writing unaided.** [[chatgpt-writing-cognitive-impact-2026|Wagner-Kobayashi (2026)]] ran a between-subjects experiment in which 35 psychology students spent 20 minutes elaborating an input text, with GPT-3.5 available only to the experimental group, and sat an identical knowledge test before and after. The hypothesized group × time interaction was significant (F(70) = 5.889, p = .018) but pointed the other way — the no-ChatGPT group learned more (posttest M = 8.29 vs 6.94; b = −1.450) — falsifying the [[writing-education|writing-to-learn]] prediction that ChatGPT would amplify [[metacognition|elaboration]]. What leaked away was the elaborative effort itself: participants produced only 2.25 examples and 0.88 connections in dialogue with the tool and reported prioritizing the word count over depth, the authors' *utilization deficiency* reading of a tool available but not deployed, and in line with the reduced mental effort and copy-paste behavior the page records elsewhere. Once [[motivation]] entered the model the group × time effect collapsed (p = .280), and higher motivation and topic interest predicted gain only *inside* the ChatGPT group, so low-motivation and low-interest students were the ones the tool disadvantaged — experimental support for the motivational pathway to over-reliance this page documents. The authors read the result narrowly (a small, assumption-violating effect under time pressure and without training) and recommend teaching students how to deploy the tool rather than banning it, which keeps the finding on the scaffold-versus-substitute boundary rather than making it a verdict on the tool. - **The surrender–offloading–agency continuum.** The Sydney PreK-12 rapid review (Arthars et al. 2026, 271 papers) frames GenAI use across cognitive, metacognitive, and affective dimensions: *[[cognitive-surrender|surrender]]* (responsibility for learning-relevant work shifts to GenAI, often unknowingly), *offloading* (deliberate, possibly productive delegation that becomes learning only if checked/elaborated), and *agency* (retaining responsibility for effort and judgment). It also warns of **metacognitive inequity**: weaker metacognitive students are more susceptible to detrimental offloading and less able to recognize it.([[young-people-learning-generative-ai-rapid-review-2026]]) - **Dependency is governed by self-efficacy as much as capability.** AI dependency is not simply a product of technical skill: [[student-dependency-on-ai-literacy-self-efficacy-2026|Maizel et al. (2026)]] show that skill-based AI literacy predicted higher dependency (consistent with the enabling-capacity/offloading view), while AI and academic self-efficacy buffered against overreliance — indicating offloading is governed by motivational self-efficacy beliefs as much as by capability. - **Offloading changes the threshold to respond, not just capacity.** [[ai-advice-suppresses-ikt-suspension-2026|Marcoccia et al. (2026)]] show that merely having access to AI advice nearly eliminated people's willingness to say "I don't know" — even when the advice was wrong — while nearly doubling confidence and cutting accuracy to a third; incentives restored accuracy (by reducing reliance) but not suspension. - **The Safety Gap as the cost of offloading struggle.** [[wang-safety-gap-productive-struggle-2026|Wang & Shan (2026)]] formalize the divergence between a student's AI-assisted performance and their unassisted capability as the "Safety Gap" — the epistemic risk when AI does the cognitive work and the learner cannot reproduce it. [[kim-ai-productive-failure-adult-2026|Kim et al. (2026)]] and [[puech-pedagogical-steering-llm-productive-failure-2025|Puech et al. (2025)]] show [[productive-failure]] design (withholding answers, preserving struggle) is the countermeasure. - **The ICAP/SAMR spectrum frames offloading as a continuum of modes, not a binary.** [[thermomix-genai-education-analogy-2026|Rummel, Nachtigall, and Panadero (2026)]] map the Thermomix kitchen-machine analogy onto learning with [[generative-ai]] within the [[icap-framework|ICAP]] and [[samr-model|SAMR]] frameworks, showing four scenarios that progress from **passive substitution** (fully outsourcing assignments — the offloading end, risking skill loss and limited [[creativity]]) to **interactive redefinition** (AI as a [[pedagogical-agent|dialogue partner]] providing real-time [[feedback|adaptive feedback]] and co-construction — the engaged end). This reframes offloading's harm as conditional on *mode of use* rather than mere frequency, converging with the pervasive "scaffold vs. substitute" boundary: whether AI use sits at the substitution/Substitution or redefinition/Redefinition end of the spectrum determines whether it displaces or supports the cognition that builds learning. ## Over-reliance: when offloading becomes harmful The ladder this page tracks has three rungs, and only the first two belong here. **Offloading** is delegating mental work, which is often productive; **over-reliance** is doing it chronically and without calibration, which is the behavioral failure this section documents; **[[cognitive-surrender]]** is a different failure — the evaluative step never happens, so the learner adopts the AI's answer without any judgment of its quality. The boundary is visible in what the AI is asked to supply: [[du-yuan-epistemic-dependence-2026|Du and Yuan's]] *instrumental* assistance produces output, while *judgment-bearing* assistance supplies the standard by which output is judged, and it is the second kind that displaces the work expertise depends on. [[young-people-learning-generative-ai-rapid-review-2026|The Sydney rapid review]] draws the same three-way split across 271 papers. Surrender therefore has its own experimental signature and its own page; over-reliance remains the frequency-and-calibration problem, and the two are separable outcomes in the same task — in [[shaw-nave-cognitive-surrender-2026|Shaw and Nave's]] trials, 73.2% of incorrect-AI trials ended in surrender while 19.7% ended in strategic offloading. **Over-reliance** is the excessive or uncalibrated dependence on AI tools where students delegate cognitive work they should perform themselves, resulting in reduced learning, diminished [[agency]], and the displacement of skill development. It is the behavioral manifestation of excessive cognitive offloading: when offloading becomes the default rather than a strategic choice. Over-reliance is not simply about using AI too much — it is about using AI in ways that substitute for rather than complement learning processes. Conceptual work urges keeping this educational over-reliance distinct from relational attachment and [[medical-education|clinical]] dependence: [[yan-conversational-ai-engagement-dependence-synthesis-2026|Yan (2026)]] shows that trust, reliance, over-reliance, attachment, and problematic use are routinely conflated in the conversational-AI literature, and that frequent delegation should not be labeled dependence without impaired control or harm. - **The efficiency paradox: mastery goals collapsing into offloading "by default."** [[yan-cognitive-outsourcing-genai-assessments-2026|Yan et al. (2026)]] identify learners who articulated mastery-oriented goals yet enacted surface-like processes, offloading "not by intention, but by default": their three intention groups (Outsourcing Tool n = 16, Learning Assistant n = 31, Cognitive Partner n = 8) show the modal case is neither deliberate cheating nor deliberate learning. Outsourcing-group students recognized their own disengagement (copy-pasting paragraph by paragraph was described as "it completely replaced my brain") and reported fast forgetting, and all 38 interviewees reported low confidence using GenAI effectively — evidence that over-reliance can be driven by missing presage conditions ([[ai-literacy]], prompting competence, task-specific guidance) rather than by motivation to avoid work. - **Agentic coding offloads comprehension in the field:** [[spec-driven-development-ai-agents-sdpbl-2026|Tanaka et al. (2026)]] report that undergraduates using [[agentic-ai|AI agents]] in a software project course wrote more code year-over-year but showed comprehension dips under heavy use - recoverable through one-on-one instructor verification of AI-generated code. most consequential risks of AI in education: - **Learning displacement:** [[ai-making-us-stupid|Research on AI's cognitive effects]] documents how AI availability reduces effortful processing — the "Google effect extended to reasoning." [[stamatoulis-genai-use-patterns-2026|Stamatoulis et al. (2026)]] isolate this as a distinct *pattern* of use: **low-verification uptake** (uncritically accepting AI output) predicted worse [[learning-gains|academic performance]], whereas **evaluative integration** (using AI to support understanding) predicted better performance — and usage **frequency** alone predicted neither. Over-reliance is therefore a *mode of use* that can be separated from how much students use AI. - **The agency problem:** [[aied-unfinished-mission-bypass|AIED's unfinished mission]] frames over-reliance as an agency and motivation crisis — students bypass learning not because AI is compelling, but because learning tasks feel pointless when AI can complete them effortlessly. - **Motivation erosion:** [[ai-availability-student-motivation|Student motivation research]] finds that knowing AI is available reduces the perceived value of learning the skill yourself, a motivational calculus that particularly affects novice learners. - **Literacy debt:** [[agentic-literacy-debt|Agentic literacy debt]] describes the cumulative skill deficit that develops when students habitually rely on AI rather than developing their own competencies, analogous to technical debt in software. - **Fatigue cycles:** [[ai-fatigue-academic-contexts|AI fatigue]] research identifies a paradox where over-reliance leads to cognitive fatigue from constant AI interaction management, which in turn drives MORE reliance — a vicious cycle. - **The placement rule:** [[brcic-effortless-trap-productive-struggle-2026|The Effortless Trap]] reframes allow-vs-ban as a placement question — an unguarded AI helper left high-school students ~17% worse on an unaided exam, while the same model rebuilt to withhold answers erased the harm. Its diagnostic — *"if letting AI in makes the task feel effortless, it is in the wrong place"* — secures the first hard attempt and the final unaided check as the moments where over-reliance most readily hides as an "illusion of learning." - **Metacognitive preservation:** [[vibe-compiler-metacognition-genai-agency-2026|The Synthesis-Analysis Reciprocity Model]] proposes tools that preserve human epistemic agency by structuring AI interaction around human analysis cycles rather than AI generation cycles. - **The metacognitive mechanics of overuse:** the beliefs-vs-experiences framework explains *why* students over-offload even when it hurts them — people offload impulsively, and pre-existing metacognitive *beliefs* anchor behavior faster than task *experiences* can correct it — so the antidote to over-reliance is metacognitive, not merely restrictive. - **Over-reliance is trainable via calibration training:** [[metacognitive-training-optimal-cognitive-offloading-2026|Ngai & Gilbert (2026)]] show reminder bias — the laboratory analogue of over-reliance — can be reduced with a brief metacognitive intervention (five practice trials pairing a prediction with feedback), correcting calibration in both directions. This implies the antidote to over-reliance is not merely restrictive rules but **calibration training that makes students accurate about what they can actually do unaided**. - **Field evidence: AI that coaches vs. AI that answers.** [[making-ai-tutoring-productive-mastery-math-2026|NUMI (Oreopoulos et al. 2026)]] found that AI support that coached rather than gave answers slowed students down but reduced effort-avoidance — improving next-attempt correctness after mistakes with more time per question (a "productive slowdown") — while [[one-click-away-khanmigo-two-year-school-experiment-2026|Khanmigo (Oreopoulos & Low 2026)]] showed that without structure making mistakes consequential, students default to shallow use (bare answers, prompt clicks) and gains match practice without AI. Both confirm that offloading's harm is contingent on **how** AI is used and designed, not just on access. CoMeT (Hou et al. 2026) supplies the design dimension that boundary otherwise leaves implicit — a tutor can carry *more* of the labor without buying more delegation. Its ladder climbed one rung each time a learner did not use its support and dropped to the lightest rung on take-up, holding [[metacognition|metacognitive demand]] statistically equivalent to a tutor that withheld answers by design (p_TOST = .004) while an artifact reached the workspace in 48.1% of sessions against 23.7% under that withholding tutor, and learners asked it to build in 50.4% of sessions against 51.1% under an unrestricted answer-on-request tutor (OR = 0.95, p = .82). Delegation tracked the demand the tutor placed on the learner rather than the amount of work it took over. - **Large-scale field evidence: the "[[generative-ai]] learning penalty."** [[stromberg-generative-ai-learning-penalty-secondary-2026|Strömberg, Lei, & Wu (2026)]], tracking 26,811 Chinese secondary students over 30 months, found that self-directed generative-AI adoption raised homework scores 18% and cut homework time 30% while *lowering* [[summative-assessment|closed-book exam]] scores 20% within six months and entrance-exam scores 18–24% after two years — concentrated among the ~81% of users whose behavior indicated homework outsourcing (short completion time + inflated homework scores). This is direct large-scale evidence that unguarded offloading (using AI as a homework substitute rather than a tutor) produces the learning penalty cognitive-offloading predicts, often undetected by students themselves. - **Over-reliance erodes self-directed learning via motivation and self-efficacy.** [[genai-thoughtless-use-self-directed-learning-2026|Zhao & Gu (2026)]] model the mechanism directly: across 487 Chinese undergraduates, **thoughtless use of GenAI (TUGA)** — adopting AI outputs without critical evaluation — significantly undermined [[self-directed-learning]] (β = −0.42) both directly and through partial mediation of [[motivation]] (β = −0.54 path from TUGA) and [[self-efficacy]] (β = −0.37). The model explained 75.3% of SDL variance. Because motivation was the strongest positive driver of SDL (β = 0.68), and thoughtless use suppresses it, this is [[quantitative-research|quantitative]] evidence that over-reliance damages the motivational and self-efficacy resources autonomous learning depends on — with gender differences (stronger motivation harm for males, stronger self-efficacy harm for females). - **Offloading tendency predicts lower higher-order outcomes, and verification literacy needs metacognition to pay off.** [[davor-ai-supported-learning-higher-order-outcomes-2026|Davor, Larbi and Boateng (2026)]] surveyed 533 university students in Ghana and modeled offloading tendency, AI task scaffolding and [[ai-literacy|AI verification literacy]] against higher-order outcomes through [[metacognition|metacognitive self-regulation]]. Offloading tendency was the strongest negative predictor of [[critical-thinking|critical thinking]] (−.240) and technical [[problem-solving|problem solving]] (−.312) and also depressed metacognitive self-regulation (−.294), while scaffolding predicted both outcomes positively (.185 and .170). Verification literacy had no significant direct effect on either outcome and worked only through metacognitive self-regulation, a full mediation pattern the authors call a metacognitive activation mechanism: teaching students to check AI output is not enough on its own, because the evaluative habit pays off only when it is embedded in planning, monitoring and reflection during the task. - **Dependence operates as a boundary condition on the benefit of use.** [[shojaei-genai-dependence-critical-thinking-employability-2026|Shojaei et al. (2026)]] surveyed 412 undergraduate business students in Oman and decomposed the association between [[generative-ai|GenAI]] use and self-reported critical-thinking disposition: the zero-order correlation was near zero (r = 0.050) while the adjusted coefficient was positive (.185), because a negative indirect path through dependence (−0.147) offset it. Dependence was negatively associated with both disposition (−.389) and self-perceived employability (−.328) and weakened both use-to-outcome links, and simple slopes fell from 0.424 at low dependence to −0.054 at high dependence, attenuation rather than reversal. The reading keeps the finding on the mode-of-use side of the boundary: dependence marks where use shifts from augmentation toward substitution, and frequency alone masks the opposing pathways. - **AI overreliance as a complex [[adaptive-learning|adaptive system]].** Rather than studying overreliance one user at a time, [[ai-overreliance-complex-adaptive-system-2026|a modeling paper]] frames it as a population-level process in which agents update Bayesian beliefs about AI quality and, when networked, learn from peers. Social proof can turn reliance into a feedback cascade (visible unverified use suppresses verification), while social learning creates consensus rather than overreliance — a framing that shifts intervention targets from individual calibration to the networked dynamics of trust and reliance. - **Six diagnostic criteria separate productive reliance from harmful dependence.** [[du-yuan-epistemic-dependence-2026|Du & Yuan (2026)]] give over-reliance a more granular diagnostic vocabulary than usage frequency, distinguishing instrumental assistance (AI helps produce output) from judgment-bearing assistance (AI supplies the standards by which output is assessed). Their review operationalizes the productive/harmful boundary through six criteria — **contestability, recoverability, transfer, traceability, distributed responsibility, and epistemic plurality** — and traces four sociotechnical pathways (fluent authority, frictionless delegation, opaque synthesis, institutionalized dependence) through which offloading either preserves or displaces the epistemic work that develops judgment. Because dependence is contextual and institutional rather than merely individual, the framework directs attention to assessment incentives, interface design, and procurement alongside learner self-regulation. - **The unmeasured aftermath: cognitive washout.** [[cognitive-washout-ai-skill-decay-2026|Yajee (2026)]] names the field's biggest open question as *post-withdrawal*: almost all offloading research measures cognition during AI use, but almost nothing measures what happens after an assistant is withdrawn for days or weeks (an exam, an outage, a license review). It formalizes **cognitive washout** with a Washout Curve Model — estimable parameters for recovery time constant, recovery completeness, residual growth, and a hysteresis index comparing relearning to original effort — and four possible outcomes (elastic rebound, partial plateau, latent scaffold, over-recovery). Because reversibility determines whether an induced deficit is an inconvenience or a cohort-level injury, the paper argues withdrawal deserves the same methodological standing as adoption, and specifies a three-arm, three-domain, twenty-two-week protocol to adjudicate between outcomes. This turns the "coach vs. crutch" and substitutive-vs-duplicative distinctions above into a *testable longitudinal* research agenda on whether and how quickly offloaded skills return. - **Reliance without domain knowledge degrades into guessing.** [[ai-particle-physics-education-redesign-2026|Mikhasenko et al. (2026)]], redesigning the introductory nuclear and particle [[physics-education|physics]] course at Ruhr University Bochum, describe a failure mode of reliance in the absence of domain knowledge: when a student could not judge whether a generated answer was physically sound, the intended "conversation with AI" degraded into guessing against plausible but unreliable output. The same course's mid-semester survey (n=30) found frequent LLM use (24 of 29 used them often or always) alongside low [[self-report-measures|self-reported]] preparedness for the computing fluency the assignments required, and open responses raised AI dependence and unequal access to paid models among the friction points. - **The Daoist counter-argument: offloading is not just impractical but self-defeating.** [[daoism-ai-education-philosophy-2026|Xie (2026)]] supplies a normative, anti-delegation counter-argument from Daoist self-cultivation: in Neidan (內丹) practice "there are no cognitive shortcuts," and the practitioner cannot outsource the labor to external devices, so AI should be "not a cognitive surrogate but an instrumental adjunct," akin to an alchemical furnace — a framing that aligns with the "coach vs. crutch" boundary above rather than with substitutive offloading. ### The CLT framework Cognitive Load Theory (Sweller) provides a contested theoretical lens on [[cognitive-psychology|working memory]] and instruction: intrinsic load (task complexity), extraneous load (presentation friction), and germane load (schema-building effort). Well-designed AI should reduce extraneous load while preserving germane processing; poorly integrated AI reduces all three, leaving students with completed tasks and empty learning. Note that the theory's claims are contested in the wider literature, but its framing remains influential in how offloading effects are discussed. **The profession-level stakes of offloading.** The Cognitive Commons framework ([[cognitive-commons-ai-expertise-regeneration|Lovett 2026]]) extends offloading from an individual to a collective level: when AI lets junior workers skip the cognitive struggle that builds deep expertise, it can deplete a profession's shared expertise pool over time — the "Validation Tether" means effective AI oversight depends on the very mastery that AI adoption may undermine. This reframes individual-level offloading and skill-decay as a systems-level regeneration problem with [[governance]] implications. ### Connections to related concepts Cognitive offloading (and its harmful form, over-reliance) connects fundamentally to [[trust-calibration]] — knowing when to trust and when to question AI — and [[ai-literacy]], which includes the metacognitive skill of knowing when to offload and recognizing one's own reliance patterns. It connects to [[scaffolding]] (structured support that reduces load without eliminating cognitive demand) and [[prompt-engineering]] (the primary mechanism through which offloading is enacted in LLM interactions). It intersects with [[metacognition]] and [[self-regulated-learning]] — effective learners calibrate their offloading decisions — and with [[critical-thinking]], [[agency]], and [[student-experience]]. [[online-teaching-and-learning]] is a particularly vulnerable context: the medium already distances learners from immediate accountability, and self-paced, screen-based work invites the "ask for the answer" shortcut that offloading research identifies as the core harm mechanism (see [[ai-misuse-learning-harm]]). **Not all offloaded friction is excess friction.** [[zohar-bloom-inzlicht-against-frictionless-ai-2026|Zohar, Bloom and Inzlicht (2026)]] draw the distinction that the offloading literature needs: previous [[ai-technologies]] removed *excess* friction — tedious or insurmountable obstacles with little learning or meaning value — whereas generative AI in intellectual work also strips away *beneficial* friction, letting a learner move from ideation to evaluation without questioning the output. They marshal the associative evidence bluntly: people who use AI struggle to accurately recall or reproduce their own work, acquire fewer skills, show less transfer, and perform worse when AI support is removed — converging with the cognitive-debt findings from EEG studies of essay writing with an AI assistant. Their argument also supplies the motivational half of the mechanism that pure cognitive accounts miss: because effort signals that our actions matter, offloading it reduces appraised purpose and meaning, and as AI substitutes for effort in a domain the motivational payoff of effort there erodes, deepening reliance further. The paper's corrective is a gradient rather than a prohibition — preserve moderate struggle, remove what only overwhelms ([[desirable-difficulties]], [[motivation]]). - **Dependent versus autonomous offloading — the distinction that sets the outcome:** [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] organize the GenAI evidence around a sharpened version of this boundary, drawing on a three-wave study of 589 students and early-career knowledge workers: *dependent* offloading delegates core thinking to the tool and was associated with transferred [[agency]], lower intrinsic motivation and poorer perceived cognitive outcomes, while *autonomous* offloading keeps epistemic agency with the learner and showed the opposite pattern. The finding that matters most for detection is that immediate performance benefits did not differ between the two modes, so a learner resolving tasks fluently on any given day gives no signal about which mode they are in. The same review records the wider split in the evidence — a three-level [[meta-analysis-systematic-review|meta-analysis]] of moderate overall benefit (g = 0.499; g = 0.669 for comprehension, cognition and creativity) against associations between dependence, fatigue and weaker [[critical-thinking]] — and traces maladaptive use to externalized self-regulation rather than to technology addiction. - **Three interaction pathways, not one behavior.** The Neuroplasticity-AI Interaction Model (NAIM) separates direct bypass, where the model supplies the solution and generative effort disappears; cognitive offloading, where specific sub-processes are delegated and the effect depends on whether those sub-processes are the learning target; and scaffolding, where the model constrains help to preserve effortful processing ([[naim-bypass-offload-scaffold-llm-learning-2026]]) Two controlled studies in the recent batch pin down the two halves of this claim — whether the behavior shifts, and what it was protecting. [[metacognitive-feedback-anti-deskilling-offloading-2026|Maier et al. (2026)]] made the learning consequence of offloading visible before each choice in a preregistered experiment with 704 participants practicing fraction arithmetic: odds of offloading an answer fell to OR = 0.47 and odds of answering a later unaided test item correctly rose to OR = 1.51, while an effort-based reward moved neither outcome. Offloading compounded within the session — after offloading one item, participants offloaded the next in 69.7% of cases without the feedback and 55.9% with it — and a ten-percentage-point rise in offloading was associated with 32% lower odds of unaided success (OR = 0.68). [[chatgpt-programming-performance-retention-ownership-2026|Bergh et al. (2026)]] supply the outcome-side complement: 55 computer science undergraduates who coded with ChatGPT scored 89% against 69% without it, recalled less of the same material immediately (41% vs. 53%) and at 48 hours (39% vs. 52%), and attributed only 45% of the submitted code to themselves against 81%. Read together, the feedback study shows the behavior can be shifted without restricting access to the tool, and the programming study shows what the shifted behavior was protecting.. The model is calibrated against the strongest available field evidence: unrestricted GPT-4 access in a study of nearly 1,000 high-school [[math-education|mathematics]] students raised practice performance by 48% yet left a 17% deficit on the unassisted exam, whereas hint-constrained GPT Tutor produced a 127% practice gain with the exam deficit largely eliminated. Offloading is therefore harmful only in the bypass configuration, and the operative design variable is whether the model substitutes for the graded target skill. ([[naim-bypass-offload-scaffold-llm-learning-2026]]) ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[ai-literacy]] — Knowing when to offload and recognizing reliance patterns - [[agency]] — Diminished when AI substitutes for the learner's cognition - [[critical-thinking]] — Degraded by uncalibrated offloading - [[creativity]] — Undermined when AI replaces the learner's generative act - [[distributed-cognition]] — The efficiency–regulation trade-off of delegation - [[embodied-learning]] — Situated cognition as the alternative to outsourcing - [[generative-ai]] — The tool context of AI-mediated offloading - [[metacognition]] — Calibrating when to offload vs. engage - [[online-teaching-and-learning]] — A medium vulnerable to the "ask for the answer" shortcut - [[prompt-engineering]] — The primary mechanism of offloading in LLM use - [[scaffolding]] — Reducing load without eliminating cognitive demand - [[self-directed-learning]] — Eroded by thoughtless AI use - [[self-regulated-learning]] — Regulating offloading decisions - [[trust-calibration]] — Knowing when to trust and when to question AI - [[retrieval-spacing-interleaving]] — the counter-practice to letting a model retrieve on the learner's behalf - [[cognitive-surrender]] ## Connected Articles - [[yan-cognitive-outsourcing-genai-assessments-2026]] — From cognitive outsourcing to reallocation: 3P analysis of student–GenAI engagement in unsupervised assessments (Yan et al. 2026) - [[family-school-autonomy-support-genai-2026]] — Family-School Autonomy Support for Children's Responsible Use of Generative AI - [[zohar-bloom-inzlicht-against-frictionless-ai-2026]] — Excess vs beneficial friction: why AI offloading differs from earlier tools - [[thermomix-genai-education-analogy-2026]] — With a Thermomix You Lose the Ability to Cook: a kitchen-machine analogy for generative AI in education (Rummel, Nachtigall & Panadero 2026) - [[cognitive-washout-ai-skill-decay-2026]] — cognitive washout: post-withdrawal dynamics of AI-induced skill decay - [[yan-conversational-ai-engagement-dependence-synthesis-2026]] — A Critical Narrative Synthesis of Conversational AI Engagement and Dependence - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[du-yuan-epistemic-dependence-2026]] — Epistemic dependence in AI-mediated learning: six diagnostic criteria separating productive reliance from harmful dependence (Du & Yuan 2026) - [[seung-basham-cognitive-offloading-swld-2026]] — GenAI cognitive offloading for students with learning disabilities - [[lodge-loble-cognitive-offloading-2026]] — AI, cognitive offloading and implications for education - [[cognitive-offloading-metacognitive-review-2026]] — Meta-cognitive insights into cognitive offloading (Guo & Ye 2026) - [[metacognitive-training-optimal-cognitive-offloading-2026]] — Metacognitive training facilitates optimal cognitive offloading (Ngai & Gilbert 2026) - [[critical-thinking-paradox-genai-learning-2026]] — The critical-thinking paradox in GenAI-integrated learning - [[cognitive-offloading-speedup-illusion]] — The speedup illusion of AI-assisted work - [[misiejuk-cognitive-offloading-prompting-2026]] — Co-occurrence network analysis of prompt patterns - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and cognitive surrender: how AI reshapes human reasoning (Shaw & Nave 2026) - [[young-people-learning-generative-ai-rapid-review-2026]] — Surrender-offloading-agency continuum for GenAI - [[ai-making-us-stupid]] — Research on AI's cognitive effects and learning displacement - [[brcic-effortless-trap-productive-struggle-2026]] — The Effortless Trap: AI replacing cognitive work (Brcic & Frljic 2026) - [[genai-thoughtless-use-self-directed-learning-2026]] — Thoughtless AI use erodes self-directed learning - [[generative-ai-guardrails-harm-learning]] — High-school math RCTs on unguarded GenAI - [[stromberg-generative-ai-learning-penalty-secondary-2026]] — The generative AI learning penalty: homework outsourcing harms learning - [[coach-not-crutch-ai-writing]] — AI can work less and learn more (Lira et al. 2025) - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[cognitive-commons-ai-expertise-regeneration]] — The tragedy of the cognitive commons: AI and expertise regeneration - [[pause-ai-cognitive-offloading-self-reflection-2026]] — PAUSE: A Privacy-Preserving Self-Reflection Tool for AI-Associated Cognitive Offloading - [[chatgpt-writing-cognitive-impact-2026]] — a writing-to-learn experiment in which ChatGPT-assisted writing produced less knowledge gain than an unaided control - [[student-cognitive-offloading-ai-higher-ed-2026]] — Patterns of student cognitive offloading to AI in higher education: naturalistic ChatGPT message-level evidence (Piatnitckaia et al. 2026) - [[metacognitive-feedback-anti-deskilling-offloading-2026]] — Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants - [[chatgpt-programming-performance-retention-ownership-2026]] — Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership - [[adaptive-scaffolding-contingency-comet-tutor-2026]] — Adaptive Scaffolding Needs Contingency: An AI Tutor That Escalates and Fades on What the Learner Does - [[davor-ai-supported-learning-higher-order-outcomes-2026]] — Offloading tendency and verification literacy predicting higher-order outcomes through metacognitive self-regulation (Davor et al. 2026) - [[ren-metacognitive-awareness-genai-reliance-2026]] — A reflection prompt reduces acceptance of incorrect AI advice and improves awareness calibration (Ren 2026) - [[shojaei-genai-dependence-critical-thinking-employability-2026]] — GenAI dependence as a boundary condition on critical-thinking disposition and self-perceived employability (Shojaei et al. 2026) - [[niu-genai-children-creative-thinking-cognitive-development-review-2026]] — GenAI and children's creative thinking: over-reliance and prompt dependence in a scoping review (Niu et al. 2026) --- ## [Cognitive Surrender](https://edtechdev.github.io/aied/concepts/cognitive-surrender/) > **Cognitive surrender** — adopting an AI system's output without doing the reasoning that would let you evaluate it. Where [[cognitive-offloading]] asks *how much* mental work a learner delegates and whether the delegation is calibrated, surrender asks whether the learner keeps the evaluative role at all: the answer arrives, is accepted, and becomes the person's own position without the checking that would have constituted [[critical-thinking|judgment]]. The term comes from [[shaw-nave-cognitive-surrender-2026|Shaw and Nave's (2026) Tri-System Theory]], which locates it as a distinct failure from both offloading and [[ai-misuse-learning-harm|over-reliance]]. ## Questions to Consider 1. When a student accepts an AI answer that happens to be right, has the class succeeded? What would have to be true for the acceptance itself to be the thing that mattered? 2. Is a calibrated delegator who knows their own limits different in kind from a student who simply cannot tell? How would [[teacher-role|an instructor]] the difference in a single piece of work? 3. If confidence rises while accuracy falls — as it does when AI advice is available — what does that do to the learner's own error detection? 4. Where is verification teachable, and where is it a workload problem [[educational-policy-ai|policy]]? Which parts of a [[curriculum-design|curriculum]] would need to change for checking to be the norm rather than the exception? 5. Should institutions measure surrender at all, or does measuring it invite the same surveillance that [[ai-use-disclosure|disclosure rules]] struggle with? 6. What would a task look like that rewards a student for overriding a wrong AI answer rather than for producing a polished one? ## Introduction Surrender is the disposition behind the most common failure in AI-supported study: not using the tool too much, but taking its word for it. The [[cognitive-offloading]] page holds the older construct and its evidence, and [[ai-misuse-learning-harm|misuse-of-AI]] holds the harm inventory. This page exists because a distinct cluster of work now names and measures something those two do not quite capture. Three things distinguish surrender. It is about *judgment*, not effort: a student can spend considerable time [[prompt-engineering|prompting]], reformatting and refining output while never once recruiting their own standard for whether the output is any good. It is often *invisible to the learner*: the signaling that normally attends effort — hesitation, uncertainty, the sense of not knowing yet — is short-circuited by fluent text. And it is *not reliably fixed by practice or incentives alone*, which is where the evidence gets pointed. ## Offloading, over-reliance, surrender: three different failures | Construct | Core question | Typical harm mechanism | Where it is measured | | --- | --- | --- | --- | | [[cognitive-offloading]] | How much mental work is delegated, and is that calibrated? | Displaced practice: the learner skips the [[desirable-difficulties|beneficial friction]] that builds a skill | RT and accuracy under tool-present versus tool-absent conditions; response-time panels | | [[ai-misuse-learning-harm|Over-reliance]] | Is delegation chronic and uncalibrated rather than strategic? | Skill erosion, motivation and self-efficacy loss, literacy debt | Self-report scales, log data, performance on unaided transfer tasks | | Cognitive surrender | Is the learner's own [[evaluative-judgment|evaluative judgment]] engaged before the answer is accepted? | Unchecked acceptance of wrong output, inflated confidence, loss of authorship | Trials where AI advice is correct versus faulty; override and verification behavior | The three sit on one ladder rather than replacing each other. Offloading is often productive; it turns into over-reliance when it stops being strategic; surrender is the version in which the evaluative step goes missing. The distinctions are drawn in practice as well as theory: the [[young-people-learning-generative-ai-rapid-review-2026|Sydney rapid review]] of 271 empirical papers on PreK-12 [[generative-ai|generative AI]] organizes student behavior into exactly these three categories — surrender, offloading, and [[agency]] — and spreads them across cognitive, [[metacognition|metacognitive]] and affective dimensions. [[du-yuan-epistemic-dependence-2026|Du and Yuan (2026)]] cut the same boundary from the other side: instrumental assistance produces output, judgment-bearing assistance supplies the *standard by which output is judged*, and only the second kind reliably displaces the work that develops expertise. Their six criteria — contestability, recoverability, transfer, traceability, distributed responsibility and epistemic plurality — are a surrender diagnostic rather than a usage measure. ## The Tri-System account [[shaw-nave-cognitive-surrender-2026|Shaw and Nave (2026)]] extend dual-process reasoning with a third system. System 1 is fast and intuitive; System 2 is slow and deliberative; **System 3** is artificial cognition that runs outside the brain. System 3 can supplement internal processing, supplying candidate answers and flagging contradictions for System 2, or it can supplant it, in which case deliberation never begins. Surrender is the supplanting case, and the theory predicts it should be detectable as a signature rather than a feeling: adopt the AI's answer when it is right, adopt it when it is wrong, and report more confidence either way. ## The experimental signature Three preregistered experiments adapted the Cognitive Reflection Test and randomized *AI accuracy* using hidden seed prompts, so participants could not tell a good assistant from a bad one from the interface (N = 1,372; 9,593 trials). - **People consult, then adopt.** Participants chose to consult the assistant on a majority of trials (>50%). - **Accuracy tracked the AI, not the person.** Against a brain-only baseline, accuracy rose about 25 percentage points when the AI was accurate and fell about 15 points when it erred, the contrast the authors call the behavioral signature of surrender (Cohen's h = 0.81, 95% CI [0.72, 0.91]; per-study h = 0.83, 0.86, 0.78). - **The two behaviors separate cleanly.** Across incorrect-AI trials, 73.2% ended in surrender, 19.7% in [[cognitive-offloading|offloading]] (overriding the wrong advice and answering correctly), and 7.1% in failed overrides. - **Confidence rises with access.** Despite roughly half of the AI answers being faulty, access to the assistant raised stated confidence by 11.7 percentage points (77.0% versus 65.3%). - **Situational pressure does not remove it.** Per-item incentives plus feedback increased offloading by about 19 points to 37.1% and reduced surrender to 57.9%; time pressure cut offloading by about 12 points to 6.2% and raised failed overrides to 13.8%. Both conditions shifted baselines without closing the accurate-versus-faulty gap (OR = 14.28 under time pressure; OR = 11.05 under incentives and feedback). - **Susceptibility is a disposition, not just a situation.** [[trust-calibration|Higher trust]] predicted a wider accurate-versus-faulty gap (OR = 2.81) and, in the head-to-head model, surrender over offloading (OR = 4.36). Higher need for cognition cut the gap (OR = 0.83) and shifted outcomes toward offloading (OR = 0.46); higher fluid intelligence was protective in the same direction (OR = 0.69), and its strongest single effect was a lower tendency to follow faulty advice once the assistant had been engaged (OR = 0.23). The practical reading is that surrender is a **discrimination failure** rather than a motivation failure. The participants who surrendered were not trying to avoid work; they could not tell good advice from bad, and the interface gave them no reason to try. ## Surrender outside the lab The lab findings line up with field and observational work that independently reached for the same word. - **Population-level displacement.** [[generative-ai-reduced-study-time-math|Rismanchian et al. (2026)]] analyze 3.2 million ALEKS learning interactions and 12.2 million [[assessment|placement-assessment]] response times across a ten-year panel, with graph-based problems serving as controls for text-based ones. Students completed AI-susceptible work faster and scored higher on it while proctored retention items showed a 25% cumulative decline in the odds of a correct response; the authors call the pattern cognitive surrender at population scale. - **Verification as the discriminating variable.** [[stamatoulis-genai-use-patterns-2026|Stamatoulis et al. (2026)]] find that *how* students use generative AI predicts [[learning-gains|performance]] while *how much* they use it predicts neither: low-verification uptake — uncritically accepting output — was associated with worse outcomes, and evaluative integration with better ones. - **Thoughtless use erodes the resources self-direction needs.** [[genai-thoughtless-use-self-directed-learning-2026|Zhao and Gu (2026)]] model adopt-without-evaluating directly: across 487 undergraduates, thoughtless use of generative AI was associated with lower [[self-regulated-learning|self-directed learning]] (β = −0.42), partly through reduced [[motivation]] and [[self-efficacy]]. - **The neural and ownership trace.** [[your-brain-on-chatgpt-cognitive-debt-essay-writing|the cognitive-debt EEG study]] measured [[writing-education|essay writing]] with an [[llm]], a search engine and no tool: the LLM group showed the weakest neural connectivity, reported the lowest authorship of their own work, and struggled to quote what they had written. - **Surrender can arrive unsolicited.** [[ai-advice-suppresses-ikt-suspension-2026|Work on judgment suspension]] found that the mere availability of AI advice collapsed people's willingness to say "I don't know" (0.06 versus 0.36, and 0.03 versus 0.44 across two studies), while correctness fell and confidence rose from about 30 to 76. Stakes did not restore suspension, though incentives improved accuracy by encouraging overrides. A fourth study showed the effect without any consultation at all, the pattern of an autocomplete or a search summary rather than a chat prompt. - **Dependent and autonomous modes are distinguishable only in consequence.** [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] follow 589 students and early-career workers across three waves: dependent offloading delegated core thinking and tracked with lower intrinsic motivation, while autonomous offloading did not, and immediate performance looked identical in both. - **The model's own contribution.** [[sycophantic-ai-social-interaction-2026|Sycophancy]] is the complement to surrender: an assistant that agrees, hedges toward the user's framing, or states a claim fluently gives the learner nothing to push against. - **Withdrawal is the unresolved half.** [[cognitive-washout-ai-skill-decay-2026|Work on cognitive washout]] points out that almost all of this research measures cognition *during* AI use, and almost none measures what happens when the assistant is taken away, which is exactly the question an exam, an outage or a license review poses. ## Designing and teaching against surrender - **Make verification the assignment, not the advice.** Require a stated first answer before the AI is consulted, and require students to name the points where they overrode it. This turns an invisible judgment step into a gradeable artifact. - **Reward overrides, not usage.** In the experiments, incentives attached to accuracy raised offloading toward 37% without eliminating surrender; incentives work when they make checking consequential, which is also the finding behind [[brcic-effortless-trap-productive-struggle-2026|the Effortless Trap]]'s placement rule — an unguarded helper left students about 17% worse on an unaided exam, and a rebuilt model that withheld answers erased the harm. - **Teach discriminating, not just distrusting.** Because need for cognition and fluid intelligence predicted resistance while self-reported trust predicted vulnerability, the trainable target is the ability to tell good output from bad on a specific task. [[metacognitive-training-optimal-cognitive-offloading-2026|Calibration training]] with prediction-plus-feedback trials is the worked prototype: a brief intervention reduced reminder bias in a laboratory analogue of over-reliance. - **Surface disagreement rather than confidence.** Because access inflated confidence even on wrong answers, the useful design move is friction that presents conflicting evidence, uncertainties or counter-arguments, so the learner has something to evaluate instead of something to accept. [[human-in-the-loop-ai|Human-in-the-loop]] designs with verification prompts and [[ai-use-disclosure|disclosure requirements]] are the institutional counterparts. - **Keep an [[educational-measurement|measurement]].** If assisted performance and unassisted capability are not measured separately, surrender is invisible to course evaluation: the semester looks better and the retention numbers quietly fall. - **Treat it as a systems and design problem.** [[agentic-ai-pedagogical-best-practice-2026|Agentic designs]] note that the more [[agentic-ai|an agent]], the less cognitive effort the learner spends, converting productive delegation into surrender by default; [[ai-overreliance-complex-adaptive-system-2026|modeling work]] adds that reliance spreads through social proof, so cohort norms and interface defaults matter as much as individual habit. ## Open questions Does surrender habituate, or does it decay with disuse? [[cognitive-washout-ai-skill-decay-2026|Washout research]] formalizes the post-withdrawal question but the long-run data do not exist yet. Can surrender be measured in the wild without invasive logging — and do self-reports of checking behavior predict anything? Does the confidence inflation persist once a learner has been burned by a confident wrong answer, or is [[trust-calibration|miscalibration]] sticky? And at the level of a profession rather than a person, does widespread surrender deplete the shared expertise pool that oversight of AI systems itself depends on, the concern raised by [[cognitive-commons-ai-expertise-regeneration|the cognitive commons framework]]? ## Limitations The evidence for surrender as a named construct rests heavily on one program of laboratory work: three preregistered experiments on an adapted Cognitive Reflection Test with convenience samples, which establishes the mechanism and the dispositional moderators but [[transfer-of-learning|transfer]] high-stakes professional judgment. The design also captures single exposures, so nothing in it shows whether surrender habituates or compounds. The field evidence is observational: the ALEKS panel shows faster completion with worse proctored retention, but a causal attribution to surrender rather than to study-strategy change requires assumptions the design cannot test. Definitions are still unstable across the literature — offloading, over-reliance, dependence, attachment and problematic use are routinely conflated, and [[yan-conversational-ai-engagement-dependence-synthesis-2026|at least one synthesis]] argues that frequent delegation should not be labeled dependence without impaired control or harm. Finally, [[self-report-measures|self-reports]] of trust, need for cognition and checking behavior carry the usual limits, and much of the classroom evidence concerns [[higher-ed|higher education]] and [[k-12|secondary]] rather than early schooling. ## Connected Concepts - [[cognitive-offloading]] - [[ai-misuse-learning-harm]] - [[critical-thinking]] - [[agency]] - [[metacognition]] - [[self-regulated-learning]] - [[ai-literacy]] - [[trust-calibration]] - [[human-in-the-loop-ai]] - [[ai-sycophancy]] - [[reducing-ai-misuse]] - [[theory-development-aied]] - [[generative-ai]] ## Connected Articles - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and the experimental signature of surrender - [[generative-ai-reduced-study-time-math]] — population-level displacement in ALEKS - [[young-people-learning-generative-ai-rapid-review-2026]] — the surrender/offloading/agency framework in PreK-12 - [[du-yuan-epistemic-dependence-2026]] — instrumental versus judgment-bearing assistance - [[ai-advice-suppresses-ikt-suspension-2026]] — the collapse of "I don't know" - [[your-brain-on-chatgpt-cognitive-debt-essay-writing]] — neural engagement and ownership - [[genai-thoughtless-use-self-directed-learning-2026]] — thoughtless use and self-directed learning - [[stamatoulis-genai-use-patterns-2026]] — low-verification uptake versus evaluative integration - [[cognitive-washout-ai-skill-decay-2026]] — what happens after withdrawal - [[ai-overreliance-complex-adaptive-system-2026]] — reliance as a population process - [[brcic-effortless-trap-productive-struggle-2026]] — placing AI so the task stays effortful - [[metacognitive-training-optimal-cognitive-offloading-2026]] — calibration training - [[family-school-autonomy-support-genai-2026]] — dependent versus autonomous offloading --- ## [Framing AI Use for Students](https://edtechdev.github.io/aied/concepts/framing-ai-use-for-students/) > **Framing AI use for students** — the persuasive and communicative craft of shaping how learners understand the value, purpose, and boundaries of AI tools and policies, so that they adopt productive and [[ethics|ethical]] use rather than rejecting, avoiding, or gaming it. It is the "buy-in" lever that [[reducing-ai-misuse]]'s educative interventions depend on: structural [[guardrails]] change the environment, but [[scaffolding]], literacy training, and AI-use policies only take hold when students are actually convinced of their point. ## Questions to Consider - Simply making AI available to students was associated with ineffective or even unethical use, while explicitly framing and embedding appropriate use improved outcomes. If the presence of the tool matters less than how it's framed, what does that suggest about the emphasis on simply 'giving students AI'? - Students' knowledge of [[governance|institutional]] AI rules shows only weak links to what they actually do — most use generative AI, many are unsure if their usage complies, and they lean on privately accessed tools. Why do rules fail to change behavior, and what might persuade instead of just inform? - When institutions respond to AI with fear and condemnation, students may hide or rationalize their use rather than learn to use it well. Have you seen a 'moral panic' response push students into secrecy? What would a reframed, opportunity-focused approach look like? - Anxiety about AI isn't purely a barrier: students who worried about accuracy and plagiarism were *more* likely to verify and revise AI output rather than accept it uncritically. How might productive anxiety be channeled into evaluative competence instead of being suppressed? - Students construct their own sense of what's acceptable through 'sites' — faculty intentions, course documents, peer norms, and institutional messages — that often diverge. When those messages conflict, which one do you think actually wins, and how does framing close that gap? ## Introduction The concept sits between two more familiar ones. Where [[technology-acceptance-model]] *predicts* uptake from perceived usefulness and ease of use, framing is the *active practice* of shaping those perceptions. And where [[student-experience]] describes how students currently perceive AI, framing is about changing that experience deliberately. It is the communication-side partner to [[educational-policy-ai]]: a policy is only as effective as students' willingness to buy into it. ## Why framing matters Framing matters because **policy and rules do not reliably change behavior on their own**. Survey [[research-methods-aied|research]] on [[regulation|regulatory]] awareness finds that students' knowledge of institutional GenAI rules shows only weak-to-moderate associations with what they actually do — most students use [[generative-ai|generative AI]] tools, over half are unsure whether their usage complies with institutional regulations, and they lean on privately accessed tools rather than institutionally provided ones.([[student-regulatory-awareness-genai]]) Knowing the rules is necessary but not sufficient; the message has to *persuade*, not just inform. The frame also shapes whether students experience AI as a **threat to be avoided or evaded** versus a **resource to be used deliberately**. When institutions respond to AI with fear and condemnation — what one line of work calls a recurring "moral panic" — they push students into hiding or rationalizing their use rather than learning to use it well.([[moral-panic-genai-classroom]]) Reframing anxiety and condemnation into structured opportunity changes the whole dynamic of [[student-engagement|student engagement]] with AI. ## Strategies and evidence ### Frame AI as a productive tool, not a threat to be banned The most directly tested framing intervention in the knowledge base is a six-year natural experiment tracking a data-[[visualization]] course across three conditions: pre-GenAI, GenAI-available (present but unintegrated), and GenAI-integrated (explicit instruction + encouragement to use AI on applied work, banned only on the knowledge-check portion). The finding: simply *making* AI available was associated with students using it ineffectively or unethically, while *explicitly framing and embedding* its appropriate use recovered and improved outcomes on applied questions.([[moral-panic-genai-classroom]]) The lesson is that the framing (how AI is positioned and taught) matters as much as the tool's presence — encouraging appropriate use beats condemning or ignoring it. ### Message expectations clearly and repeatedly — and design around students Because rule-awareness alone does not shift behavior, effective framing pairs clear expectations with structural reinforcement. Students construct their own sense of what is acceptable through what one interview study calls the "sites" where AI policy is interpreted — faculty intentions, course documents, peer norms, and institutional messages often diverge, producing rationalizations like "copying AI text is victimless."([[student-rationalization-ai-writing]]) Framing must therefore close the gap between what faculty intend and what students infer, and acknowledge the social and emotional context — shame and guilt regulate when and how students make AI use visible, driving hiding behaviors and selective disclosure rather than honest engagement.([[shame-guilt-ai-regulation-computing-education]]) ### Say why, not just what The most direct framing lever in the knowledge base is supplying the *reason* for a boundary task by task, composed as though speaking to students ("you will…", "we will…"). [[mccorkle-aligned-genai-course-policy-2025|McCorkle's (2025)]] design case documents the mechanism: after office-hour conversations revealed that students who had read a prohibition did not believe it applied to them, she rebuilt the policy so that every allowed or unallowed use of GenAI was justified by the specific [[assessment]] it protects — for example, GenAI help with learning objectives is unallowed because "I am assessing your ability to compose learning objectives that are specific, measurable, and at an appropriate level" — and delivered that rationale both in the syllabus and as just-in-time reminders inside assignment instructions. Framing by rationale is an [[equity-in-ai-education|equity]] move as much as a persuasive one: it assumes less shared background knowledge about authorship, originality, and attribution, and it turns policy into a dialogue rather than a verdict ([[academic-integrity]]). ### Reframe anxiety and uncertainty into evaluative competence A [[mixed-methods-research|mixed-methods]] study of academic [[writing-education|writing]] found that AI anxiety is not simply a barrier: students who worried about accuracy and plagiarism were *more* likely to verify, cross-check, and revise AI output rather than accept it uncritically. The study frames [[ai-literacy|AI literacy]] less as acceptance and more as **regulatory competence** — the capacity to question outputs, revise selectively, and maintain authorship responsibility.([[ai-anxiety-strategic-regulation-writing-2026]]) Framing AI use for students means channeling productive anxiety toward evaluation, not suppressing it. ### Use targeted messages to shape specific behaviors Small, well-designed messages can shift behavior. An **inoculation message** about ChatGPT's fallibility increased students' intentions to verify AI-provided information and their actual verification behavior.([[chatgpt-inoculation-training-verification-2026]]) Likewise, simply warning students about AI fallibility increased [[help-seeking]] in an [[intelligent-tutoring|intelligent tutoring]] system — a frame of *calibrated caution* rather than blanket distrust.([[ai-fallibility-warning-help-seeking]]) These point to a general principle: frame the tool's limits honestly, and students calibrate their behavior accordingly rather than either over-trusting or rejecting it. ### Secure buy-in and take-up, not just access Framing is upstream of take-up. Field experiments on AI tutoring found the binding constraint was not capability but *engagement*: despite dedicated session time, nearly half of students never used the platform, and users averaged only 2–5 minutes per week — until a low-cost human-support intervention (a brief in-person onboarding) improved take-up.([[access-not-enough-ai-tutoring-2026]]) Getting students to buy into the value and purpose of a tool is a precondition for any learning benefit; framing includes selling that value, not just removing access barriers. ## Framing and motivation Framing connects to [[motivation]] through [[self-determination-theory]]: how a tool or policy is *presented* determines whether students experience AI use as autonomous and purposeful or as controlled and imposed. Students' engagement with GenAI is shaped by whether the tool supports their sense of competence, autonomy, and relatedness — a framing question as much as a feature question.([[students-engagement-with-generative-ai-in-academic-learning-a-self-determination]]) When AI availability erodes the perceived point of effort — "why put in this much effort?" — the frame must rebuild a *purpose* for that effort, linking AI use to durable learning rather than task completion.([[ai-availability-student-motivation]]) ## Media and public framing Students are also framed by the wider media and public discourse around [[ai-education|AI in education]], which shapes their baseline expectations before any instructor message. Analyses of public discourse and how platforms like YouTube frame ChatGPT use in education show that prevailing frames — hype, doom, or pragmatism — influence how learners and [[teacher-role|educators]] approach the technology.([[youtube-frames-chatgpt-education]])([[ai-ethics-education-public-discourse]]) Effective framing by instructors often means deliberately countering or redirecting these ambient narratives. ## Practical guidance - **Position AI as a productive resource** by explicitly teaching *when* and *how* to use it, rather than banning or ignoring it — integrated framing beats both condemnation and laissez-faire. - **Message expectations repeatedly and from every "site"** — syllabus, assignment prompts, [[feedback]], and peer norms should tell a consistent story so students don't invent their own rationalizations. - **Channel anxiety into evaluation.** Frame uncertainty about accuracy as a reason to verify and revise, not as a reason to avoid or cheat. - **Use honest, targeted messages** (e.g., inoculation and fallibility warnings) that build calibrated caution rather than blanket trust or distrust. - **Sell the purpose.** Connect AI use to durable learning and learner [[agency]], and pair messages with the support that converts intent into take-up. ## Connected Concepts - [[reducing-ai-misuse]] - [[ai-literacy]] - [[student-experience]] - [[technology-acceptance-model]] - [[educational-policy-ai]] - [[academic-integrity]] - [[motivation]] - [[self-determination-theory]] - [[agency]] - [[trust-calibration]] - [[student-engagement]] - [[misconceptions]] ## Connected Articles - [[mccorkle-aligned-genai-course-policy-2025]] — Transparent rationale for each allowed and unallowed GenAI use, task by task (McCorkle 2025) - [[moral-panic-genai-classroom]] — Encouraging appropriate use of GenAI rather than condemning it as disruption - [[student-rationalization-ai-writing]] — The "five sites" where students rationalize AI use in academic writing - [[student-regulatory-awareness-genai]] — Knowing the rules is not enough: regulatory awareness and actual use - [[ai-anxiety-strategic-regulation-writing-2026]] — From AI anxiety to strategic regulation - [[chatgpt-inoculation-training-verification-2026]] — Inoculation messaging and verification behavior - [[ai-fallibility-warning-help-seeking]] — Warning about AI fallibility increases help-seeking - [[shame-guilt-ai-regulation-computing-education]] — Shame and guilt as social regulators of AI use - [[access-not-enough-ai-tutoring-2026]] — Human support improves engagement with AI tutoring - [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination]] — SDT-based engagement with GenAI - [[ai-availability-student-motivation]] — How AI availability shapes student motivation - [[youtube-frames-chatgpt-education]] — How YouTube frames ChatGPT use in education - [[ai-ethics-education-public-discourse]] — Public discourse on AI ethics in education - [[ithaka-sr-ai-skills-college-graduates-2026]] — Instructors frame AI as critical/responsible use; employers frame it as productivity - [[ssaho-ai-academic-integrity-review-2025]] — Building a culture of academic integrity via clear expectations - [[generative-ai-mediational-agent-sociocultural-2026]] — Human-first habits of participation in AI-mediated learning - [[caeai-ai-companions-learning-over-performance-2026]] — designing companions that protect effortful learning --- ## [Reducing AI Misuse](https://edtechdev.github.io/aied/concepts/reducing-ai-misuse/) > **Reducing AI misuse** — the design, [[pedagogy|pedagogical]], and [[educational-policy-ai|policy]] levers that prevent students from substituting [[generative-ai|generative AI]] for their own [[cognitive-offloading|cognitive work]] and instead steer them toward [[ethics|ethical]], productive use. Effective approaches are sorted by impact rather than popularity, and the strongest evidence favors **structural levers** — tool [[guardrails]] and assessment redesign — that change the environment so misuse is harder, regardless of a student's motivation, over **educative levers** that rely on building durable capacity and [[framing-ai-use-for-students|student buy-in]]. ## Questions to Consider - Students who outsource homework to AI can see their homework scores *rise* while their [[summative-assessment|closed-book exam]] scores *fall*. Before you read, why might performance and learning diverge so sharply — and what does that gap suggest about what 'success' with AI actually means? - The page ranks interventions by causal evidence and reach, and it places *structural* levers (tool guardrails, assessment redesign) above *educative* ones ([[teacher-role|teaching]] good practice) — precisely because structural levers work regardless of a student's motivation. Do you agree that changing the environment beats changing minds? What's the risk of relying on each? - A guardrailed 'hint-not-answer' tutor eliminated the learning harm that an unguarded one caused, even though both seemed helpful. Think of a learning tool you've seen that gives answers too readily. Where's the line between a hint that scaffolds and an answer that substitutes — and can you state it before reading? - The strongest fix includes assessment redesign: unassisted in-class exams, oral defenses, process artifacts, reasoning rewarded over surface fluency. How would you feel taking a course graded this way, and does that feeling tell you something about why this lever is both effective and unpopular? - AI declaration frameworks that force students to map their use to cognitive stages (planning vs. content generation) shift emphasis from policing to professional practice. Do you think such reflection genuinely builds better judgment, or does it just teach students how to describe misuse more cleverly? - Set a goal before reading: identify one concrete change in your own course or workflow that would make AI misuse harder, and one that would make productive use easier. Which of the page's tiers would each belong to? ## Introduction The concept rests on the evidence that AI misuse actively harms durable learning — the performance–learning gap documented in [[ai-misuse-learning-harm]] — even while inflating immediate performance. Interventions therefore target the mechanisms of that harm: answer-copying, [[cognitive-offloading]], motivation erosion, and learning displacement. They are not mutually exclusive; a robust approach combines a structural floor with educative capacity-building. ### Why structural levers matter most Interventions can be sorted by **causal evidence × structural reach × scalability × [[sustainability]]**. On this basis, the two *structural* levers rank highest because they work whether or not students choose the right behavior — they constrain the environment rather than depending on internal motivation. The *educative* levers are essential but only effective when students buy in, so they are treated as the second tier despite their conceptual promise. [[coates-governing-academic-integrity-indicators-2025|Coates, Croucher and Calderon (2025)]] push the structural argument up a level, treating governance rather than student behavior as the binding constraint and proposing a 130-item integrity indicator framework so that academic governors can see whether assessment is authentic, whether students are known individually by the teachers who assess them, and whether integrity features in induction and orientation. Their reforms target governance architectures, people, and [[ai-technologies|technologies and resources]], and they propose borrowing [[guardrails|red teaming]] from cyber security to expose assessment vulnerabilities before students find them — a reminder that the structural floor is maintained by institutions that need both the information and the will to maintain it. ### Tier 1 — Directly proven to reduce learning harm **Guardrailed AI tool design ("hint-not-answer" [[scaffolding]]).** In the strongest causal finding in the knowledge base, a field [[rct]] showed an unguarded ChatGPT-style tutor raised assisted practice performance **+48%** but reduced unassisted exam scores **−17%**, while a guardrailed tutor (hints instead of answers, plus teacher-authored problem information) eliminated the harm entirely. This mechanically prevents the answer-copying "crutch" behavior behind the damage. Activities include hint-not-answer tutoring, seeding prompts with correct solutions and common [[misconceptions]], and requiring a student attempt before AI output is revealed. **Assessment redesign (AI-resistant + unassisted measures).** Because misuse harm is assessment-dependent — surfacing on proctored, closed-book, and unassisted measures while inflating ordinary graded coursework — changing what counts as achievement both deters misuse and surfaces it. Activities include unassisted in-class exams and oral defenses, requiring process artifacts (drafts, reflections, annotated reasoning), rewarding reasoning over surface fluency, and designating AI-free zones. [[ivory-psychology-assessment-integrity-2026|Ivory et al. (2026)]] give the process-artifact requirement a concrete instrument: mandated version histories and reproducible analysis documents, so a suspected submission can be inspected as a timeline of how the work developed, and the amount of fabricated material a student would have to generate rises far beyond what outsourcing saves them. Large-scale field evidence underscores this: [[stromberg-generative-ai-learning-penalty-secondary-2026|Strömberg, Lei, & Wu (2026)]] found that homework outsourcing raised homework scores 18% while *lowering* closed-book exam scores 20% — exactly the signal that unassisted, proctored measures are designed to surface, and the study recommends weighting closed-book in-person assessment more heavily. [[leaton-gray-ai-digital-cheating-ethical-pedagogies-2025|Leaton Gray, Edsall and Parapadakis (2025)]] give the same logic its most explicit situational-crime-prevention statement, arguing that the failure belongs to assessments a model can answer convincingly rather than to students, and reporting a case in which 25 prevention techniques applied to an Australian business capstone — tracked student interactions, red-flag detection of too-expert work, random team reassignment, weekly in-class invigilated tests, takedown notices for course materials students had published — reportedly cut misconduct cases from 183 to 27 within a year. Their prescriptions are motivational rather than investigative: raise the perceived purpose of assessment, build [[self-efficacy]], and increase the perceived social cost of cheating, with the authors warning that addressing only one or two of the three will fail, supported by five discipline-specific redesigns that turn timed problem-solving exams into open-book [[problem-solving|problem solving]] with a written reflection on the student's own process, summative essays into [[collaborative-learning|collaborative]] archival research projects, and [[eportfolio|portfolios]] into iterative peer-reviewed design processes. The structural levers have an enforcement arm, and its evidence base is weaker than its reach suggests. [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] coded 1,162 generative-AI misconduct cases and found that the evidence most often cited at the point of allegation — [[ai-detection|detector]] output, similarity reports, AI-typical content patterns — carried the weakest probative value, while admissions, observed exam behavior and independently verified fabricated references were the strongest; because institutional procedures set no minimum evidentiary threshold, evidence quality bore no reliable relationship to case outcomes, and students whose cases rested on thin evidence had recourse mainly to appeals. [[wright-transcription-not-generation-2026|Wright (2026)]] shows a second cost of the same imprecision: prohibitions written around "generative AI" rather than around function capture tools that merely convert the format of work a student has already authored, so transcription-only use can be sanctioned as misconduct — an over-inclusion that falls hardest on disabled and [[equity-in-ai-education|equity]]-exposed [[learners]] and does nothing for [[assessment-validity|validity]]. [[harerimana-remote-proctoring-nursing-scoping-2026|Harerimana, Mtshali and Mchunu (2026)]] add that [[remote-proctoring|remote proctoring]], the most widely adopted integrity control, rests its deterrence case on how students say they feel rather than on reduced misconduct, and imposes anxiety, concentration loss and infrastructure-related exclusion that is not shown to be answered by lower misuse. [[li-genai-assessment-language-equity-2026|Li (2026)]] adds the rule-design dimension: because the same interface performs permitted language editing and prohibited substantive drafting, undifferentiated GenAI rules impose cohort-skewed compliance burdens on students who use English as an additional language, so the more defensible prevention is a purpose-based support-substitution boundary anchored in the assessment construct, with calibrated disclosure, proportionality staged across threshold, evidentiary and sanction decisions, and design levers — staged submissions, short construct-aligned oral verification, critique-based questions — doing the work detection cannot. Prevention also has to cover the grader itself: [[humble-prompt-injection-ai-grading-red-team-2026|Humble's (2026)]] red-team evaluation found that two of five indirect prompt-injection strategies embedded in a submission file raised a failing essay to a pass undetected, at reported success rates of 100% and 94%, so where AI tools grade student work, [[human-in-the-loop-ai|human review]], clearer separation of instructions from submitted content and resilience testing belong in the structural floor rather than in an optional tier. ### Tier 2 — Strong framework support, high potential **Scaffolded use sequences: think first, AI second, reflect third.** Eight design principles for integrating LLMs without displacing [[critical-thinking]]: preserve cognitive friction, position AI as a provisional thinking partner rather than an authority, embed evaluation checkpoints, require [[metacognition|metacognitive]] journaling and prompt logs, and balance AI-mediated with AI-free phases. Correlational evidence (independent work before AI produces stronger outputs) is strong; it is the pedagogically complete version of Tier 1. **[[ai-literacy|AI literacy]] and [[prompt-engineering|prompting]] literacy with deliberate practice and immediate [[feedback]].** A [[k-12]] module using scenario-based prompt practice with an [[llm]] auto-grader improved actual prompting skills and raised confidence in using AI for learning **+10.4%**, with 87% reporting they learned how to use AI responsibly. Demonstrated skill gains; the open question is whether these convert into downstream [[learning-gains|learning outcomes]]. It also addresses the [[equity-in-ai-education|equity]] gap in prior AI access. **Structured AI-use declaration frameworks.** Replacing generic "I used AI" checkboxes with [[discipline-specific-aied|domain-specific]] declarations that map use to cognitive stages (structural planning vs. content generation) forces reflection on the learning process and clarifies the boundary between acceptable assistance and misconduct, shifting the emphasis from policing to professional practice. Educative levers look different again once integrity is treated as a practice to be taught rather than a rule to be enforced. [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] argues that detection- and verification-based integrity regimes are misaligned with GenAI-augmented work, and reframes integrity as pedagogical practice enacted through [[evaluative-judgment|evaluative judgment]] — the learner's capacity to weigh options, justify academic choices and take responsibility under uncertainty — made visible through artifacts such as annotated decision trails, documented verification, oral defense and draft-to-draft version histories, with detection demoted to one supplementary layer rather than the primary infrastructure. [[mulisa-students-genai-integrity-perspectives-2026|Mulisa and Mezgebu (2026)]] find the same opening from the students' side: across 27 interviews at an Ethiopian [[higher-ed|university]], use was near-universal, most participants credited GenAI with raising their achievement, and the sharpest complaint ran the other way — AI users scoring above diligent independent workers, which participants described as demotivating. The authors conclude that students' beliefs predict ethical use more strongly than institutional rules do, so awareness, clear policy and assessment redesign have to arrive together — which is the page's educative tier stated as a condition: capacity-building changes behavior only through [[framing-ai-use-for-students|student buy-in]] accompanied by a structural floor. [[ji-student-voices-academic-integrity-scoping-2026|Ji's (2026)]] [[meta-analysis-systematic-review|scoping review]] of 38 studies of student voices supplies the field-level version of the same conclusion: the reviewed studies converge on a shift from retrospective detection to proactive ethical reasoning, on teaching AI ethics and AI-giarism in the [[curriculum-design|curriculum]] rather than in one-off [[ai-literacy|literacy]] sessions, and on detailed guidelines co-created by leaders, faculty and students, because unclear guidance is read as tacit permission rather than as caution. Ji's own reading is that ownership and moral reasoning motivate students more than fear of punishment, and that students are neither passive recipients of GenAI nor culprits without moral constraint but moral agents working a gray zone — this tier's premise, stated as a finding. ### Tier 3 — Promising, lower direct causal evidence **Metacognitive and self-assessment interventions.** Reflective journals, prompt logs, and calibration training rebuild the "absent cognitive baseline" of AI-native students who cannot locate their own cognitive boundary because AI-generated fluency masks it. Conceptually central but not yet causally tested. **Motivation redesign.** Because AI availability erodes [[motivation|autonomous motivation]] ("why put in the effort?"), restructuring tasks around goals AI cannot fulfill and around [[agency|learner agency]] directly targets the persistence erosion that compounds the direct harm. **Critical AI literacy.** A power-knowledge framing that teaches learners to interrogate, challenge, and participate in [[governance|AI governance]] rather than consume it. Long-term, equity-oriented, and structural in its ambitions, though its learning effects are largely untested. **Recognizing [[ai-sycophancy|sycophancy]] to prevent uncritical acceptance.** Because [[ai-sycophancy|sycophantic]] AI validates rather than challenges the user, it is a direct misuse vector: students who receive affirming agreement for incorrect thinking are encouraged to substitute AI for their own [[cognitive-offloading|cognitive work]]. [[contextual-sycophancy-ai-literacy|Contextual sycophancy]] shows AI literacy and prompting training reduce but do not eliminate the error loop, so misuse prevention must pair educative recognition training with system-level corrective-friction design (see Tier 1 guardrails). - **[[desirable-difficulties|Productive friction]] and pedagogical function.** The Sydney rapid review argues GenAI undermines learning when it lets students bypass the cognitive/metacognitive friction needed to learn, and that purposefully designed tools introduce *productive friction* (withholding answers, prompting explanation). It distinguishes four pedagogical functions — learning *from*, *with*, *about*, or *by shaping* GenAI — each with different demands on agency and assessment.([[young-people-learning-generative-ai-rapid-review-2026]]) ## Connected Concepts - [[ai-misuse-learning-harm]] - [[cognitive-offloading]] - [[academic-integrity]] - [[assessment]] - [[ai-literacy]] - [[scaffolding]] - [[self-regulated-learning]] - [[metacognition]] - [[motivation]] - [[prompt-engineering]] - [[ai-sycophancy]] - [[trust-calibration]] - [[framing-ai-use-for-students]] - [[cognitive-surrender]] ## Connected Articles - [[ivory-psychology-assessment-integrity-2026]] — Version-control evidence trails and reproducible analysis documents as misuse deterrents (Ivory et al. 2026) - [[brcic-effortless-trap-productive-struggle-2026]] — The Effortless Trap: placement rule for AI use (Brcic & Frljic 2026) - [[generative-ai-guardrails-harm-learning]] — GenAI Without Guardrails Can Harm Learning - [[genai-performance-vs-learning]] — Distinguishing Performance Gains from Learning - [[ai-assessment-scale-reform]] — The AI Assessment Scale and Assessment Reform - [[critical-thinking-genai-scaffolding]] — Scaffolding Critical Thinking with Generative AI - [[aaai2026-prompting-literacy-k12]] — Learning to Use AI for Learning (K-12 AI Literacy Module) - [[genai-declaration-frameworks-higher-education]] — Structuring Transparency: GenAI Declaration Frameworks - [[absent-cognitive-baseline-2026]] — The Absent Cognitive Baseline - [[ai-availability-student-motivation]] — AI Availability and Student Motivation - [[ai-literacy-power-knowledge]] — AI Literacy: An Exercise in Power-Knowledge - [[contextual-sycophancy-ai-literacy]] — The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention - [[sycophantic-ai-social-interaction-2026]] — Sycophantic AI makes human interaction feel more effortful and less satisfying over time - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[ssaho-ai-academic-integrity-review-2025]] — Culture-building and assessment redesign over detection policing - [[young-people-learning-generative-ai-rapid-review-2026]] — Cognitive surrender, productive friction, and metacognitive inequity - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant assessment - [[stromberg-generative-ai-learning-penalty-secondary-2026]] — The generative AI learning penalty: homework outsourcing harms learning - [[ai-overreliance-complex-adaptive-system-2026]] — AI overreliance modeled as a complex adaptive system - [[burneo-can-edtech-close-learning-gaps-2026]] — Guardrails removed harm without improving exam scores - [[munoz-misconduct-allegation-evidence-2026]] — What misconduct allegation files actually contain as evidence - [[wright-transcription-not-generation-2026]] — Over-inclusive AI rules and the students they catch - [[sharma-judgment-visible-genai-assessment-2026]] — Integrity as evaluative judgment rather than compliance - [[mulisa-students-genai-integrity-perspectives-2026]] — Students on whether GenAI is a cheating tool or a learning partner - [[harerimana-remote-proctoring-nursing-scoping-2026]] — What remote proctoring does to students, and to equity - [[leaton-gray-ai-digital-cheating-ethical-pedagogies-2025]] — Situational crime prevention against AI-facilitated cheating: 183 to 27 cases, and five discipline-specific redesigns (Leaton Gray, Edsall & Parapadakis 2025) - [[ji-student-voices-academic-integrity-scoping-2026]] — What 38 studies of student voices recommend: co-created clarity and educative reasoning over detection (Ji 2026) - [[li-genai-assessment-language-equity-2026]] — Where language support ends and substitution begins for EAL writers (Li 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Prompt injection as a misuse vector against AI grading (Humble 2026) - [[coates-governing-academic-integrity-indicators-2025]] — Governance as the binding constraint on integrity: a 130-item indicator framework (Coates, Croucher & Calderon 2025) --- ## [Academic Integrity](https://edtechdev.github.io/aied/concepts/academic-integrity/) > **Academic integrity** — the ethical framework governing honest academic work in the age of AI. The knowledge base documents how the concept has been reframed by [[generative-ai|generative AI]]: from a problem of [[ai-detection|detecting dishonest output]] to a design problem of making honest work visible, verifiable, and worth producing. Academic integrity [[research-methods-aied|research]] in this space has evolved from detection-focused approaches toward fundamental assessment redesign, pedagogy-led governance, and [[ai-literacy|teaching students how to use AI well]] rather than merely policing whether they do. ## Questions to Consider - Think of the last time you heard 'AI cheating' discussed. Was the conversation about catching students or about designing assignments students would want to do honestly? Which emphasis feels more familiar to you, and why? - A polished, plausible piece of work can now be generated in seconds. If you can no longer judge a student's capability by the product they hand in, what would you need to see or hear to feel confident they actually learned it? - Research finds students often rationalize AI use ('copying AI text is victimless') rather than misusing it out of malice. What assumptions about students' motives does a purely punitive integrity policy make — and what might those assumptions get wrong? - One line of research treats students' AI use as a coordination problem: behavior shifts when peer expectations and assessment incentives change, not when rules are restated. What peer or design factors in your own context might be quietly shaping whether AI use is honest? - Studies show fear of penalty drives students to hide AI use, and that transparent students can even draw suspicion. If you were designing an 'AI use disclosure' form, what would make a student actually want to fill it out truthfully? - The same institutional AI policy is interpreted differently by students across cultures — culture, not policy wording, drives what feels wrong. How should an integrity policy be communicated to a culturally diverse cohort so expectations are actually understood? ## Introduction The arrival of generative AI has not created the need for academic integrity — it has made weaknesses in existing approaches harder to ignore. A polished, plausible product can now be generated in seconds, so **product resemblance is an increasingly unreliable signal of capability**. This shifts the integrity question from *"can we catch AI use?"* to *"can our [[assessment|assessments]] still warrant the inferences we draw about student learning?"* ### The evolution from detection to redesign - **Detection skepticism:** [[ai-detection]] research and [[governance|institutional]] analyses increasingly find that AI detection tools are unreliable and procedurally unfair. [[detecting-llm-generated-text-latent-prompt|LLM text detection]] faces fundamental limitations. Fully AI-generated submissions can pass through live examination systems largely undetected, and experienced markers do not reliably spot GenAI-authored work. Detection, at best, is a limited, situational tool — not a strategy of first resort. - **Detection's narrow margin, and cheating as planned behaviour:** [[leaton-gray-ai-digital-cheating-ethical-pedagogies-2025|Leaton Gray, Edsall and Parapadakis (2025)]] assemble the case that AI amplifies a vulnerability the sector already had — AI-generated text has passed as human-authored in up to 80% of cases, while the evidence they cite puts machine detection at about 80% against 78.4% for human reviewers, too narrow a margin to carry a misconduct finding. Their motivational analysis is the more distinctive contribution: working from Ajzen's Theory of Planned Behavior and Bandura's self-efficacy theory, they report Krou et al.'s (2021) meta-analysis finding that [[self-efficacy]] correlates negatively with cheating while actual ability does not correlate inversely with it at all, so a capable, confident student who reads an assessment as unfair may cheat to regain control. The conclusion follows the same inversion this page documents elsewhere — an assessment a model can answer convincingly is an intellectually trivial assessment, and the failure belongs to the assessment rather than to the student. - **Assessment redesign:** [[authentic-assessment]], [[beyond-detection-authentic-assessment-ai-2025|beyond-detection approaches]], and [[ai-assessment-scale-reform|the AI Assessment Scale]] shift the focus from catching AI use to designing assessments where AI use is either irrelevant, transparent, or required to demonstrate a specific capability. - **Structural vulnerability of grading:** [[biology-degree-integrity-genai-cheating-2026|Chan et al.]] provide a concrete case study of how current grading is structurally exposed to AI-mediated dishonesty. In a [[biology-education|biology]] department, instructors perceived only in-person proctored exams as minimally vulnerable; outside-of-class assignments were seen as highly vulnerable, leaving about a third of a student's grade highly vulnerable and 80% at least somewhat vulnerable. This frames the integrity problem as partly a *grading-design* problem, motivating rebalancing toward proctored or in-class assessment and [[authentic-assessment|authentic assessment]] designs that are harder to outsource. [[ivory-psychology-assessment-integrity-2026|Ivory et al. (2026)]] supply the discipline-level counterpart with a whole three-year psychology program: 36 of 40 assessments across 16 types produced passable content at minimum effort, and the four that failed were those requiring presence, visual media, or the student's own dataset. Their reading of *why* it passes is the part that generalizes — marking that rewards fluent structure and good-faith use of the right analysis condones fabricated values and hallucinated references, so the pass boundary rather than detection decides whether AI output is graded as achievement, and the assessment vulnerability "cannot be blamed solely upon student usage." - **The submission as an attack surface: prompt injection against [[automated-assessment|AI grading]].** [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] red-teams an everyday grading workflow and shows that a student submission can carry hidden instructions that change the grade an AI tool produces: two of five indirect injection strategies embedded in the submission file raised a failing essay to a pass with no visible warning to the user, at reported attack success rates of 100% (9/9) and 94% (17/18), both combining instruction manipulation, role-playing and obfuscation across docx, pdf and htm formats. The most basic attack failed completely, and its refusal was silent — the tool disabled the chat without reporting the attempt — while a single detected injection produced a reassurance that the tool would grade "only according to the official assignment instructions" before six further runs of the same file raised the grade anyway. The integrity consequence runs in both directions: a grade raised by a hidden instruction carries no [[assessment-validity|validity]] claim, the same technique could be used to degrade a submission with no durable trace in the output, and the human marker is left as the only real check on manipulation designed not to be visible. - **Validity as the organizing frame:** [[assessment-validity]] reframes integrity as an evidential problem. [[authentic-products-authenticated-processes-2026|Authentic assessment research]] introduces **construct substitution** — an AI-generated product is attributed to the student, so the assessment infers the tool's capability rather than the student's. The evidential question survives any AI policy: whether use is prohibited, permitted, or required, the assessment must still generate evidence warranting the inference being drawn. - **Variation-at-scale as a no-surveillance integrity mechanism:** [[varia-construct-equivalent-assessment-variant-generation-2026|VARIA (Lee 2026)]] [[benchmark]] the premise behind AI-Integrated Authentic Assessment (AIAA) — replacing surveillance-based proctoring with per-student task variation so copying is structurally useless. The integrity guarantee is conditional on LLMs generating variants that are surface-distinct and [[assessment-validity|construct-equivalent]]; VARIA's 600-variant pilot finds frontier models satisfy this only at the margin (joint score 0.81–0.88) while non-frontier models collapse (0.50–0.55), so "variation-at-scale cannot be solved by [[prompt-engineering|prompting]] alone." This gives the detection-vs-redesign debate an empirical, falsifiable check: the no-surveillance promise of authentic, task-varied assessment now depends on a measured (and still narrow) generation capability rather than an assumed one. - **Authenticating student reflection against GenAI:** [[5p-reflection-model-genai-2026|Kadel et al. (2026)]] argue that traditional reflection models can no longer authenticate student reflection once GenAI can author reflective prose, and embed integrity directly inside a reflection model — the 5P framework's Pitfalls stage explicitly handles plagiarism, [[hallucination-risk|hallucination]], and over-reliance, while its Process and Product stages require documenting prompts and validating outputs so that [[learners]]' own reasoning is distinguishable from AI-generated contributions. - **Policy development:** [[genai-policies-higher-ed-computing|Institutional AI policies]] and [[educational-policy-ai]] research examine how universities develop and communicate integrity expectations — and why abstract policy statements so often fail. - **Integrity as a governance problem, not a detection problem.** [[coates-governing-academic-integrity-indicators-2025|Coates, Croucher and Calderon (2025)]] relocate the binding constraint from student behavior to academic [[governance]], which they call resilient yet "not well positioned or poised" to meet GenAI-related threats to the authentication of student assessment. Their answer is an integrity indicator framework of 130 items across eight dimensions, running from Designing (19 items) to Improving (6), and comprising governance questions rather than psychometric scales — whether the institution's top board receives assessment-quality updates, what percentage of assessment resembles relevant and meaningful problems, what percentage of students are known individually by the teachers who assess them. It was developed through a multiyear research review, five Australian university case studies, prototyping, expert confirmation with 60 invited experts across six world regions and [[quantitative-research|quantitative]] piloting, and it is paired with reforms to governance architectures, people, and technologies and resources that the authors expect to pay out only under external pressure from [[regulation]], benchmarking and cross-institutional competition; they present the framework as formative and call for psychometric validation before it is used for comparison or [[educational-measurement|measurement]]. - **What the documents actually say: integrity dominates institutional AI policy.** [[institutional-ai-policy-health-informatics-2026|Eldredge et al. (2026)]] analyzed the AI policy and guidance documents of all 48 CAHIIM-accredited health informatics and health information management master's programs in the US: 40 (83%) had at least one qualifying document, and those documents framed AI overwhelmingly as a matter of preventing student misuse in coursework, centering academic integrity and acceptable use. Academic integrity was by far the most frequent keyword (n = 139, ahead of citation at 59, assessment at 50, and plagiarism at 38), while privacy, intellectual property, and regulatory themes were present but less common (HIPAA n = 5, FERPA n = 11, IRB guidance n = 8). Topic modeling over the same corpus returned four themes, the first being academic integrity and appropriate AI use, alongside generative AI in teaching, AI and data-tool use in research, and student engagement with conversational AI systems. The evidential value for this page is the quantification: integrity-centered framing is not just the rhetorical default of institutional AI policy, it is empirically the dominant content of it. The authors' sharpest observation is about where that attention stops, noting that comparatively little guidance covers AI use in applied learning, research, and simulated environments, where curricular and data-governance concerns intersect, so the documents govern coursework while the settings the degree is preparing students for go largely unaddressed. - **Integrity framing is giving way to task-based regulation:** [[chirikov-regulate-ai-syllabi-2026|Chirikov's (2026)]] longitudinal study of 31,000+ course syllabi shows instructors' AI policies shifting away from a purely integrity-based frame: academic-integrity mentions in syllabi fell from 63% (Spring 2023) to 49% (Fall 2025), while references to AI's impact on learning rose from 1% to 29%. Instructors increasingly regulate AI **by task type** — restricting it for drafting/reasoning (where AI would displace learning) and permitting it for editing/proofreading and study support — rather than applying a blanket integrity prohibition. This reframes integrity policy as a task-level design decision rather than a binary rule. ### The misconduct procedure and its evidentiary collapse The enforcement machinery that sits on top of integrity policy is itself a design decision, and generative AI has dismantled its evidentiary foundation. [[teichmann-detecting-undetectable-misconduct-2026|Teichmann (2026)]] argues that the misconduct procedure universities imported from plagiarism assumes prohibited use can be detected and proved — a premise the technology dissolves. Unlike text-matching software, which points to a copied source, an AI-text classifier identifies no source (none exists); it outputs a probabilistic judgment about style that degrades under paraphrase, systematically mislabels non-native English writing, and cannot be explained or cross-examined, while skilled or lightly edited use leaves no trace at all. Persisting anyway reverses the burden of proof (the student is asked to prove a negative), strains every element of procedural justice, and lands the harm of false accusation hardest on the already disadvantaged. The proposed remedy is twofold: an evidentiary standard under which a detector score alone never grounds a finding, with graduated, education-first responses, and a shift of institutional effort into validity-centered and [[authentic-assessment|authentic assessment]] design — the answer to undetectable AI being better assessment rather than better surveillance. The limit case has no machinery at all: [[ai-written-admissions-essays-penalized-2026|Isley, Gaebler and Goel (2026)]] document a US public policy master's program that prohibited generative AI assistance in a signed attestation and screened no submissions, where 56.1% of 2025 applicants submitted at least one essay a commercial detector classified as primarily AI-written and each flagged essay was associated with a 1.5 percentage point lower admission probability (p < .01), rising to 2.6 points (p < .001) once essay quality was held constant. With no detection, adjudication or enforcement step between the rule and the outcome, the penalty was delivered entirely through admissions readers' unwritten judgment — prohibition plus unaided reader judgment reproducing the discretionary, unexplained sanction a fair procedure exists to prevent. [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] supply the empirical counterpart by coding every GenAI misconduct case at one regional Australian university over three years: 1,162 cases carrying 1,855 evidence items, each rated for relevance, credibility and inferential force. Detector output was the evidence most reached for and least able to bear weight — Turnitin or similarity reports carried 100% Weak inferential force and standalone detector output 100% Low credibility, and detector evidence fell to 0.5% of items by 2025 as institutions absorbed what it could not prove — while the strongest evidence types were those that do not rest on probabilistic text classification: student admissions, observed prohibited exam behavior, and independently verified fabricated references, which were the largest structural shift at 22.3% of items. Their sharpest criticism is structural rather than evidentiary, since "there is no requirement for investigators to assess the probative quality of evidence before progressing an allegation, no minimum evidentiary threshold at any stage of the pipeline," so evidence quality bore no reliable relationship to case outcomes. They also record how policy caught up with practice: the assessment policy in force through 2024 said nothing about AI, and from January 2025 a revised policy permitted approved authenticity software while prohibiting the upload of student work to third-party AI-detection tools. A second failure is one of scope rather than proof. [[wright-transcription-not-generation-2026|Wright (2026)]] shows that prohibitions written at the level of platform identity rather than function capture non-generative format conversion — speech-to-text transcription, OCR, plain text to LATEX — alongside the generative drafting they mean to bar, even though the peer-reviewed computer science literature treats recognition and generation as distinct operations. The cost lands unevenly: students with conditions affecting fine motor control, handwriting legibility or typing accuracy have relied on exactly those tools, and as standalone voice-to-text products are discontinued or degraded, AI-powered transcription is filling the functional gap, so an over-inclusive rule removes a primary means of producing legible work, making the dispute an [[accessibility]] and [[equity-in-ai-education|equity]] one before it is an integrity one. Wright's remedy is a function-based definition of generative AI plus four operational criteria — fidelity, non-augmentation, traceability and attestation — that give a student a structured route to rebut a transcription-only allegation while leaving the burden of proof with the institution. Detection's evidentiary problem is also a validity problem. Because AI assistance is iterative and interwoven with drafting rather than outsourced wholesale, authorship and ownership of meaning come apart, and a sound submitted product may not establish that the student exercised the judgment the task was meant to target. Detector performance varies across tasks, disciplines and model versions, and disclosed or suspected AI use can act as a biasing cue for markers, so a detection response adds construct-irrelevant variance rather than removing it. Where integrity debates ask who produced the words, a validity account asks what claim about the student the performance warrants. ### The rationalization problem Students do not generally [[ai-misuse-learning-harm|misuse]] AI out of malice; they rationalize it. [[student-rationalization-ai-writing|Interview research]] identifies at least **five disconnect sites** where students' interpretation of AI policy diverges from faculty intent, and a taxonomy of **20+ distinct rationalizations** — from "copying AI text is victimless" to "text reflecting my beliefs is my own writing." These rationalizations are ad hoc, post hoc, and internally inconsistent, and they describe a "steep, ethical slippery slope" on which students slide far outside [[pedagogy|pedagogical]] goals. This is why [[misconceptions|student misconceptions]] about AI are the upstream cause of integrity violations, and why integrity education must address [[ethics|ethical reasoning]], not just technical skill. [[ji-student-voices-academic-integrity-scoping-2026|Ji's (2026)]] scoping review of 38 empirical studies of student voices supplies the synthesis this section rests on. Its four themes are ambiguity (generating a whole task reads as cheating, grammar checking and brainstorming are largely acceptable, and paraphrasing, outlining and translation sit in a gray area), ethical [[agency]] in which students build personal rules in a guidance vacuum they read as a silent approval of their actions, the gap between ethical awareness and practice, and diversity by gender, level, discipline and culture. The awareness-practice gap is the rationalization problem measured at scale: Huang et al.'s Ethical Dissonance Index sorted 522 Chinese students into four clusters, one of which frequently did what it viewed as illegitimate, and Ofem et al.'s structural-equation modeling of 4,679 Nigerian students found that positive perceptions of ChatGPT predicted dishonest use while positive integrity attitudes acted as a significant negative mediator. Ji reads students as neither passive recipients of GenAI nor culprits without moral constraint but active agents navigating an ethical gray zone, which is why the reviewed studies converge on co-created, educative and context-sensitive responses rather than punitive ones. ### Why policy alone fails: the coordination problem [[ethical-ai-higher-ed-game-theory|A coordination-game framework]] provides a mechanism-level account of why policy pronouncements rarely change behavior: students' AI use is a **coordination problem**, where individual choices depend on peer expectations and assessment design. The model's key finding is **non-linear threshold dynamics** — small, well-calibrated changes to reflective-assessment incentives can trigger rapid cohort-wide shifts toward responsible use, while weak or misaligned incentives let opportunistic practice persist. In practical terms, modest redesign (e.g., requiring students to reflect on their AI interactions) can have disproportionate effects where abstract rules have none. Peer accountability does not always point toward integrity, as [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] found in graded [[group-work|group assessment]]: seven of fifteen student groups deliberately reduced their GenAI use partly to avoid free-riding on groupmates, since group consequences were shared rather than self-contained, yet in five groups a permissive collective climate — "everyone in my group is using GenAI" — lowered the perceived [[ai-misuse-learning-harm|risk of misuse]] and inverted the very accountability mechanism group work is meant to create. The same study found students reframing originality as faithfulness to the understanding their classmates built together, and its practical recommendation follows directly: make the negotiation of acceptable AI use an explicit, documented, and assessable outcome rather than leaving the norm to emerge from peer pressure or perceived risk. Policy inconsistency across the institution produced exactly the guessing the coordination account predicts. Nine of the eleven interviewees in [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al. (2026)]] framed their own course's explicit permission to use generative AI as a possible "trap" — a lure to identify students who could not resist — even though the policy was open and written into the assessment guidelines; students generalized from bans and warnings encountered elsewhere rather than reading their own course's rules on their own terms. The authors' conclusion follows the mechanism rather than the wording: per-course clarity was not enough, and [[educational-policy-ai|program- or institution-level]] consistency is needed so that students stop inferring intent from the surrounding culture. Procedures matter as much as policies, and their evidentiary basis has collapsed for AI. [[teichmann-detecting-undetectable-misconduct-2026|Teichmann (2026)]] argues that the misconduct procedure universities imported from plagiarism rests on a premise generative AI dismantled: that prohibited use can be detected and proved. Unlike text-matching software, which can point to a copied source, AI-text classifiers identify no source because none exists — they output a probabilistic judgment about style that degrades under paraphrase, misclassifies non-native speakers systematically, and cannot be explained or cross-examined, while skilled or lightly edited use leaves no trace at all. Persisting anyway reverses the burden of proof (the student is asked to prove a negative), strains every element of procedural justice, and lands the harm of false accusation hardest on the already disadvantaged. The proposed remedy is twofold: publish an evidentiary standard under which detector output alone never grounds a finding, with graduated education-first responses, and move institutional effort into validity-centered and authentic assessment design. [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi (2026)]] supply the mechanism-level counterpart from the student's side: because each assessment environment makes some response most attractive, prohibition leaves concealment attractive when verification is thin, monitoring makes hidden use costlier without making disclosure safe, and only redesign — lowering the payoff from outsourcing while raising the value of visible reasoning — moves students toward responsible use. Their sharpest result is that deterrence runs through a detector's discrimination between hidden use and legitimate work rather than its catch rate, so when false positives rise faster than true positives, stronger monitoring can make concealment relatively more attractive. ### The socio-emotional dimension Integrity enforcement has a neglected emotional cost. [[shame-guilt-ai-regulation-computing-education|Shame-and-guilt research]] with students shows these emotions regulate when and how AI use becomes visible, producing **hiding behaviors and selective disclosure** — and that they coexist with continued use, creating cycles of reduced agency and moral tension rather than behavior change. That is why [[social-norms-ai-use|the norms around AI use]] regulate visibility more effectively than they regulate use. Students even describe their AI use in language of addiction. The implication: detection-heavy, surveillance-oriented policy risks **driving [[ai-misuse-learning-harm|misuse]] underground** rather than addressing it, undermining the candid negotiation that productive use requires. The obverse of hiding is declining, and it carries its own emotional cost. Among the 85 student teachers [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al. (2026)]] surveyed under an assessment policy that explicitly permitted generative AI, 62.4% (53) chose not to use it at all, and the adopters' use was shallow and corrective — proofreading (43.8%) and clarity checks (34.4%) far outnumbered text generation (18.8%) and brainstorming (12.5%). Fear of [[legal-issues-and-risks|wrongful accusation]], not technical difficulty, did much of the work: 41.5% of non-adopters named it, against 13.2% citing missing knowledge or skills, while 77.4% framed non-use simply as a preference for working alone. When a permissive policy is read as a trap, the safe response is refusing the offer — which costs the institution the AI-integrated work it was trying to invite. ### AI use disclosure statements [[ai-use-disclosure|AI use and disclosure statements]] are the concrete mechanism through which integrity expectations are operationalized — and the research shows they often fail when treated as neutral compliance forms. [[gonsalves-student-non-compliance-ai-declarations-2025|Gonsalves (2025)]] found 74% of students failed to declare AI use on a mandatory coursework coversheet, driven by fear of penalties, guideline ambiguity, inconsistent enforcement, and peer norms. [[kirsanov-beyond-detection-ai-online-assessments-2026|Kirsanov et al. (2026)]] and [[vetter-hidden-cost-disclosure-genai-2026|Vetter et al. (2026)]] confirm that fear of retribution and unclear policy chill disclosure — and that transparent students can even draw suspicion. [[chang-should-i-tell-my-teacher-ai-disclosure-2026|Chang et al. (2026)]] reframe disclosure as a [[help-seeking]]/[[self-regulated-learning|self-regulation]] behavior that anxiety redirects toward peers. The collective lesson: disclosure policies must address the [[affective-computing|affective]] and social barriers, be clear and consistent, and treat disclosure as [[formative-assessment|formative]] pedagogy rather than surveillance. **[[luo-dawson-value-judgments-grading-2026|Luo & Dawson (2026)]]** add the teacher-side of the equation: teachers' grading of GenAI-assisted work is driven by value judgments about student honesty, diligence, and trust, and many teachers penalize (or are tempted to penalize) students who disclose GenAI use — even when the work quality is strong. This is the "two-way [[explainable-ai|transparency]]" problem: students are expected to declare use, but teachers rarely clarify how that declaration will affect grades, so honest disclosure can carry an unstated grading penalty. The study grounds the disclosure problem in the value-laden reality of teacher grading and argues that transparency must run both directions. Students' own reports quantify how badly that message is landing. In a survey of 504 sociology undergraduates, 81 percent said their instructors or teaching assistants had given guidance on AI use, yet only 46 percent called those instructions very clear; 19 percent reported receiving no guidance at all and the rest described it as at best somewhat clear ([[student-genai-use-views-writing|Kuznetsov, Sheely & Baker, 2026]]). Fear of committing an academic offense was the second most common concern students raised (28 percent) — a cost imposed by ambiguity rather than by enforcement, since only 3 percent reported using GenAI to generate assignment text and 2 percent to produce a full draft. The practical implication is that clarity of communication is itself an integrity mechanism, and that a student who cannot tell what is permitted bears a risk the institution never intended to impose. Course-level permission does not resolve the disclosure dilemma either. In [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al.'s (2026)]] study of 85 student teachers whose assessments explicitly permitted generative AI, self-declarations totaled 28 users against 32 in the anonymous survey, and the two sources disagreed in opposite directions by course — Course A recorded 14 survey users but only 8 declarations, while Course C reversed the pattern (8 survey, 12 declarations). The authors read the divergence as graded consequences shaping what students were willing to make visible, adding student-teacher evidence that a declaration system measures disclosure behavior as much as it measures use. ### Cultural and contextual variation Policy text does not equal policy perception. [[cross-cultural-student-perceptions-genai-computing|Cross-national research]] found that, despite functionally identical institutional policies, students at different universities rated the same AI-assisted practices differently — **culture, not policy wording, drove perceived wrongness**. Policy harmonization does not produce perception harmonization, so culturally diverse cohorts interpret the same rules differently, an equity concern for enforcement and grading that argues for scenario-based clarification over abstract rule statements. Language background rather than culture is a second axis of variation. [[li-genai-assessment-language-equity-2026|Li (2026)]] argues that because the same interface performs permitted language editing and prohibited substantive drafting, integrity rules treating GenAI as one category of unauthorized assistance convert linguistic disadvantage into integrity risk: an outright ban removes a scalable form of language support, a permit-but-disclose regime loads compliance work onto EAL students whose use is more frequent and iterative, and prohibitions on editing beyond minor changes concentrate suspicion on writers whose fluency has shifted most. The boundary Li proposes is purpose-based rather than tool-based — permitted support adds no new ideas, no new sources and no material re-ordering of analysis, while substitution creates or materially reshapes the intellectual work being evaluated — and it is anchored in the assessment construct because rubrics that award marks for fluency and idiomatic expression introduce construct-irrelevant variance for EAL writers ([[assessment-validity]]). The framework demotes [[ai-detection|detector]] output to a triage signal that rarely constitutes proof, preferring triangulation through staged submissions, draft histories, verifiable source trails and a brief construct-aligned conversation, and it calibrates [[ai-use-disclosure|disclosure]] so that routine support does not carry compliance costs exceeding those borne by monolingual peers. ### From policing to pedagogy The knowledge base documents a paradigm shift: from AI as an integrity threat to be policed, to AI as a tool whose appropriate use must be taught. This is the ethical dimension of [[ai-literacy]] and is [[embodied-learning|embodied]] in practical design: - **Task-specific AI-use declarations:** [[genai-declaration-frameworks-higher-education|Domain-specific declaration frameworks]] replace generic "I used AI" checkboxes with structured declarations mapping use to cognitive stages (e.g., structural planning vs. content generation), forcing reflection and shifting focus from policing to professional practice. - **Process-transparent assessment:** architectures such as [[credential-cognitive-stewardship-ai-assessment|cognitive stewardship]], staged submissions, oral defenses, and the [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene|AI Viva]] (a [[conversational-ai|conversational agent]] probing whether students understand their submissions) make [[human-in-the-loop-ai|human judgment]], verification, and responsibility visible. [[miles-prompt-literacy-human-centered-genai-framework-2026|Miles, Haber-Curran and Arar (2026)]] push the same logic down to the prompt itself: their sample rubric grades iterative refinement, critical interpretation of output, and reflective revision, so what is assessed is the student's engagement with the tool rather than the artifact it produced, and they argue that teaching students only to optimize output leaves the ethical and epistemological dimensions of use untouched. - **[[reducing-ai-misuse|Reducing misuse]]:** integrity sits alongside [[ai-misuse-learning-harm]] (the learning cost of misuse) and [[reducing-ai-misuse]] (the interventions that prevent it), tying honesty to genuine learning rather than rule-following. - **Transparency as an integrity strategy, not just a courtesy:** [[mccorkle-aligned-genai-course-policy-2025|McCorkle's (2025)]] design case treats the *rationale* for each allowed or unallowed GenAI use as the integrity mechanism itself. Students interviewed after ignoring a prohibition policy explained that they did not see themselves as behaving dishonestly, which reframes the failure as ambiguity rather than noncompliance — so the redesign answered it by justifying every restriction with the specific [[assessment]] it protects, and by designing for McCabe's "20-60-20" persuadable middle rather than the determined few. The case also names the equity cost of vague policy: expectations that are unclear and uneven across instructors are what convert policy failure into disciplinary action ([[equity-in-ai-education]]). - **Design, not detection: the Bochum case.** The Ruhr University Bochum redesign of introductory nuclear and particle [[physics-education|physics]] ([[ai-particle-physics-education-redesign-2026|Mikhasenko et al., 2026]]) linked its AI policy to a practical integrity failure rather than to detection: because tutorial problems were disclosed in advance, some students prepared AI-generated solutions and copied them onto the blackboard without engaging in the intended reasoning. The authors treat this as a design problem — moving tutorial problems to prepared in-class discussion and making a written exam the grade determinant — rather than a policing problem, while explicitly permitting AI in study, explaining why verification is the student's responsibility, and noting that open-ended AI-permitted homework also raised dependence, uneven access to paid models, and teaching-assistant workload. The clearest statement of the pedagogical inversion comes from [[ai-agents-joyful-assessment-third-space-2026|El Khoury and Ma (2026)]], who argue that reform starting from suspicion narrows the educational imagination to control, compliance and surveillance, and that integrity should be a **consequence** of assessment designed for engagement rather than its starting point. Their supporting observation is a mechanism already visible elsewhere in this knowledge base: disengagement is one of the conditions under which dishonesty becomes more likely, so agendas that stress-test assessments or constrain AI use without addressing engagement leave part of the problem untouched. Students outside Anglophone higher education describe the same tension in their own terms. [[mulisa-students-genai-integrity-perspectives-2026|Mulisa and Mezgebu (2026)]] interviewed 27 undergraduates at an Ethiopian university and found the student body divided against itself: almost all used GenAI or watched peers use it, most credited it with improving their academic achievement, a small minority called such use outright misconduct, and nearly everyone described an uneven playing field in which AI-assisted work earned better grades than honest effort while independent workers lost their sense of diligence. The sharpest feature of the account is the students' definition of plagiarism, which is technically defensible and incomplete — if plagiarism means reproducing someone else's words, a machine-written text that copies nothing is not plagiarism — which is why the authors argue the definition must expand beyond copying and pasting, following Ka and Chan's (2025) "AI-plagiarism", and why they place students' own beliefs, not only institutional rules, at the center of ethical use. [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] takes the pedagogical inversion a step further by extending Eaton's postplagiarism frame from an ethical orientation into assessment design. Integrity, on this account, is enacted through [[evaluative-judgment|evaluative judgment]] — the learner's capacity to weigh options, justify academic choices, and assume responsibility under epistemic uncertainty — and made visible through four practices: annotated decision trails, verification and accountability practices, oral defense and dialogic accountability, and draft differences with version history. Detection is retained only as a supplementary layer for clear misrepresentation or deliberate outsourcing of intellectual labor, never as the primary infrastructure of integrity, because it asks whether GenAI was used rather than how the decisions behind the work were made; the argument also names the equity risk that detection-centered models fall hardest on learners who rely on generative tools for linguistic or cognitive [[scaffolding]]. ### Connections Academic integrity connects to [[assessment-validity]], [[ai-literacy]], [[ai-detection]], [[authentic-assessment]], [[assessment]], [[educational-policy-ai]], [[regulation]], [[ethics]], and [[equity-in-ai-education]]. It is the ethical dimension of [[ai-education|AI in education]], inseparable from [[cognitive-offloading|Over-Reliance]] and the broader question of how [[generative-ai]] reshapes [[higher-ed]] and [[k-12]] learning. - **Systematic-review synthesis.** A PRISMA review of 25 studies (Balalle & Pannilage 2025) finds AI acts as both a threat (AI-generated writing, paraphrasing tools) and a detection tool (Turnitin AI scores), that detection software is unreliable for AI-generated work, and that institutions must build a culture of academic integrity through clear policy, assessment redesign, and ethics training rather than policing alone.([[ssaho-ai-academic-integrity-review-2025]]) - **From policing to dialog: learning verification.** A practitioner account of Grand Canyon University's institution-wide framework ([[best-response-student-ai-dialog-2026|Mandernach 2026]]) argues detection is unreliable and formal integrity processes rarely reach resolution, leaving faculty with "suspicion without recourse." GCU instead adopted **learning verification** — asking students to demonstrate understanding of their submitted work in a brief conversation — reframing integrity from a compliance problem to an assessment problem. It restores faculty authority, shifts students from "how not to get caught" to genuine [[student-engagement|engagement]], and treats AI use as acceptable when the student can demonstrate learning; students' initial anxiety about verification underscores that surveillance-heavy policy can corrode [[trust]]. - **GenAI defeats autogradable homework (2026):** ChatGPT passed every one of 150 test sessions on deliberately hardened, autogradable Qiskit (quantum computing) homework designs — [[personalized-learning|personalization]], hidden references, reflections, [[simulation|simulator]] execution — showing that rubric-based graders cannot reliably distinguish AI-completed from student-completed work and arguing for direct assessment of understanding ([[chatgpt-qiskit-homework-autogradable-2026]]). - **Evaluation in the age of AI — output as evidence (2026):** a university-level analysis argues the AI assessment crisis is a misalignment between assessment design and [[learning-gains|learning outcomes]], not just dishonesty; it documents surveillance harms (lockdown browsers, eye-tracking), a "Disclosure Trap" (students fear declaring AI use lowers marks), a performance gap that grades socioeconomic status (paid vs. free [[llm]] tiers), and "pedagogical burnout" among faculty policing AI — recommending process-based evaluation over detection ([[evaluation-age-ai-output-evidence-2026]]). ### Newer evidence: ethical reasoning, detection limits, AI marketing, and ghost students A wave of recent research sharpens the picture of academic integrity in the age of generative AI: - **Secondary students reason about AI-giarism situationally, not as a fixed rule.** [[chan-rethinking-aigiarism-secondary-integrity-2026|Chan (2026)]] shows that secondary students' ethical reasoning about "AI-giarism" is nuanced and context-dependent — many see AI-assisted work as acceptable when it supports understanding but problematic when it substitutes for their own effort — challenging the assumption that students simply lack integrity or that a single policy can capture their ethics. - **Authentic assessments alone cannot safeguard integrity.** [[kofinas-generative-ai-authentic-assessment-integrity-2025|Kofinas et al. (2025)]] find that markers generally **cannot distinguish** assessments with GenAI input from those without, and that the level of assessment authenticity has **no impact** on the ability to safeguard against or detect GenAI use. The higher-education sector "cannot rely on authentic assessments alone to control the impact of GenAI" — a direct challenge to the [[authentic-assessment|assessment-redesign]] strategy, which must be paired with other measures. - **Integrity guidance must extend into the research process, not just [[teacher-role|teaching]].** [[dai-chan-responsible-genai-research-ai-literacy-2026|Dai & Chan (2026)]] find postgraduate researchers enact [[ai-literacy]] across research tasks and argue that responsible-use policies, which currently focus on teaching and assessment, must [[scaffolding|scaffold]] the ethical dimensions of GenAI use in research — where concerns center on originality, authorship, [[privacy|data privacy]], and skill degradation rather than plagiarism alone. - **Purpose must precede policy.** [[taylor-lacroix-purpose-before-policy-academic-integrity-2026|Taylor & LaCroix (2026)]] argue that whether GenAI use constitutes misconduct depends on the university's *purpose*. Rising misconduct cases reflect structural incoherence in the neo-liberal university, where technological enthusiasm, corporate influence, and policy enforcement conflict — leaving students accountable for behaviors implicitly shaped by the institution. Universities cannot credibly enforce integrity without coherence between stated mission, pedagogy, and technology practice. - **Psychological and behavioral determinants.** [[psychological-mechanisms-academic-integrity-ai-2026|Frontiers research]] maps the psychological mechanisms and behavioral determinants of academic integrity under AI — how attitudes, self-efficacy, norms, and perceived consequences shape honest use — connecting integrity to [[motivation]], [[self-efficacy]], and [[ai-literacy]] as behavioral constructs rather than pure rule-following. [[predictors-ethical-genai-use-higher-ed-2026|Tabares-Cruz et al. (2026)]] quantify these in a SEM of 980 Ecuadorian university students, ordering academic-integrity and transparency dispositions first, followed by AI literacy, [[critical-thinking|critical verification]], institutional guidance, self-regulation, and data protection — with integrity dispositions and [[ai-literacy]] together explaining a substantial share of ethical GenAI use and outpacing institutional rules alone. - **AI marketing normalizes use and framess "cheating vs. competing."** [[sobo-cheating-competing-ai-marketing-literacy-2025|Sobo et al. (2025)]] show AI is marketed to students as a practical necessity ("Make your writing sound more natural to avoid being mistakenly flagged"), and students feel compelled to adopt it to stay competitive even while worrying about dependency and learning forfeiture — an internalized entrepreneurial imperative. This points to the need for [[reducing-ai-misuse|marketing literacy]] as part of AI integrity education. - **AI humanizers expose the performative cycle of detection.** [[roe-ai-humanizers-legitimacy-assessment-2026|Roe et al. (2026)]] catalog 55 AI-humanizer websites that alter AI-generated text to evade detection, framed through Goffman's dramaturgy. Humanizers make misconduct discursively absent and perform legitimacy, demonstrating that the detection-vs-circumvention arms race is structurally unending — reinforcing the shift from [[ai-detection|policing]] to assessment design and [[ai-literacy]]. - **Ghost students and the agentic-AI verification gap.** [[bozkurt-ghost-students-agentic-ai-2026|Bozkurt, Crompton & Fell Kurban (2026)]] introduce the **"ghost student"**: a digital surrogate created by coupling LLMs (the "mind") with agentic AI browsers (the "body") that can navigate LMS, engage content, and complete assessments with human-like mimicry, making the actual learner's presence optional. This creates a **verification gap** that traditional proctoring and detection are structurally unable to close — an integrity threat that grows as AI becomes [[agentic-ai|agentic]] rather than merely generative. The clearest disciplinary case for redesign over detection comes from computing education. A [[meta-analysis-systematic-review|systematic review]] of 72 studies of generative AI in computing and [[cs-education|programming education]] ([[kumar-genai-computing-education-systematic-review-2026|Kumar, Wongsirichot and Nanthaamornphong 2026]]) found only **three** studies examining AI-detection mechanisms — the thinnest evidence base of the 14 themes it consolidated — while course and assessment redesign was supported by 25. The review treats the imbalance as diagnostic rather than incidental: institutions have largely updated policy documents without redesigning assessments, and most instructors sit at a *tolerance* rather than *transformation* level of integration, with 70% of one national faculty sample explicitly requesting training on AI-resistant assessment design. Its recommended response is concrete: add an oral component or other process-visible element to at least one high-stakes assessment per course, and make critical engagement with AI output (reading, testing, modifying, explaining, critiquing) a graded, observable component of student work rather than an aspiration left to student discretion ([[assessment-validity]]). - **Fabricated references have reached the published computing-education record.** [[citation-errors-hallucinations-computing-education-2026|Denny et al. (2026)]] traced 113,588 references from 5,225 computing education papers in the ACM Digital Library against the full corpus of 723,930 publications and 15,872,533 references, manually verified 828 suspicious records, and confirmed 30 references containing verifiably fabricated bibliographic information across 14 papers, all published in 2025 or 2026. At the SIGCSE Technical Symposium the verified count rose from 3 in 2025 to 17 in 2026, sitting in 2.3% of 2026 proceedings papers, and hallucinated references appeared across five SIGCSE-sponsored or in-cooperation venues in 2025. The number is intact only with its counterweight: most flagged references were benign — 229 were ACM metadata mismatches where the PDF was correct and 188 were valid bibliographic variants — so the venue figure is a deliberate lower bound. It lands on authors, not only on [[peer-assessment]], because reviewers checking reference lists cannot verify every entry, and [[llm]]-assisted drafting makes an invented but plausible citation cheap to produce. - **Clarity without redesign moves misuse sideways rather than removing it.** [[petricini-zipf-ai-use-ethics-matrix-2026|Petricini & Zipf (2026)]] plot AI use on two axes — students' intention and effort against the clarity and support the environment provides — and report that the most populated quadrant in their interview data was *anxious compliance*, where students hide legitimate help (grammar support, concept explanations, organizing their own ideas) to avoid [[legal-issues-and-risks|false accusation]]. Their warning is directional: where rules become clear but [[assessment]] still rewards speed and product, policy-aware students shift into *efficient circumvention* rather than into *virtuous tool use*. [[austin-ai-agents-assignment-redesign-2026|Austin (2026)]] reaches the same place from the assignment side — when agents satisfy every rubric criterion without visible reasoning, grading the decision trail (confidence calibration, rejected AI suggestions, course-specific constraints) replaces detection, which she notes misfires in both directions. ## Connected Concepts - [[ai-use-disclosure]] — AI use and disclosure statements - [[assessment-validity]] - [[ai-literacy]] - [[ai-detection]] - [[authentic-assessment]] - [[assessment]] - [[educational-policy-ai]] - [[regulation]] - [[ethics]] - [[equity-in-ai-education]] - [[cognitive-offloading]] - [[ai-misuse-learning-harm]] - [[reducing-ai-misuse]] - [[misconceptions]] - [[generative-ai]] - [[higher-ed]] - [[k-12]] - [[ai-education]] - [[legal-issues-and-risks]] - [[social-norms-ai-use]] — the informal rules that sit under formal policy ## Connected Articles - [[ivory-psychology-assessment-integrity-2026]] — A whole psychology program passable at minimum effort, and the marking criteria that let it through (Ivory et al. 2026) - [[zou-is-this-a-trap-student-teachers-genai-2026]] — Student teachers declined GenAI under a permissive policy; "trap" framing and survey–declaration gap - [[ai-agents-joyful-assessment-third-space-2026]] — AI agents, joyful assessment, and third space - [[kumar-genai-computing-education-systematic-review-2026]] — Detection evidence is thin (3 studies of 72); redesign carries the weight - [[mccorkle-aligned-genai-course-policy-2025]] — Aligned GenAI course policy: assessment-derived permissions, transparent rationale (McCorkle 2025) - [[varia-construct-equivalent-assessment-variant-generation-2026]] — Construct-equivalent assessment variant generation (Lee 2026) - [[chirikov-regulate-ai-syllabi-2026]] — How instructors regulate AI across 31,000 course syllabi; integrity framing declining (Chirikov 2026) - [[biology-degree-integrity-genai-cheating-2026]] — Can students cheat their way to a biology degree? A case study of the vulnerability of biology course grades to academic dishonesty in the era of generative AI - [[gonsalves-student-non-compliance-ai-declarations-2025]] — Student non-compliance with AI use declarations - [[detecting-llm-generated-text-latent-prompt]] — Detecting LLM-Generated Text - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond Detection: Authentic Assessment in an AI-Mediated World - [[ai-assessment-scale-reform]] — The AI Assessment Scale and Assessment Reform - [[authentic-products-authenticated-processes-2026]] — From Authentic Products to Authenticated Processes - [[student-rationalization-ai-writing]] — It's OK Because… Student Rationalization of AI Use - [[ethical-ai-higher-ed-game-theory]] — Coordination Game Framework for Ethical AI Use - [[shame-guilt-ai-regulation-computing-education]] — Shame and Guilt as Social Regulators of AI Use - [[cross-cultural-student-perceptions-genai-computing]] — Did Alice Do Wrong? Cross-Cultural Perceptions of AI Use - [[luo-dawson-value-judgments-grading-2026]] — Value judgments in grading GenAI-assisted work: honesty, trust, validity, and two-way transparency (Luo & Dawson 2026) - [[chen-zou-genai-group-assessment-agency-2026]] — Peer accountability and originality in GenAI-mediated group assessment - [[teichmann-detecting-undetectable-misconduct-2026]] — Detection’s evidentiary collapse and the case for procedural justice, proportionality, and design - [[mohamed-temimi-assessment-imperfect-information-disclosure-2026]] — Deterrence, disclosure, and redesign as an assessment-design problem under imperfect information - [[munoz-misconduct-allegation-evidence-2026]] — What evidence 1,162 GenAI misconduct files actually rested on, and the missing evidentiary threshold - [[wright-transcription-not-generation-2026]] — Over-inclusive AI rules and the transcription-versus-generation distinction - [[leaton-gray-ai-digital-cheating-ethical-pedagogies-2025]] — AI-based digital cheating and prevention-based ethical pedagogies: the 183-to-27 misconduct case and five discipline-specific redesigns (Leaton Gray, Edsall & Parapadakis 2025) - [[coates-governing-academic-integrity-indicators-2025]] — Integrity as a governance problem: a 130-item indicator framework for academic governors (Coates, Croucher & Calderon 2025) - [[ji-student-voices-academic-integrity-scoping-2026]] — Scoping review of 38 studies of higher education student voices on academic integrity and GenAI (Ji 2026) - [[li-genai-assessment-language-equity-2026]] — Drawing the support-versus-substitution line for EAL writers, and the inequity of rules that ignore language background (Li 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Prompt injection against AI-mediated grading: hidden instructions that change the grade undetected (Humble 2026) - [[petricini-zipf-ai-use-ethics-matrix-2026]] — The AI-Use Ethics Matrix: anxious compliance, and why clarity alone can push students into efficient circumvention (Petricini & Zipf 2026) - [[austin-ai-agents-assignment-redesign-2026]] — Grading the reasoning trail when AI agents can complete the assignment (Austin 2026) - [[learning-analytics-genai-secondary-writing-2026]] — Using Learning Analytics to Support Secondary School Students' Writing with Generative AI - [[ai-written-admissions-essays-penalized-2026]] — AI-written admissions essays are widespread but penalized --- ## [Teaching](https://edtechdev.github.io/aied/concepts/teacher-role/) > **Teaching** — how AI reshapes the work, identity, and agency of educators. With 50+ articles examining this dimension, the knowledge base documents a fundamental transformation: from sole knowledge authority to orchestrator of human-AI learning environments. This page goes beyond describing that shift — it details what teachers actually *do* differently, how they can adapt their practice, and how they connect to [[learning-design]], [[ai-literacy]], and [[academic-integrity]]. ## Questions to Consider - The page describes the teacher's shift 'from sole knowledge authority to orchestrator of human-AI learning environments.' What do you gain and what do you lose as an educator when you stop being the main source of content? - It argues AI reallocates a teacher's scarcest resource — attention — from producing materials to interpreting learners. Does that reframing ring true, and what would you actually do with the reclaimed time? - [[research-methods-aied|Research]] cited on the page shows teachers as active AI designers — writing the prompts and scaffolds that shape how AI behaves for their students — rather than passive consumers. What would it take for you (or teachers you know) to feel like a designer of AI rather than a user of it? - At the far end of the 'teacher-AI teaming' spectrum, the teacher orchestrates a team of human learners, AI tutors, and curriculum resources. Where does human judgment remain irreplaceable in that team, and what might quietly erode if the orchestration is left mostly to AI? - If the teacher's attention is reallocated from producing materials to interpreting learners, what new competencies and risks does that introduce — and who supports the teacher in developing them? ## Introduction Teaching here names the professional role as AI reshapes it: less content delivery, more orchestration of human and machine participants, interpretation of what learners actually understand, and design of the conditions under which AI use is legitimate. The pages collected under this concept document teacher–AI teaming and workflow change, the [[teacher-ai-competency|competencies]] the role now demands, and the tension between efficiency gains and the attention that individual support requires. It sits alongside [[educational-development]], [[ai-literacy]] and [[assessment]] as one of the roles that determines whether AI integration changes practice or merely decorates it. The [[learning-sciences|learning sciences]] are the research field whose evidence reshapes this position: they describe and test how learning happens and which designs change it, while this page covers the classroom role that has to act on that evidence in the moment. ## How AI transforms teaching - **From instructor to orchestrator:** [[teacher-ai-teaming-five-levels|Five levels of teacher-AI teaming]] and [[teacher-student-agency-orchestration|agency orchestration research]] map the spectrum from AI as tool to AI as teaching partner. At the far end, the teacher orchestrates a team that includes human learners, [[intelligent-tutoring|AI tutors]], and curriculum resources rather than delivering all content themselves. - **Workflow transformation:** [[ai-changing-teaching-workflows]] documents how AI shifts teacher time from content delivery to higher-value activities like individual support, [[feedback]], and [[curriculum-design]]. The teacher's scarcest resource — attention — is reallocated from producing materials to interpreting learners. A [[li-language-educators-genai-review-2026|systematic review of language educators]] (Li et al. 2026) corroborates this: educators value GenAI most for preparatory work (lesson planning, materials creation, writing support) while hesitating on live classroom use, and their adoption is shaped by professional-identity, pedagogical, technical, institutional, and academic-integrity factors — with competency gaps mapping to episteme (understanding AI's capabilities/limits), techne ([[prompt-engineering]], AI-enhanced task/assessment design, detecting AI-generated text), and phronesis (ethical judgment, bias/privacy handling, context-sensitive judgment). documents how AI shifts teacher time from content delivery to higher-value activities like individual support, [[feedback]], and [[curriculum-design]]. The teacher's scarcest resource — attention — is reallocated from producing materials to interpreting learners. - **Competency demands:** [[teacher-ai-competency|Teacher AI competency frameworks]] define what educators need to know, from basic tool fluency to pedagogically grounded orchestration. [[teacher-ai-adoption-confidence|Adoption studies]] identify the real barriers: confidence, [[governance|institutional]] support, and workload concerns. [[ai-teaching-innovation-ai-tpack-2026|Bai & Hsieh (2026)]] add an empirical weight to this in an SEM of 898 Chinese university teachers: GenAI-supported teaching innovation depends less on isolated AI technological knowledge and more on pedagogically and disciplinarily embedded AI competence ([[ai-literacy]] and professional identity), both of which partially mediated the competence→innovation relationship — a signal to build faculty development around embedded, identity-linked AI competence rather than tool-only training. - **Co-design and agency:** [[teacher-authored-prompts-student-ai-dialogue|Teacher-authored prompts]] and [[gaide-vibe-coding-k12-teachers|vibe coding for teachers]] show educators as active AI designers, not passive consumers — writing the prompts and scaffolds that shape how AI behaves for their students. In children's STEAM arts lesson planning, [[luo-tahir-chatgpt-steam-lesson-planning-2026|Luo and Tahir (2025)]] experimentally confirmed that *how* the teacher delegates to ChatGPT matters: filling content gaps in a self-outlined lesson (the method favored by 60% and recommended) preserved teacher design autonomy while still raising expert-rated plan quality (median 20.5 vs. 17.6, p = .002, large effect), outperforming having AI generate whole plans or merely checking a finished one. AI served chiefly as an inspiration and gap-finding collaborator rather than a source of fully original innovation — reinforcing that the teacher remains the design agent and judge of output, including catching child-safety and cultural-bias flaws the tool missed. A [[meta-analysis-systematic-review|systematic review]] of [[wang-teacher-ai-co-design-review-2026|teacher–AI co-design of learning tasks]] (Wang, Liu & Islam 2026) finds this co-designer stance remains uneven — the dominant mode is still AI as assistant/content generator — and [[talebzadeh-ai-group-activity-roles-2026|Talebzadeh (2026)]] shows the teacher's agency as "[[multilingual-learning|bilingual]] learning designer" is what turns AI-designed group activities into differentiated, ZPD-aligned instruction. - **[[teacher-education|Preservice]] preparation:** [[ai-tpack-preservice-math-teachers|TPACK-based training]] and [[educational-development]] programs prepare future teachers for AI-augmented classrooms before they enter them. [[zhuang-zhang-chatgpt-math-teacher-education-2026|Simulated-student role-play]] extends this into hands-on practice: a custom ChatGPT bot (Student GPT) playing a mathematically [[misconceptions|misconception]]-ridden middle schooler lets preservice teachers rehearse diagnosing and guiding student reasoning in a low-risk setting, preparing them for a central teaching task — reading and remediating student thinking. - **A teacher-facing frontier the evidence base has barely entered:** [[edustories-classroom-case-studies-2026|Štefánik et al. (2026)]] assemble **Edustories**, 1,492 teacher-written case studies of real elementary and high-school classrooms — challenging student behavior, the intervention the teacher attempted, and what followed — precisely because most AI-in-education research has targeted individualized student assistance while most teaching happens in collective classrooms. [[benchmark|Benchmarking]] four language-model families on predicting whether an intervention succeeded, the strongest models reached 58% accuracy against 64% for human experts; the authors read the gap in both directions, as a limit on current [[teacher-ai-competency|teacher-facing]] assistance and as evidence of emerging potential. What the page gains from it is a measured ceiling rather than a capability claim: advice to teachers needs to be better than expert judgment before it can be trusted, and today the ordering is the other way. - **Where AI actually enters the work, by task:** [[where-ai-enters-teacher-work-2026|Holster (2026)]] separates adoption from allocation using TALIS 2024 data from 56,669 teachers, with an allocation sample of 24,058 AI users across 46 education systems. A task rated one point above a teacher's own mean demand was more likely to receive AI help (OR = 1.163), but the direction differs by task: planning rises with demand (OR = 1.084), assessment and marking falls (OR = 0.921), and special-education support and adaptation rises steeply (OR = 1.442, a predicted spread from 34.0% to 50.3% across the demand range). The [[teacher-ai-competency|competency]] and policy reading is that teachers already keep high-stakes marking human while reaching for AI on individualized adaptation work, so guidance that treats teacher AI use as a single behavior will miss where support is actually needed. ### What the changing role actually looks like: concrete examples The abstraction "orchestrator" is easier to grasp through concrete, day-to-day shifts in instructor work: - **From writing every handout to curating AI-generated drafts.** An instructor who once spent an evening building a differentiated worksheet now prompts an AI to produce three versions at different difficulty levels, then spends the same evening *reviewing and adapting* them — a shift from authoring to editorial judgment. [[curriculum-as-code-instructional-design-2026|Curriculum-as-Code]] workflows make this pipeline explicit, and [[kibar-ilgaz-ai-instructional-design-review-2026|Kibar & Ilgaz's systematic review]] shows AI acting as a "co-worker" that drafts content the designer then refines. Expert validation of AI-generated science lesson plans makes explicit what that editorial judgment must weigh: [[karaismailoglu-ai-lesson-plans-science-experts-2026|Karaismailoglu, Surmeli and Yildirim (2026)]] found eleven [[science-education]] specialists rated ChatGPT-4 and an education-focused tool's sixth-grade plans as usable drafts — 7 of 11 judging them "applicable by correction," only 3 directly "Applicable" — and preferences diverged from raw scores (7 preferred the higher-scoring education-focused tool, 4 the general-purpose plan) because judgment weighed [[affective-computing|affective]] and contextual dimensions beyond structural fidelity. Teachers remain the arbiters of whether and how an AI draft becomes a teachable, contextually grounded lesson. - **From lecturing to real-time coaching.** Rather than delivering the same explanation to the whole class, teachers increasingly route routine questions to an AI tutor and focus face-to-face time on the students who need human judgment — [[ai-changing-teaching-workflows|workflow studies]] document exactly this reallocation of attention. - **From grading to designing assessments that AI can't game.** Instructors stop relying on recall-based tasks that [[llm|LLMs]] trivially solve and instead design [[authentic-assessment|authentic assessments]], [[ai-assessment-scale-reform|assessment scales]], and [[formative-assessment|formative]] tasks that preserve [[learning-gains]]. - **From teaching content to teaching *use*.** The teacher's job increasingly includes modeling how to prompt, evaluate, and responsibly use AI — an [[ai-literacy]] curriculum woven into every course, not a separate subject. - **From reading [[visualization|dashboards]] to acting on them — contextually.** [[mejia-domenzain-ml-findings-teachers-blended-2026|Mejia-Domenzain et al. (2026)]] show that how teachers actually use analytics dashboards diverges by context: flipped-classroom (university) teachers followed a sequential exploration and favored course-level adaptation and showing dashboards in class, whereas vocational teachers revisited summary pages and used the tool mainly for individual coaching sessions. The actions teachers proposed were shaped by the content represented and their teaching level rather than the plot type — university teachers favored weekly tests and course adaptation, vocational teachers direct, individualized coaching. This points to context-aware dashboard design and differentiated teacher support needs rather than a one-size-fits-all analytics interface. - **From gatekeeper to principled facilitator of collaboration feedback.** The Community Builder ([[breideband-community-builder-cobi-2026|CoBi]]) classroom pilots show the teacher role decisively shapes *implementation integrity*: teachers who used the AI's noticings and visualizations to spark [[metacognition]] and critical reflection about [[collaborative-learning|collaboration]] achieved high-integrity use, while those who let the system drift into performance-monitoring — or veer off into general AI discussions — did not. Teachers also worried about being put "on the spot" by real-time feedback and preferred pre/post-action review over live display, underscoring that orchestrating classroom-wide collaboration AI demands substantial [[professional-training|professional learning]], not just tool fluency. ### The orchestration metaphor The dominant metaphor in the knowledge base is *orchestration*: teachers coordinate human learners, [[intelligent-tutoring|AI tutors]], and curriculum resources. This contrasts with replacement narratives — AI augments rather than substitutes for human teaching. Empirical evidence supports this stance: a PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 studies (2023–2025) found LLMs match human raters on short, well-structured tasks but that performance declines on longer, multilingual, and nuanced work — concluding LLMs cannot fully replace teachers and that hybrid, human-in-the-loop assessment systems achieve the highest grading effectiveness ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). The orchestration lens reframes the instructor's core skill as *judgment*: deciding when a human, an AI, or a designed learning activity is the right instrument for a given learner and moment. A 2026 PRISMA review of 29 studies of [[teacher-intervention-k12-ai-based-instruction-2026]] sharpens the metaphor into a four-phase cycle — monitoring, judgment, intervention, orchestration — and shows why no phase can be assumed. Teachers preferred shared control, accepting, modifying, rejecting or overriding AI suggestions; some deferred intervention deliberately so students could struggle productively first; and dashboards that expanded awareness could also overload attention or exceed what one teacher could physically act on. Its three strategies — pedagogical translation of AI output, design of learning support, and reconstruction of interaction structures — describe the work as recontextualization rather than approval, which places the teacher closer to mediator than to the human-in-the-loop reviewer role. ### Evolving and critical teacher roles Recent work expands the orchestration metaphor into richer role conceptualizations: - **Co-orchestrator across activity transitions.** [[teacher-student-agency-orchestration|Yang et al. (2026)]] use participatory speed dating with 17 teachers and 13 students to map how control should be distributed across the stages of a classroom activity — the before, during, and after transitions between individual and collaborative work — in a co-orchestration tool supporting real-time dynamic pairing. The design principle that emerges is that control is *not* a fixed point: teachers and students want different amounts of agency at each stage, and the tool should let the teacher adjust grouping, activity transitions, and intervention timing in real time while retaining a teacher override. This positions the teacher as a **co-orchestrator of the whole live learning environment** — a role more specific than the general orchestration metaphor, focused on the granular, moment-to-moment decisions of who works with whom, when, and with what support — and it connects individual-tutoring research to classroom-level design. - **The "cognitive choreographer."** Posthumanist frameworks recast the teacher as a *cognitive choreographer* who orchestrates cognition distributed across [[biology-education|biological]] and artificial systems, moving beyond instrumentalist models like [[tpack]] and SAM.([[elsayed-pedagogical-symbiosis-posthuman-learner]]) - **Facilitator, co-investigator, [[ethics|ethical]] supervisor.** In science learning, teachers' roles shift from knowledge transmitters to facilitators and co-investigators, and gain new responsibilities as ethical supervisors of students' responsible AI use.([[li-ai-science-situated-learning-teachers-2025]]) - **Mediator of learning principles.** Educators operationalize age-old learning principles (experiential, situated, and distributed cognition) through AI, treating AI as a tool that enhances rather than replaces the educator's guiding role.([[fowlin-operationalizing-learning-principles-ai]]) - **Unsettled and vulnerable professional identity.** [[farazouli-navigating-uncertainty-teachers-genai-2026|Farazouli et al. (2026)]] document the *emotional* side of role reconfiguration: 24 Swedish university teachers experienced GAI's emergence as alarming and overwhelming, reporting a "state of vulnerability" (low confidence, insecurity, fear of "not being ahead of students") and feeling "stuck" between utopian and dystopian discourses. GAI prompted them to re-evaluate their priorities — cultivating [[critical-thinking]], [[evaluative-judgment]], and ethical GAI use — and to question their own and the university's future role in a landscape that felt "out of control." This frames the teacher-role shift as identity-level and emotionally charged, not merely a skills or workflow change. - **Academic developers as digital mediators, including as a brake.** [[beyond-the-algorithm-academic-developers-digital-mediators-2026|Sithole (2026)]] interviews twelve academic developers and learning designers at two South African Historically Disadvantaged Institutions and finds the role shifting from facilitating reflective teaching toward technological intermediary and problem-solver, with mediation spanning pedagogy, ethics and institutional politics. The distinctive move is treating slowing down as part of the job: one participant describes translating between "what management wants, what the technology can do, and what lecturers are actually worried about," and naming plagiarism, student dependency and the outsourcing of thinking before enthusiasm sets the pace — "It is not just technical support; it is also negotiation." The same study documents the affective cost ("I feel like an impostor. I am learning AI as I go, but the institution expects expertise") and the professional-learning response: peer spaces that test and deliberately break tools, aiming at judgment rather than skill, and an explicit refusal of the adoption-or-resistance binary — "I am not anti-AI; I am pro-context." - **Principled selectivity instead of adoption or resistance.** [[ai-integrated-teaching-identity-tensions|Adiozaman and Segar (2026)]] interviewed two academics with more than ten years' experience three times across a single semester and found identity work organized around three interrelated tensions — pedagogy versus platform, educator versus facilitator, and care versus compliance — which resolved not into a settled position but into *principled selectivity*: context-sensitive decisions guided by pedagogical values, ethical commitment and professional judgment. They describe a trajectory from implicit orientation, through mid-semester strategies such as asking students to explain their thinking and redesigning tasks and assessment, to stances that stayed deliberately open rather than fixed, with refusal of a particular use treated as a legitimate exercise of judgment rather than a failure to adopt. This supplies the constructive counterpart to the vulnerability documented above: what the role shift demands is judgment, so institutional responses built around tool training address the wrong problem. - **Semi-autonomy as the preferred operating point.** [[dai-genai-frenemy-teaching-autonomy-2026|Dai et al. (2026)]] surveyed 287 GenAI-experienced teachers in 27 countries and regions and found them placing the tool at the *semi-autonomous* levels of an automation scale — 130 chose Level 2 ("teacher assistance"), 94 Level 3 ("partial automation"), and exactly one chose Level 6 — describing it as "auxiliary", "scaffold", "complement" or "partner". [[technology-acceptance-model|Perceived usefulness]] dominated their adoption intention (β = .832, p < .001), while perceived artificial autonomy influenced intention only indirectly through usefulness, and risk aversion depressed intention without denting usefulness because their worries attached to *students'* use — cheating and integrity, weakened foundational knowledge and higher-order thinking, [[hallucination-risk|hallucinations]], reduced human interaction, and ethical, legal and [[equity-in-ai-education]] risks. The authors' "frenemy" label captures the stance precisely: GenAI valued as a support tool but distrusted as an autonomous agent. The corollary for the teaching role is that teachers want to remain the deciding agent and to settle use case by case — discipline, student needs, task type, timing — so orchestration tools should be built around [[human-in-the-loop-ai|human oversight]] and case-level judgment rather than autonomy promises. - **Authority stays with the teacher; capability expands only inside it.** [[reichert-human-centered-llm-chatbot-design-teachers-2026|Reichert, Briceno, Tabarsi & Barnes (2026)]] gave six secondary teachers across grades 6-12 (social studies, [[math-education|mathematics]], animal science, science, and [[cs-education|computer science]] and robotics) the task of paper-prototyping an LLM [[conversational-ai|chatbot]] for their own classroom, and found that every prototype encoded the same stance: the system was a *bounded expert*, specialized but confined to a defined domain and supervised by the teacher. Teachers drew two boundaries. *Authority* boundaries followed from professional and legal responsibility — control over what students learn and how they are kept safe could not be delegated, so oversight (complete conversation logs, real-time alerts for inappropriate queries, and manual override) was treated as a duty of the role rather than a check on the tool. *Expertise* boundaries followed from what the model cannot know: individual students' histories, classroom dynamics, and institutional norms. Delegation was then allocated selectively across instruction — teachers welcomed AI to present content, supply practice problems, scaffold, and give formative [[feedback]], but kept informing students of objectives and [[summative-assessment|summative]] [[assessment]] themselves. This supplies the semi-autonomy preference documented above with a concrete architecture: the teacher's authority is the fixed element and AI capability grows only within it. - **Teachers as value-laden graders.** [[luo-dawson-value-judgments-grading-2026|Luo & Dawson (2026)]] show that grading GenAI-assisted work is not a neutral, criteria-based act but a value judgment shaped by teachers' conjecture about who the student is (honesty, diligence), what they are capable of (independence, GenAI skill, disciplinary mastery), how they relate to others (trust), and whether the decision leads to good outcomes ([[bias-mitigation|fairness]], beneficence). Teachers often lack awareness of the value orientations underlying their grading, and these vary by discipline ([[humanities-education|humanities]] more critical of GenAI use; hard/applied sciences fewer grading challenges) and epistemological stance (absolutist/multiplist/evaluativist). This positions the teacher as an evaluative professional whose judgment — not just tool use — must be supported and made transparent. - **Teacher as the arbiter of pedagogically meaningful explanations.** Because teachers' [[trust|trust in AI]] recommendations rises when explanations are framed in *their* curricular/pedagogical language rather than in raw model internals, effective tools must "speak" the teacher's domain. In a within-subject experiment, [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor et al. (2025)]] found domain-driven explanations of an AI grouping tool were trusted and accepted far more than data-driven feature-importance ones, positioning the teacher's pedagogical vocabulary as the interface through which AI earns trust and use — and their judgment of whether an explanation is pedagogically sound as a deciding factor in adoption. These roles connect teacher work to [[distributed-cognition]], [[situated-learning]], [[embodied-learning]], and [[critical-pedagogy]], and reframe the teacher as a designer and ethical guide of AI-[[sociocultural-learning|mediated learning]] rather than merely a user. - **AI-supported decisions cluster where behavioral data is available.** [[ai-supported-lecturer-decision-making-2026|Köroğlu et al. (2026)]] reviewed 27 empirical studies (2016–2025) and identified eight lecturer decision types, finding AI support concentrated on instructional, [[feedback]] and [[assessment]] decisions while curriculum, learning-environment, emotional, ethical and administrative decisions were rarely supported. [[learning-analytics]] Dashboards were the most common system, processing text and log data into behavioral indicators of performance and engagement; [[multimodal]] and interaction-based data, agentic systems and cognitive, metacognitive, [[motivation|motivational]] and affective outcomes were all comparatively rare in the reviewed corpus. The practical reading for instructors is that the decision space an AI tool opens is bounded by the data it displays, so tool selection is also a choice about which teaching decisions are being resourced. ## How instructors should adapt their teaching practices to AI The knowledge base's evidence converges on a set of concrete adaptations: **Design the AI's pedagogical wrapper, not just the tool.** The same AI yields large [[learning-gains]] or net harm depending on how the activity is designed around it. [[kibar-ilgaz-ai-instructional-design-review-2026|Kibar & Ilgaz]] and the [[learning-gains]] evidence are consistent: the instructor's job is to *design the learning experience*, treating AI as a component within a structured activity rather than the answer engine. Scaffold a student attempt first, then let AI coach — this is the difference between [[formative-assessment|assessment]] and [[cognitive-offloading|performance inflation]]. **Require a [[human-in-the-loop-ai]] checkpoint.** Rather than accepting AI output at face value, teach students to evaluate, correct, and take ownership of AI-assisted work. [[mechanical-compliance-human-flourishing-ai-literacy-2026|Fair-use AI literacy]] frames this as balancing AI's productivity against genuine human flourishing and authorship. Design assignments so AI helps with drafting but the learner remains the agent of evaluation and revision. **Adapt assessment to the age of AI.** Move away from tasks LLMs can complete verbatim toward [[authentic-assessment|authentic]], process-based, and in-person assessments. [[ai-assessment-scale-reform|Assessment scales]] give instructors a rubric for deciding how much AI assistance is legitimate at each stage, and [[zhao-genai-higher-order-thinking-meta-2026|meta-analytic evidence]] shows AI can support higher-order thinking when the assessment design demands it. **Protect the conditions for durable learning.** Because AI answers make learning look easy, instructors must design to counter the [[cognitive-offloading|performance-learning gap]]. [[lodge-loble-cognitive-offloading-2026|Lodge & Loble]] warn that effortless success with AI can mask the absence of learning; instructors should build in [[productive-failure|productive struggle]], require unassisted demonstration of mastery, and treat AI-supported answers as a starting point, not the finished outcome. **Plan for the delivery medium.** Adaptation is not medium-neutral: [[online-teaching-and-learning|online teaching]] with AI multiplies both the opportunities (scalable [[personalized-learning|personalization]], always-on support) and the risks (integrity, offloading). Design the AI wrapper as deliberately in online as in face-to-face contexts.([[lopez-pernas-llm-appropriate-student-support-2026]]) ## Incorporating AI literacy into teaching AI literacy is not a separate module — it is woven into how instructors design every course. Effective approaches include: - **Model AI literacy in practice.** Teachers who themselves [[teacher-ai-competency|use, evaluate, and critique AI]] set the norm for students. [[ai-literacy-continuum-higher-education|AI literacy continua]] describe the progression from tool fluency to critical evaluation that instructors should scaffold. - **Teach evaluation and [[trust-calibration|calibration]], not just use.** Students need to know when AI is trustworthy and when it hallucinates — [[critical-thinking]] about AI outputs is a core literacy. This connects to [[reducing-ai-misuse|responsible use]] and [[framing-ai-use-for-students|how AI use is framed for students]]. - **Co-construct AI-use norms with students.** [[finkelstein-principled-ai-education-2025|Principled approaches to AI education]] and [[teacher-authored-prompts-student-ai-dialogue|teacher-authored prompts]] show that explicit, co-designed expectations work better than punitive rules. ## Ensuring academic integrity in the age of AI Academic integrity with AI is a *design* problem, not a policing problem. Instructors adapt by: - **Reframing integrity around process and authorship.** Instead of detection, emphasize [[academic-integrity]] as transparent, documented use. [[bozkurt-ghost-students-agentic-ai-2026|Agentic AI and ghost-student research]] shows the integrity risk shifts when students deploy [[agentic-ai|AI agents]] on their behalf, so instructors must define what authorship means when the "ghost" is an AI. - **Using disclosure and framing.** [[framing-ai-use-for-students|How AI use is framed for students]] — whether as a crutch or a legitimate tool — shapes whether they disclose it. [[explainable-ai|Transparency]] norms (e.g., AI-use disclosure) reduce the incentive to hide AI use. - **Designing out the incentive to cheat.** Authentic, process-based, and in-person assessments (oral exams, [[eportfolio|portfolios]], observed [[problem-solving]]) make outsourcing less attractive than AI-detection tools do. [[ai-assessment-scale-reform|Assessment scales]] and [[authentic-assessment]] are the constructive alternative to [[ai-detection|detection]] arms races. - **Task-level [[regulation]] is how instructors actually govern AI.** [[chirikov-regulate-ai-syllabi-2026|Chirikov's (2026)]] study of 31,000+ course syllabi shows instructors increasingly acting as *task-level regulators* rather than applying blanket rules: they restrict AI for drafting/revising (79% of courses) and reasoning/problem-solving (65%), permit it for editing/proofreading (83%) and study support (80%), and leave ideation/planning most contested (46% permit / 54% restrict). This differentiation — built on which tasks AI displaces versus augments — is a concrete, instructor-driven alternative to adoption-or-ban policy and gives the teacher-role a central place in [[educational-policy-ai|policy]] formation. ## Connecting teaching to learning design Teaching and [[learning-design]] are two sides of the same coin — the instructor's daily judgments are the live execution of the designed learning experience. AI tightens this connection: - **Teachers as learning designers.** [[learnai-just-in-time-ai-cocreation-university-2026|LearnAI]] and [[teacher-authored-prompts-student-ai-dialogue|teacher-authored prompt design]] show instructors functioning as designers: specifying the activity, the AI's role, and the scaffolds that structure learning. This is [[learning-design]] in action at the point of use. - **Learning-design principles govern AI pedagogy.** The same [[learning-design]] principles — clear objectives, aligned assessment, [[scaffolding]], and [[feedback]] — determine whether AI helps or harms. [[jeon-isd-agent-bench-2026|ISD-Agent-Bench]] empirically validates that grounding AI design in formal instructional-design models beats theory-free [[prompt-engineering|prompting]]. - **Design for the teacher's orchestration.** Effective AI learning environments are designed *with the teacher in mind* — the tools [[prezenski-human-centered-ai-aided-learning|human-centered AI]] provide should reduce teacher workload and augment judgment, not add another opaque black box. When instructors co-design AI learning activities ([[activity-theory-teachers-adoption-ai-sem-2026|activity-theory perspectives]]), adoption and quality both improve. ## Relationship to learner identity Teacher role and [[learner-identity]] are reciprocal faces of the same human process, and AI reshapes both. - **Teacher identity is a professional identity; learner identity is a learning identity.** The teacher-role page documents how AI reshapes the *work, identity, and agency* of educators — their evolving professional self-understanding (see [[laidlaw-genai-identity-crisis-faculty-2026|Laidlaw's framing of GenAI as an identity crisis]]). Learner identity is the parallel construct for students: who they are and are becoming as learners, in [[stem-education|disciplinary]], professional, creative, and academic terms. - **They are causally coupled.** Teachers who experience identity disruption (uncertainty about their professional purpose amid [[generative-ai|GenAI]]) are less able to support their students' identity development — a teacher who doubts their role struggles to validate students' emerging sense of self in the same domain. Conversely, teachers who sustain a confident professional identity are better positioned to scaffold students' belonging and authorship. - **Distinct failure modes.** Teacher identity is threatened by *role obsolescence* and *purpose* (the "what's the point of teaching?" question). Learner identity is threatened by *authorship loss* and *competence* (the "is this really mine / am I good enough?" question). Both are identity-level (not just skills-level) responses to AI. - **Both are professional-development and [[pedagogy|pedagogical]] concerns.** Supporting teacher identity belongs to [[educational-development]] and [[teacher-ai-competency]]; supporting learner identity belongs to [[student-experience]], [[authentic-assessment]], and [[agency]]. A well-designed AI-integrated system attends to both — because the teacher's identity is the condition under which learners' identities form. ## Teacher Co-Design of Early AI Literacy - **Teacher co-design in early AI literacy.** Lee (2026) shows that two pre-K and two kindergarten teachers who co-designed the Play With AI (PL-AI) curriculum experienced substantial growth in confidence and pedagogical agency, with co-design fostering curriculum ownership, reflective practice, and meaningful adaptation. This positions teachers as central co-designers — not just implementers — of developmentally appropriate AI literacy curricula, a model with implications for [[teacher-education|teacher preparation]] in [[early-childhood-elementary-ai-education|early childhood]] [[ai-education|AI education]]. ### Teachers Co-Designing AI Learning Resources - The teacher's role extends to co-designing AI learning resources: teacher-AI co-designed [[simulation|simulations]] for drone STEM instruction kept GenAI output pedagogically valid and contextually relevant, and teachers are the intended beneficiaries of the Teachers' AI Literacy Scale. Effective GenAI integration increasingly depends on teacher involvement in design. ## Connected Concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[learner-identity]] — evolving disciplinary, professional, creative, and academic learner identities - [[business-education]] - [[educational-development]] - [[teacher-ai-competency]] - [[ai-literacy]] - [[k-12]] - [[higher-ed]] - [[scaffolding]] - [[learning-design]] - [[intelligent-tutoring]] - [[llm]] - [[professional-training]] - [[distributed-cognition]] - [[situated-learning]] - [[critical-pedagogy]] - [[teacher-education]] - [[learning-sciences]] - [[academic-integrity]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) ## Connected Articles - [[edustories-classroom-case-studies-2026]] — Edustories: A Collection of Real-world Case Studies from Classroom Practices - [[reichert-human-centered-llm-chatbot-design-teachers-2026]] — Teachers design classroom chatbots as bounded experts and retain authority - [[dai-genai-frenemy-teaching-autonomy-2026]] — GenAI as "frenemy": teachers prefer semi-autonomy, usefulness drives adoption, risk aversion targets students' use (Dai et al. 2026) - [[ai-supported-lecturer-decision-making-2026]] — AI-Supported Lecturer Decision-Making in Higher Education - [[chirikov-regulate-ai-syllabi-2026]] — How instructors regulate AI across 31,000 course syllabi; task-level regulation (Chirikov 2026) - [[ai-teaching-innovation-ai-tpack-2026]] — AI teaching innovation behavior among college teachers (AI-TPACK SEM; Bai & Hsieh 2026) - [[teacher-ai-teaming-five-levels]] - [[teacher-student-agency-orchestration]] - [[ai-changing-teaching-workflows]] - [[teacher-authored-prompts-student-ai-dialogue]] - [[laidlaw-genai-identity-crisis-faculty-2026]] — GenAI as identity crisis, not skills gap - [[lodge-loble-cognitive-offloading-2026]] — AI, cognitive offloading and implications for education (Lodge & Loble 2026) - [[kibar-ilgaz-ai-instructional-design-review-2026]] — AI and Instructional Design Practice: A Systematic Review (Kibar & Ilgaz 2026) - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[bozkurt-ghost-students-agentic-ai-2026]] — Ghost students and agentic AI in assessment - [[wang-teacher-ai-co-design-review-2026]] — Teacher–AI co-design of learning tasks: trends and perspectives (Wang et al. 2026) - [[farazouli-navigating-uncertainty-teachers-genai-2026]] — University teachers' experiences and perceptions of GAI: vulnerability, rethinking assessment, student learning at risk (Farazouli et al. 2026) - [[luo-dawson-value-judgments-grading-2026]] — Value judgments in grading GenAI-assisted work: honesty, trust, validity, and two-way transparency (Luo & Dawson 2026) - [[mejia-domenzain-ml-findings-teachers-blended-2026]] — Making ML findings accessible to teachers in blended classrooms - [[li-language-educators-genai-review-2026]] — Language educators' practices and development with GenAI - [[luo-tahir-chatgpt-steam-lesson-planning-2026]] - [[breideband-community-builder-cobi-2026]] - [[zhuang-zhang-chatgpt-math-teacher-education-2026]] - [[xai-teachers-trust-edtech-recommendations-2026]] - [[karaismailoglu-ai-lesson-plans-science-experts-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] - [[ai-integrated-teaching-identity-tensions]] — Identity tensions and principled selectivity in AI-integrated teaching (Adiozaman & Segar 2026) - [[teacher-intervention-k12-ai-based-instruction-2026]] — Teacher intervention in K-12 AI-based instruction: a systematic review - [[beyond-the-algorithm-academic-developers-digital-mediators-2026]] — Academic developers as digital mediators: role reconfiguration, affective labor and judgment in South African HDIs - [[where-ai-enters-teacher-work-2026]] — Where Artificial Intelligence Enters Teacher Work --- ## [Teacher AI Competency](https://edtechdev.github.io/aied/concepts/teacher-ai-competency/) > **Teacher AI competency** — the knowledge, skills, and dispositions teachers need to effectively, ethically, and equitably integrate AI into teaching and learning. It extends beyond technical tool use to include [[pedagogy|pedagogical]] integration, [[assessment|assessment literacy]], ethical judgment, and the confidence to [[ai-literacy|use AI well]]. Teacher AI competency is the teacher-side counterpart to [[ai-literacy]], and is developed through [[educational-development|professional development]]. It is central to how [[teacher-role|the teacher's role]] is transforming in AI-augmented classrooms. ## Questions to Consider - The page argues the teacher is 'the decisive factor' in whether AI improves learning — that tools only help when teachers can plan for them, scaffold use, and evaluate outputs. Does that match your experience, or do you think the tool itself matters more than the teacher? - Teacher AI competency spans technical proficiency, pedagogical integration, assessment literacy, and ethical judgment. Which of these do you think teachers most lack, and which is hardest to train? - The [[research-methods-aied|research]] documents a gap between teachers' self-perceived and actual AI skill. Why do you think people overestimate their readiness, and what would it take to close that gap honestly? - An intensive professional-development program produced large gains in AI pedagogical skill in the research cited, suggesting technical-pedagogical skill is trainable. If that's true, why do so many teachers still seem unprepared — what's standing in the way? - If a teacher can 'use AI well,' what does 'well' mean to you — and how would you know a teacher has achieved it rather than just adopted the tool? ## Introduction Teacher AI competency matters because the teacher is the decisive factor in whether AI improves learning. Research consistently shows that AI tools only translate into better outcomes when teachers can plan for them, scaffold student use, evaluate outputs, and integrate them into coherent instruction. The knowledge base's literature examines the *dimensions* of this competency, the *gaps* between self-perception and actual skill, and the *professional development* that builds it. ## Core competency dimensions The knowledge base's research converges on several interconnected dimensions: - **Technical proficiency:**- **Pedagogical knowledge is the decisive layer.** A cross-level study of [[k-12|secondary]] [[ai-education|AI education]] ([[pedagogy-first-technology-second-teacher-knowledge-2026|46 teachers, 2,832 students]]) found technical AI knowledge alone was insufficient — and could even slightly reduce students' perceptions of AI for social good — whereas pedagogical AI knowledge drove students' perceptions and their intention to learn AI. Competency frameworks should therefore weight pedagogical AI knowledge as the pivotal dimension, not a soft add-on. crafting effective [[prompt-engineering|prompts]] for educational objectives, [[ai-ed-evaluation|evaluating AI]] tools for pedagogical fit and safety, and troubleshooting failures in real time. [[genai-pd-ai-pck-learning-gain-2026|An intensive GenAI PD program]] documented significant gains across all five AI-PCK components (overall *d* = 2.36), showing technical-pedagogical skill is trainable. - **Pedagogical integration:** mapping AI use to learning objectives, designing [[scaffolding]] that supports student [[metacognition]] and self-[[regulation]], and integrating AI into [[learning-design|instructional design]]. [[ai-tpack-teacher-multi-agent-workflow|AI-TPACK research]] models how teachers combine technological, pedagogical, and content knowledge through multi-[[agentic-ai|agent]] workflows, while [[teacher-ai-teaming-five-levels|a five-level teacher-AI teaming framework]] (transactional → synergistic) captures how [[generative-ai|GenAI]] may replace, complement, or augment teacher competence. - **Assessment literacy:** evaluating AI-generated content and student AI outputs, and understanding how [[assessment-validity|validity]] shifts when students use AI. This connects to [[automated-assessment]], [[ai-detection]], and the broader [[assessment]] redesign agenda. For recommendation systems, this extends to judging whether an AI's explanations are genuinely understandable and pedagogically meaningful: [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor et al. (2025)]] found teachers trusted AI grouping recommendations more when explanations were framed in domain (curricular) language than when they exposed raw model features — a skill for demanding and appraising [[explainable-ai|explanation quality]] rather than accepting opaque output. - **Model selection and prompt design as demonstrable skills:** instrumental [[ai-literacy]] includes knowing which model to use and how to phrase the task. [[teacher-ai-literacy-prompt-feedback-quality-2026|Jacobsen et al. (2026)]] had 153 pre-service teachers' lesson-planning goals critiqued by ChatGPT-4, Claude 3 and Gemini Advanced under four systematically varied prompts (N = 240 feedbacks in Study 1, 345 in Study 2), with [[feedback]] rated on nine quality categories. Model choice alone explained 26.9% of the variance in rated quality (18.4% in Study 2) and prompt design added a significant 15.9% (5.7%) on top; the single decisive prompt feature was domain-specific technical language, whose removal significantly lowered quality (β = −0.412), whereas adding examples and dropping the chain-of-thought instruction made no significant difference in Study 1. Their reading — that subject terminology both unlocks relevant training-data content and frames the request professionally, like role prompting — plus a rule of thumb to use the most capable current frontier model, makes model choice a didactic decision rather than a technical one. - **Ethical and critical use:** recognizing [[bias-mitigation|bias]] in AI outputs, protecting student data ([[privacy]]), and ensuring equitable outcomes ([[equity-in-ai-education]]). [[llm-cultural-relevance-k12|Culturally relevant AI use]] examines how teachers can use LLMs to diversify materials rather than reinforce dominant norms. - **Confidence and attitudes:** teacher [[self-efficacy|confidence]] shapes adoption. [[teacher-ai-adoption-confidence|Adoption research]] finds confidence, support, and perceived utility drive whether teachers actually use AI, and [[ai-pedagogical-orientation|faculty orientations]] shape adoption in research and teaching. **Emotional and moral readiness is a distinct dimension.** [[vassallo-ai-guilt-complex-faculty-2026|Vassallo (2026)]] surveyed the academic staff of a Maltese [[higher-ed|university]] (109 respondents) and built an AI Guilt Index (α = 0.88) from four moral-emotion items, finding that *anticipatory* guilt outweighed remorse experienced after use: the strongest endorsement was worry that AI use undermines one's credibility (34.9% agreeing), then feeling like one is [[academic-integrity|cheating]] when using it (25.7%), while post-use remorse drew only 9.2%. The findings that matter for competency frameworks are that non-users reported *higher* guilt than users (M = 3.25 vs M = 2.32) and that guilt fell as career security rose — early-career academics reported the most (M = 2.71) and senior academics the least (M = 2.03). Emotional readiness is therefore not captured by skill or confidence measures, and the paper argues competency frameworks should treat guilt and identity concern as normal transitional responses rather than faults to correct. ## The competency gap A key finding is the **gap between [[self-report-measures|self-reported]] and performance-based competency**. [[ai-literacy-assessment-misalignment|Research on AI-literacy assessment]] documents a substantial discrepancy (up to ~40%) between what teachers *believe* they can do and what they can actually *demonstrate* — teachers confident in AI skills often lack foundational prompting and evaluation abilities. This motivates **[[assessment|performance-based assessment]]** of teacher competency rather than reliance on self-report, and connects to [[self-assessment|calibrated self-assessment]]. The gap is visible in the artifact as well as in the self-report. A [[ai-integration-instructional-design-collaboratory-2026|cross-institutional faculty collaboratory in teacher preparation]], in which teacher educators designed AI into their own methods courses, reported that candidates could produce polished AI-assisted lesson plans while being unable to explain why a plan fit the learners and the standards, since the plan itself says nothing about the reasoning behind it. Grading the justification rather than the product is one response. The [[bondurant-shaughnessy-ai-pedagogies-practice-2026|pedagogies-of-practice frame]] suggests another, treating rehearsal as an approximation of practice: AI-mediated rehearsal with structured post-rehearsal feedback raised candidates' use of probing and exploring questions, yet candidates' own judgments of their performance still diverged from what observers recorded. **The gap also shows up as non-participation.** [[watson-rainie-ai-challenge-faculty-survey-2026|Watson & Rainie (2026)]] surveyed 1,057 US college faculty in late 2025 and found 26% do not use [[generative-ai|generative AI]] tools at all, with a third choosing not to use them for teaching and non-use concentrated in the arts and [[humanities-education|humanities]] (40%). The institutional side of the gap was larger than the individual one: 68% said their schools had not prepared faculty to use generative AI for teaching and mentoring, and faculty named colleagues' resistance (82%) and unfamiliarity (83%) as the leading obstacles to departmental adoption — a picture in which capability-building, peer norms and policy all have to move together. **Pre-service training largely sidesteps the ethical dimension of the competency.** A [[meta-analysis-systematic-review|systematic review]] of AI in initial teacher training for pre-service primary mathematics teachers ([[pinto-ai-initial-teacher-training-mathematics-review-2026|Pinto et al. 2026]], 11 studies selected from 341 records) found the interventions concentrated on short-term technical and pedagogical gains, with nine of the eleven running brief interventions spanning one to six sessions, and reported that ethics was addressed in only three of the eleven studies. The competency frameworks name ethical judgment as a dimension; the pre-service literature meant to build it rarely teaches or measures it. ## Professional development that works The knowledge base's PD literature identifies effective approaches: - **Intensive, theory-grounded programs:** [[genai-pd-ai-pck-learning-gain-2026|An intensive GenAI PD program]] with 163 teachers/pre-service teachers produced significant gains across all AI-PCK components, with pre-service teachers benefiting most. [[teacher-education-ai-literacy-sdt-2026|Self-determination-theory-based PD]] shows need-supportive training improves teachers' AI literacy, attitudes, and [[student-engagement|engagement]] while reducing anxiety. - **[[design-based-research|Design-based]] and integrated approaches:** [[genai-literacy-training-teacher-education-dbr-2026|DBR-based GenAI literacy training]] addresses the overemphasis on technical knowledge and pre-GenAI tools; [[rail-ed-genai-literacy-teacher-education|integrative, developmental frameworks]] and [[sec-ai-literacy-narrative-review-2026|social-emotional competency integration]] broaden literacy beyond pure technique. - **Inquiry and authentic practice:** [[quest-ai-inquiry-preservice-teachers|AI-supported inquiry models]] build AI literacy and authentic performance in pre-service teachers. - **Context-specific readiness:** [[sangwa-epiq-ai-faculty-readiness-2026|The EPIQ-AI readiness framework]] emphasizes that faculty readiness is a sociotechnical issue requiring alignment of faculty capacity, [[governance]], and quality assurance. - **Support must be differentiated by experience and AI proficiency.** [[choi-teacher-ai-interaction-lesson-design-2026|Choi et al. (2026)]] found that how teachers actually interact with AI in lesson design depends on the *interplay* of teaching experience and AI proficiency, not either alone. Experienced teachers with high AI proficiency critically adapt AI output to context (re-prompting, elaboration), whereas novices — even technically fluent ones — tend to accept AI responses directly and rarely consider students and context. This argues for profiling-based PD: response-evaluation checklists and prompt templates for novices, and hands-on skill-building for experienced teachers with lower AI proficiency. - **Co-design and pedagogical prompt literacy are competencies, not add-ons.** A [[meta-analysis-systematic-review|systematic review]] of teacher–AI co-design of learning tasks ([[wang-teacher-ai-co-design-review-2026|Wang, Liu & Islam 2026]], 28 studies) finds the dominant collaboration mode is AI as assistant/content generator, and locates a gap in teachers' fuller co-design and dialogic partnership capacities. [[talebzadeh-ai-group-activity-roles-2026|Talebzadeh (2026)]] shows PD that pairs technical AI training with pedagogical reasoning — building "pedagogical prompt literacy" (encoding [[tpack|PCK]] into prompts) — is what lets teachers turn AI output into effective [[collaborative-learning|differentiated group activities]]. - **Structures that build competency, rather than lists of it.** [[physics-faculty-learning-community-ai-2026|Perl-Nussbaum and Finkelstein (2026)]] ran a six-session biweekly faculty learning community in a large public R1 physics department — nineteen faculty across the series, about ten at each meeting — where every session opened on local data, moved to small-group testing of AI against real anonymized student homework, and closed in collective discussion, and it produced a five-entry shared repository rather than a training package. Assessment literacy can be built the same way through a design tool: in [[authentic-assessments-generative-ai-pilot-2026|Paula et al.'s (2026)]] pilot, eight experienced STEM and Health course coordinators grew their assessment literacy by critiquing a custom GPT's assessment drafts against disciplinary standards, even though the outputs repeatedly missed disciplinary context, ignored topic sequencing, and in one case carried fabricated references through repeated prompting. All eight kept [[evaluative-judgment|academic judgment]] with themselves and refused end-to-end automation, which makes structured critique of generated drafts the development mechanism rather than tool training. - **Short sessions can move acceptance without moving adoption.** [[mesenhoeller-teachers-ai-differentiation-acceptance-2026|Mesenhöller and Böhme (2026)]] evaluated a three-hour INSIGHT session with 100 German primary and secondary teachers and found perceived usefulness and perceived ease of use both rose significantly (usefulness t(99) = -3.24, p = .002, d = .32; ease of use d = .25) while behavioral intention to use AI-based technologies for differentiation did not change. The authors note their baselines were already favorable, so a ceiling effect is plausible. The reading for PD design is narrow but useful: a short session can shift how teachers judge AI tools without shifting whether they plan to use them, which is the kind of change that needs follow-through rather than a one-off event. - **Institutional support:** [[educational-development|professional development]] must be paired with institutional infrastructure ([[educational-policy-ai|policy]], , [[institutional-change-framework-ai|institutional change]]) for sustainable adoption. A design-based counterexample treats competency as situational rather than a rung on a ladder. [[adaptive-ai-model-teacher-educators-2025|Eyal's 2025 design-based study with 22 higher-education teacher educators]] had participants examine five published assessment frameworks and co-design an alternative organized around three inter-related axes: context fit (infrastructure, socio-cultural factors, local needs, developmental stage), professional needs (discipline, pedagogy, leadership, support), and dynamic development. The model rejects fixed competency levels and allows non-linear progression, and it ships with a 20-item reflective self-assessment questionnaire rated 1 to 5. Its validation is qualitative only, with no quantitative reliability testing, so it stands as a design contribution rather than a validated instrument. ## Teacher AI competency and the transforming teacher role As AI takes over routine instructional and assessment tasks, the teacher's distinctive contribution shifts toward orchestration, judgment, and relationship: deciding when and how AI is used, scaffolding [[agency|student agency]] and critical use, ensuring equity, and providing the social and emotional support AI cannot. This reframes teacher competency around [[human-in-the-loop-ai|human-in-the-loop]] oversight, [[ethics|ethical judgment]], and [[self-regulated-learning|supporting self-regulated learning]] — connecting to [[teacher-role]] and [[cognitive-offloading|guarding against over-reliance]]. ## Implications for AI in education - **Assess performance, not just self-report:** teacher competency should be evaluated through demonstration, given the documented self-report gap. - **Train the full competency, not just tools:** PD should build technical, pedagogical, assessment, and ethical dimensions together, grounded in [[learning-theories|learning theory]]. - **Build confidence alongside skill:** attitudes and [[self-efficacy]] shape adoption, so PD should reduce anxiety and build confidence through authentic, supported practice. - **Support the institutional layer:** sustainable teacher competency requires aligned policy, governance, and capacity, not isolated training. - **Teacher digital competence for GenAI [[curriculum-design|curriculum design]].** [[guillen-curriculum-genai-teacher-competence-2026|Guillén-Gámez (2026)]] validate a TAM-based diagnostic instrument with 434 in-service teachers; behavioral intention was the main predictor of digital competence for using GenAI in curriculum planning, with self-efficacy as a root driver. ### A Psychometric Instrument for Teacher AI Competency - A psychometric study developed the Teachers' AI Literacy Scale (TAILS) to measure AI literacy specifically within [[teacher-education|language teacher education]], operationalizing the ED-AI framework's six dimensions. The instrument's development fills a gap in assessments that target students or general users, supporting the measurement of teacher AI competency. The instrument landscape itself has since been reviewed. [[assessing-teachers-ai-literacy-measurement-tools-2026|Zainal, Mohd Matore and Maat's 2026 systematic review of teacher AI literacy measurement tools]] appraised 33 instruments published between 2019 and 2025 and found the field methodologically monotonous: 31 (93.9%) are self-report scales of perceived confidence, only two (6.1%) test knowledge objectively, and none use performance-based tasks. Internal consistency was the strongest quality domain (28 of 33 at Grade A) and fairness the weakest, with five instruments (15.2%) reporting measurement invariance or differential item functioning evidence. Content also lags the technology, since 29 instruments (87.9%) target general AI concepts and only four (12.1%), all from 2025, address generative AI. ## Connected Concepts - [[ai-literacy]] - [[educational-development]] - [[teacher-role]] - [[prompt-engineering]] - [[learning-design]] - [[scaffolding]] - [[metacognition]] - [[assessment-validity]] - [[equity-in-ai-education]] - [[ethics]] - [[self-efficacy]] - [[human-in-the-loop-ai]] - [[agency]] - [[cognitive-offloading]] - [[educational-policy-ai]] - [[ai-education]] - [[tpack]] - [[teacher-education]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education ## Connected Articles - [[typology-generative-ai-tools-education-2026]] — Tool selection as an exercise of educator agency - [[generative-ai-k12-teaching-learning-systematic-review-2026]] — Systematic review of generative AI in K-12 teaching and learning (Marzano 2026) - [[pedagogy-first-technology-second-teacher-knowledge-2026]] — Teacher professional knowledge in K-12 AI education: TAIK vs TPAIK and student learning (Shen et al. 2026) - [[choi-teacher-ai-interaction-lesson-design-2026]] — Teacher-AI interaction patterns in lesson design across experience and AI proficiency (Choi et al. 2026) - [[tpack-genai-inservice-teachers-mediation-2026]] — In-service teachers' TPACK-GenAI and the mediating role of pedagogical knowledge (Mohebi & ElSayary 2026) - [[preservice-teachers-responsible-genai-2026]] — Pre-service teachers' responsible GenAI use: curriculum implications (Kohnke et al. 2026) - [[governing-unseen-ai-literacy-language-teachers-2026]] — Governing the unseen: AI literacy among language teachers - [[cdpk-pedagogy-benchmark-llms]] — Benchmarking LLM pedagogical knowledge (CDPK + SEND) - [[melo-llm-classroom-observation-teach-2026]] — LLM classroom observation for teacher professional development (Melo et al. 2026) - [[bondurant-shaughnessy-ai-pedagogies-practice-2026]] — Generative AI across representations, decompositions and approximations: rehearsal, structured feedback and the accuracy cautions - [[edurev-100741-tpack-genai-review]] — Systematic review of GenAI in student learning from a TPACK perspective - [[genai-pd-ai-pck-learning-gain-2026]] — Efficacy of an intensive GenAI professional development program - [[ai-tpack-teacher-multi-agent-workflow]] — Modeling AI-TPACK through teacher multi-agent workflows - [[teacher-ai-teaming-five-levels]] — Toward synergistic teacher-AI interactions - [[teacher-education-ai-literacy-sdt-2026]] — Teacher education for AI literacy through self-determination theory - [[genai-literacy-training-teacher-education-dbr-2026]] — Design-based research GenAI literacy training - [[rail-ed-genai-literacy-teacher-education]] — Rethinking GenAI literacy in teacher education - [[sec-ai-literacy-narrative-review-2026]] — Integrating social-emotional competencies with AI literacy - [[teacher-ai-adoption-confidence]] — AI adoption among teachers: confidence and support - [[ai-pedagogical-orientation]] — Faculty orientations shape AI adoption - [[ai-literacy-assessment-misalignment]] — Misalignment between self-reported and performance-based AI competency - [[quest-ai-inquiry-preservice-teachers]] — AI-supported inquiry for pre-service teachers - [[sangwa-epiq-ai-faculty-readiness-2026]] — EPIQ-AI faculty readiness framework - [[llm-cultural-relevance-k12]] — LLMs for culturally relevant K-12 pedagogy - [[institutional-change-framework-ai]] — Institutional change framework for AI - [[teachingcoach-chatbot-instructor-guidance]] — TeachingCoach chatbot for instructor guidance - [[laidlaw-genai-identity-crisis-faculty-2026]] — GenAI as identity crisis, not skills gap - [[raffaghelli-situated-ai-ethics-2026]] - [[guillen-curriculum-genai-teacher-competence-2026]] — Assessing Teacher Digital Competence for GenAI Curriculum Design (Guillén-Gámez 2026) - [[reflective-triangle-model-teacher-ai-2026]] — Reflective Triangle Model: AI as cognitive mediator - [[stenalt-good-education-teacher-ai-conceptions-2026]] — phenomenographic study of university teachers' conceptions of AI - [[preservice-teacher-agency-genai-design-learning-2026]] — Pre-service teacher agency during GenAI interactions in design for learning (Krushinskaia, Elen & Raes 2026) - [[questionnaire-teachers-genai-uses-validation-2026]] — Questionnaire on teachers' uses of generative AI (Pérez-Montesdeoca et al. 2026) - [[ukraine-ai-literacy-secondary-framework-2026]] — Five-level AI literacy framework + PD for Ukrainian secondary educators (Marienko et al. 2026) - [[language-teachers-ai-literacy-edai-2026]] — Teachers' AI Literacy Scale (TAILS) psychometric study (ED-AI framework) - [[wang-teacher-ai-co-design-review-2026]] — Teacher–AI co-design of learning tasks: trends and perspectives (Wang et al. 2026) - [[talebzadeh-ai-group-activity-roles-2026]] — Architecture of roles in AI-designed differentiated group activities (Talebzadeh 2026) - [[xai-teachers-trust-edtech-recommendations-2026]] - [[teacher-ai-literacy-prompt-feedback-quality-2026]] — Prompt engineering and model selection as predictors of AI-feedback quality (Jacobsen et al. 2026) - [[vassallo-ai-guilt-complex-faculty-2026]] — The AI Guilt Complex: anticipatory guilt and four moral response profiles among academic staff (Vassallo 2026) - [[mesenhoeller-teachers-ai-differentiation-acceptance-2026]] — A three-hour teacher PD session raised perceived usefulness and ease of use but not stated intention to adopt (Mesenhöller & Böhme 2026) - [[pinto-ai-initial-teacher-training-mathematics-review-2026]] — Systematic review of AI in pre-service primary mathematics teacher training: ethics addressed in only three of eleven studies (Pinto et al. 2026) - [[watson-rainie-ai-challenge-faculty-survey-2026]] — AAC&U/Elon survey of 1,057 US faculty: preparedness, non-use and the individual-vs-institutional policy gap (Watson & Rainie 2026) - [[chick-faculty-development-ethical-ai-2026]] — Six-week faculty institute from fear to ethical integration, symbiotic pedagogy and AIPACK (Chick, Morello & Staffey 2026) - [[ai-integration-instructional-design-collaboratory-2026]] — Cross-institutional faculty collaboratory: AI integration as instructional design in teacher preparation - [[assessing-teachers-ai-literacy-measurement-tools-2026]] — Systematic review of 33 instruments for measuring teacher AI literacy: 31 self-report, two objective knowledge tests, no performance tasks (Zainal, Mohd Matore & Maat 2026) - [[adaptive-ai-model-teacher-educators-2025]] — A design-based adaptive AI literacy model and 20-item reflective questionnaire co-designed with 22 teacher educators (Eyal 2025) - [[physics-faculty-learning-community-ai-2026]] — A Workshop Series for Effective Use of AI in Uncertain Times: Building a Physics Faculty Learning Community - [[authentic-assessments-generative-ai-pilot-2026]] — Designing Authentic Assessments with Generative AI: A Pilot Study of Assessment Authentifire in Higher Education --- ## [Learning Design](https://edtechdev.github.io/aied/concepts/learning-design/) > **Learning Design** (also known as *instructional design*) — the systematic process of creating effective learning experiences through the analysis of learning needs and the design, development, implementation, and evaluation of instructional materials and activities. AI is transforming learning design by automating content creation, enabling [[adaptive-learning|adaptive learning]] paths, supporting data-driven iteration, and augmenting — rather than replacing — the instructional designer's role. ## Questions to Consider - Think of a course or lesson you have experienced or designed. Where did 'what to teach' (curriculum) end and 'how to teach it' (learning design) begin — and how did the two interact? - A common assumption is that better AI fluency automatically produces better educational content. The page counters this with evidence that explicit pedagogical structure — not just AI fluency — is what determines learning effectiveness. Where have you seen impressive output that failed to teach? - If an AI tool can generate a full course from a prompt, what human decisions become more important rather than less? The page argues AI augments rather than replaces the instructional designer's role — what would that augmented role look like? - Some instructional-design models like ADDIE are used as rigid, linear steps. But the page treats them as iterative, flexible planning heuristics. When might following a process too literally undermine good design? - The page shows that pedagogically grounded prompting — for example, a five-step framework based on learning theory — significantly improved higher-order outcomes. If you were building an AI tutor, what would you encode in an explicit design layer so its teaching strategy stays traceable and reproducible? ## Introduction Learning design bridges AI capabilities and effective pedagogy. Where [[curriculum-design]] addresses *what* to teach at the program level, learning design addresses *how* to teach it at the course and lesson level. The articles in this knowledge base explore both AI as a tool for learning designers and learning-design principles for building effective [[intelligent-tutoring|AI tutoring]] systems. What that design work involves in practice is itself an empirical question. [[tang-chatbots-learning-design-2026|Tang et al. (2026)]] coded 1,378 designer-chatbot turns from five novice learning designers working with a chatbot embedded in a design tool, and found the dialogue clustered on intended learning outcomes and pedagogical approach rather than content generation. Designers returned to outcomes repeatedly as an alignment check while turning curriculum components into concrete tasks, and the assistant's role shifted across phases, from clarifying terms to supporting task design to running a verification pass before a deadline. Design support, on this evidence, is less about producing material than about keeping design intent coherent. ### Key research themes **AI-assisted content creation** is the most directly transformative application. **[[curriculum-as-code-instructional-design-2026|Curriculum as Code]]** presents a six-phase architecture integrating Generative AI with LaTeX and Python to automate [[stem-education|STEM]] materials creation, validated across 8 modules and 28 project contexts with student quality ratings of 8.5-9.9/10. **[[instructional-agents-multi-agent-course-gen|Instructional Agents]]** uses a multi-agent framework structured around the ADDIE model, with role-based agents (Teaching Faculty, Instructional Designer, Course Coordinator) collaborating to generate complete course materials. **[[courseblueprint-adaptive-video-generation|CourseBlueprint]]** provides a structured pipeline for adaptive [[pedagogy|pedagogical]] [[video-education|video generation]] grounded in course corpora, demonstrating that explicit pedagogical structure — not just AI fluency — is essential for educational content generation. [[generative-ai|Generative AI]] platforms can also embody learning-design principles in the content they produce: [[ai-modeling-problem-generation-platform-2026|an AI-powered platform for generating mathematical modeling problems]] combined established design principles with [[prompt-engineering|retrieval-augmented generation]], developed through the ADDIE approach to produce pedagogically grounded tasks and recommendations that conventional content generators lack. Yet the payoff of such AI-assisted content and lesson generation is mediated by the teacher's own expertise: [[choi-teacher-ai-interaction-lesson-design-2026|Choi et al. (2026)]] found that experienced teachers critically adapt AI-generated lesson ideas to students and context (re-prompting and elaborating on output), whereas novices tend to accept AI suggestions directly — so the pedagogical value of AI content tools depends on the teacher's experience and AI proficiency, not the tool alone. A systematic review of [[wang-teacher-ai-co-design-review-2026|teacher–AI co-design of learning tasks]] (Wang, Liu & Islam 2026) confirms the pattern at scale across 28 studies (2015–2025): GenAI is used mainly for lesson planning, prompt generation, and creative ideation, and the dominant collaboration mode is AI as assistant/content generator rather than a fuller co-designer — with efficiency, responsiveness, [[creativity]], and [[equity-in-ai-education|equity]] recurring as affordances. [[talebzadeh-ai-group-activity-roles-2026|Talebzadeh (2026)]] sharpens the teacher-expertise finding for group activity design: experienced teachers produce richer, more synergistic, better ZPD-aligned role architectures in AI-designed cooperative activities than novices regardless of AI familiarity, framing "pedagogical prompt literacy" as the lever that turns AI output into effective [[collaborative-learning|differentiated group learning]]. **Pedagogically grounded AI tutoring** applies instructional design principles to AI system design. **[[didactical-teacher-assistant-dimensional-modeling|Brisson et al.]]** built a didactically-driven [[llm]] teacher assistant where tutoring strategy is encoded in an explicit external layer — making content selection and didactic structuring traceable and reproducible, directly addressing opacity concerns in [[rethinking-scaffolding-llm-tutors]]. **[[instructional-guidance-genai-learning|Hou et al.]]** demonstrated that a five-step [[prompt-engineering|prompting]] framework grounded in Generative [[learning-theories|Learning Theory]] significantly improved higher-order cognitive outcomes, showing that instructional guidance — not just AI access — determines learning effectiveness. Both connect to [[scaffolding]] and [[intelligent-tutoring]]. **Frameworks and evaluation** provide structured approaches. **[[bridging-instructional-design-framework-math]]** and **[[cotal-formative-assessment-scoring-2026|CoTAL]]** demonstrate [[human-in-the-loop-ai|human-in-the-loop]] design principles. **[[genai-mindtool-generative-learning]]** positions AI as a "mindtool" — a cognitive partner that extends rather than replaces learner thinking — directly applying instructional design theory to AI integration. **[[ludia-udl-ai-thought-partner-2026|LUDIA]]** applies Universal Design for Learning principles to create an accessible AI thought partner for educators, connecting instructional design to [[inclusive-learning]]. **[[airis-cognitively-activated-ai-physics-2026|AIRIS]]** (Activate–Inquire–Reflect) is a task-structuring framework for cognitively activated AI use that bounds the AI's contribution so that prediction, interpretation, and evaluation remain the learner's — an AI-specific adaptation of inquiry cycles grounded in [[self-regulated-learning]], Cognitive Load Theory, and [[human-ai-collaboration]]. Complementing these design frameworks, the [[dohn-boundary-object-classifying-genai-learning-activities-2026|Dohn et al. (2026) taxonomy]] offers a *classification* rather than a design method: six categories (Learning Objective, Content, Representation Format, Epistemic [[student-engagement|Engagement]], Social Design, Artifacts) that let designers and [[research-methods-aied|researchers]] describe, compare, and imagine GenAI learning activities by making explicit why, what, how, with what, and with whom learners engage GenAI — built as a boundary object through postdigital dialogue. Smart-classroom frameworks extend this to [[teacher-education|teacher education]]: [[instructional-design-proficiency-masters-math-2026|Zhu, Liang, Mao, and Wang (2026)]] propose a three-dimensional framework for smart education — learning effectiveness, information and communication technology (ICT), and classroom organization — and instantiate it in a [[math-education|mathematics]] M.Ed. course that integrates [[automated-assessment|automated scoring]], personalized recommendations, and multi-[[ai-feedback-quality|AI feedback]] across pre-, in-, and post-class stages. A quasi-experiment showed significant gains in students' ability to formulate precise, professionally grounded instructional objectives, yielding the transferable **D-T-E Model** (Disciplinary Demand–Technological Empowerment–Evaluation Loop) — [[discipline-specific-aied|discipline-specific]] guidance for [[educational-development|teacher educators]] moving smart-education concepts into practical instructional design practice. **Rubric-guided prompting as a design lever.** [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] demonstrated that the rubric functions as a mediating interface between human pedagogical intent and machine inference: treating assessment criteria as revisable design artifacts — rather than fixed instruments — and iteratively co-refining them with the LLM raised LLM–human agreement on student design work from 54.75% to 81.25%. Rubrics engineered for LLMs must balance precision and flexibility — too vague invites free interpretation, too rigid reduces the model to pattern-matching — and role-aware prompting (instructor, peer reviewer, grant reviewer) yielded distinct evaluative feedback. This positions rubric engineering as a concrete learning-design practice for shaping AI evaluation behavior, with human-in-the-loop oversight remaining essential. **AI agents for instructional design** extend the field into [[agentic-ai|agentic AI]]. **[[jeon-isd-agent-bench-2026|ISD-Agent-Bench]]** is the first standardized, theory-grounded benchmark for evaluating LLM-based instructional design agents — its 25,795-scenario Context Matrix (51 contextual variables × 33 ISD sub-steps from ADDIE) shows that agents grounded in classical ISD frameworks (ADDIE, Dick & Carey, Rapid Prototyping ISD) outperform theory-free agents, empirically validating that instructional design is a structured discipline rather than a generic prompting task. Agents are not only *builders* of designs but also *critics* of them: [[ai-web-agents-lesson-design-2025|Wang, Mitchell & Piech (2025)]] use a single autonomous web agent that navigates a multi-step online lesson like a student to evaluate a learning design *before* real learners engage — its description of the student experience predicts where novices will drop out and surfaces actionable design feedback, outperforming every baseline and even a simulated cohort of students on a global CS1 course. This frames pre-launch, agentic evaluation as a low-cost complement to human design iteration. **[[wang-multi-agent-systems-learning-designers-2025]]** and **[[instructional-agents-multi-agent-course-gen|Instructional Agents]]** explore multi-agent frameworks that orchestrate role-based agents around instructional-design models, while **[[ai-tpack-teacher-multi-agent-workflow|AI-TPACK]]** examines how teachers and agents jointly apply technological-pedagogical-content knowledge. This work connects instructional design to [[benchmark|benchmarking]], [[ai-ed-evaluation]], and the design of [[curriculum-design|curriculum]] at scale. ### Connections to related concepts Learning design is the bridge discipline of [[ai-education|AI in education]] — it connects [[curriculum-design]] (what to teach) with [[scaffolding]] (how to support learners), [[educational-development]] (how to prepare educators), and [[generative-ai]] (the tools themselves). It is tightly coupled with [[teacher-role]] because AI tools reshape what learning designers and teachers do, and with [[ai-literacy]] because effective AI integration requires educators to understand AI capabilities and limitations. The [[learning-sciences|learning sciences]] are the research field behind these principles: where this page covers the professional practice of creating learning experiences, the learning sciences study that practice and its designs empirically and generate the cognitive, motivational and social principles that learning design then operationalizes. ### How learning design determines learning gains Learning design is the lever that decides whether AI produces [[learning-gains|learning gains]] or merely AI-inflated performance. The knowledge base's evidence is consistent on this: **the same AI tool yields large gains or net harm depending on how the learning experience is designed around it.** [[instructional-guidance-genai-learning|Hou et al.]] showed that a five-step prompting framework grounded in learning theory significantly improved higher-order cognitive outcomes, while access to AI alone did not; [[genai-mindtool-generative-learning|mindtool]] and [[airis-cognitively-activated-ai-physics-2026|AIRIS]] frameworks preserve the learner's cognitive work so that durable gains (rather than task-efficiency) result. Design choices that protect [[learning-gains]] — scaffolding that requires a student attempt, [[formative-assessment]] with unassisted outcome measures, and pedagogical structure that keeps the learner the agent — mirror the field's finding (see [[learning-gains]]) that AI is a strong gain when it coaches and a harm when it answers. Conversely, poorly designed AI-integrated lessons fall prey to the [[cognitive-offloading|performance-learning gap]], where apparent success masks no learning. ### Practical guidance for designers and developers For instructional designers, course developers, and engineers building AI-assisted learning experiences, the knowledge base's findings translate into actionable practice. One boundary is worth marking before the practices themselves: learning design as this page describes it is the design of a course for a known cohort, whereas the same principles baked into a product that many courses — taught by people the designer will never meet — will use are the work of [[educational-technology-developers]], where defaults, configurability and documentation carry pedagogical weight: **Ground AI generation in a structured instructional model.** AI content is only as good as the pedagogical structure behind it — explicit structure, not AI fluency, determines quality. Design around a recognized model (ADDIE, Dick & Carey, rapid prototyping) and encode pedagogical decisions explicitly rather than relying on the model to infer them.([[courseblueprint-adaptive-video-generation]])([[jeon-isd-agent-bench-2026]])([[didactical-teacher-assistant-dimensional-modeling]]) **Adopt a principle-level framework as well as an instructional model.** A course-level model structures one design; a published framework sets the criteria that many designs should satisfy. An example worth reading in full is Digital Promise's *Powerful Learning with Emerging Technology*, which organizes its guidance under three principles — Evidence-Based, Learner-Centered, Skill-Building — each expanded into practices and strategies, and attaches [[privacy]], [[explainable-ai|explainability]] and fairness to particular practices as safety obligations rather than optional extras.([[powerful-learning-with-emerging-technology-2025]]) **Use role-based multi-agent workflows for content production.** Instead of one generic prompt, orchestrate distinct agents/roles (teaching faculty, instructional designer, course coordinator) that collaborate through a defined pipeline — this mirrors how real course teams work and yields more complete materials than a single prompt.([[instructional-agents-multi-agent-course-gen]])([[wang-multi-agent-systems-learning-designers-2025]]) **Provide instructional guidance, not just AI access.** Whether learners interact with AI directly or with AI-generated materials, guidance built on learning theory (e.g. a stepwise prompting [[scaffolding|scaffold]] grounded in generative-learning principles) drives higher-order outcomes; access alone does not. Design the learning activity around how the mind learns, and treat AI as a cognitive "mindtool" that extends thinking rather than replacing it.([[instructional-guidance-genai-learning]])([[genai-mindtool-generative-learning]]) **Make content traceable and reviewable.** Let a human designer review and correct AI output before it reaches learners, and structure AI generation so the pedagogical rationale (why this content, in this order) is inspectable — addressing both quality and the opacity concerns that undermine [[trust]]-generated instruction.([[bridging-instructional-design-framework-math]])([[cotal-formative-assessment-scoring-2026]]) **Design for [[accessibility]] from the start.** Apply [[universal-design-for-learning|UDL]] principles when building AI tools and AI-generated materials so they serve diverse learners, rather than retrofitting accessibility after the fact.([[ludia-udl-ai-thought-partner-2026]]) **Plan for the delivery medium.** Instructional design for [[online-teaching-and-learning|online teaching and learning]] is not a neutral translation of in-person design — the medium changes what scaffolding, assessment, and interaction are viable, and AI multiplies both the opportunities (scalable [[personalized-learning|personalization]], always-on support) and the risks ([[academic-integrity|integrity]], [[cognitive-offloading|cognitive offloading]]) designers must plan for. Design the AI's pedagogical wrapper as deliberately in online as in face-to-face contexts. **Evaluate against a benchmark, not vibes.** If you're building an instructional-design agent, evaluate it against a standardized, theory-grounded benchmark (e.g. [[jeon-isd-agent-bench-2026|ISD-Agent-Bench]]) so you can measure whether grounding in a real ISD framework actually improves output over a generic LLM.([[jeon-isd-agent-bench-2026]]) - **AI is reshaping instructional design practice.** [[kibar-ilgaz-ai-instructional-design-review-2026|Kibar & Ilgaz (2026)]] [[meta-analysis-systematic-review|systematically review]] 28 studies (2020-2025) and find AI assists designers with content generation, templates, and personalization, and is conceptualized as a co-worker/collaborator/partner rather than just a tool — though pedagogical alignment and practitioner readiness remain challenges. ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[online-teaching-and-learning]] — Online Teaching and Learning - [[curriculum-design]] - [[scaffolding]] - [[educational-development]] - [[teacher-role]] - [[ai-literacy]] - [[generative-ai]] - [[intelligent-tutoring]] - [[personalized-learning]] - [[adaptive-learning]] - [[formative-assessment]] - [[higher-ed]] - [[k-12]] - [[agentic-ai]] - [[inclusive-learning]] - [[universal-design-for-learning]] - [[learning-theories]] - [[learning-sciences]] - [[learning-gains]] - [[behaviorism]] - [[educational-technology-developers]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Powerful Learning with Emerging Technology - [[claassen-learning-analytics-genai-learning-design-2026]] — LA and GenAI in learning design decision-making - [[tang-chatbots-learning-design-2026]] — Chatbot use in learning design: designers dwell on outcomes and pedagogy rather than content generation (Tang et al. 2026) - [[zhou-constructive-alignment-genai-business-2026]] - [[ai-student-engagement-online-learning-review-2025]] - [[ai-communities-of-inquiry-2026]] - [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] — Ontology-based layered hybrid knowledge model for personalized e-learning - [[rewriting-curriculum-genai-pedagogy-2026]] — Rewriting the curriculum: GenAI-driven pedagogical change - [[lin-llm-interactive-lesson-generation]] — Automatic LLM creation of interactive learning lessons (Lin et al. 2025) - [[choi-teacher-ai-interaction-lesson-design-2026]] — Teacher-AI interaction patterns in lesson design across experience and AI proficiency (Choi et al. 2026) - [[long-ai-higher-ed-engagement-teaching-methods-2026]] — AI in higher ed: engagement + mediating role of teaching methods - [[curriculum-as-code-instructional-design-2026]] - [[dohn-boundary-object-classifying-genai-learning-activities-2026]] — Taxonomy (boundary object) for classifying GenAI learning activities - [[instructional-agents-multi-agent-course-gen]] - [[didactical-teacher-assistant-dimensional-modeling]] - [[instructional-guidance-genai-learning]] - [[courseblueprint-adaptive-video-generation]] - [[bridging-instructional-design-framework-math]] - [[cotal-formative-assessment-scoring-2026]] - [[genai-mindtool-generative-learning]] - [[ludia-udl-ai-thought-partner-2026]] - [[learnity-graphs-lifelong-learning-framework-2026]] - [[pchl-he-framework-genai-content-creation-2026]] - [[jeon-isd-agent-bench-2026]] - [[ai-web-agents-lesson-design-2025]] — AI Web Agents: autonomous web agent evaluates lesson designs and predicts student dropout before students engage (Wang, Mitchell & Piech 2025) - [[airis-cognitively-activated-ai-physics-2026]] — AIRIS: A Framework for Cognitively Activated AI Augmentation in Physics - [[wang-multi-agent-systems-learning-designers-2025]] - [[ai-tpack-teacher-multi-agent-workflow]] - [[halani-designing-for-reach-2026]] — Designing for Reach: Seven Levers and the Student Alone with AI - [[vargas-situated-learning-ai-review-2024]] - [[vargas-ai-catalyst-situated-learning-2026]] - [[panciroli-ai-literacy-episodes-situated-learning]] - [[fowlin-operationalizing-learning-principles-ai]] - [[cfes-p24-multimodal-slide-auditing-2026]] — CFES-P24: Benchmarking Multimodal LLMs for Slide Auditing - [[learnai-just-in-time-ai-cocreation-university-2026]] — LearnAI: Just-in-Time AI Co-Creation Across Disciplines - [[ai-video-dual-gatekeeping-2026]] — When Saying No Makes Better Videos: Dual Gatekeeping for Pedagogically Grounded AI Content Creation - [[rhaimi-productivemath-2025]] — ProductiveMath: AI to Support Productive Failure Problem Design - [[kibar-ilgaz-ai-instructional-design-review-2026]] — AI and Instructional Design Practice: A Systematic Review (Kibar & Ilgaz 2026) - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[guillen-curriculum-genai-teacher-competence-2026]] — Assessing Teacher Digital Competence for GenAI Curriculum Design (Guillén-Gámez 2026) - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[preservice-teacher-agency-genai-design-learning-2026]] — Pre-service teacher agency during GenAI interactions in design for learning (Krushinskaia, Elen & Raes 2026) - [[making-ai-annoying-constrained-writing-2026]] — Making AI annoying on purpose: constraint in AI-supported writing (Konradt, Boote & Taub 2026) - [[instructional-design-proficiency-masters-math-2026]] — Smart-classroom model and D-T-E loop improving M.Ed. instructional design proficiency in mathematics (Zhu et al. 2026) - [[ai-modeling-problem-generation-platform-2026]] — AI-powered platform generating mathematical modeling problems (ADDIE, RAG) - [[wang-teacher-ai-co-design-review-2026]] — Teacher–AI co-design of learning tasks: trends and perspectives (Wang et al. 2026) - [[talebzadeh-ai-group-activity-roles-2026]] — Architecture of roles in AI-designed differentiated group activities (Talebzadeh 2026) - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design --- ## [Educational Development](https://edtechdev.github.io/aied/concepts/educational-development/) > **Educational Development** (also known as *faculty development*) — the processes, programs, and institutional supports that help educators develop the skills, confidence, and [[learner-identity|professional identity]] to teach effectively with AI. Educational development spans individual training, [[curriculum-design|curriculum]] redesign, [[educational-policy-ai|institutional policy]] change, and the cultural work of making sense of what GenAI means for the academic profession. ## Questions to Consider - If faculty in the same department describe AI as a human-like 'assistant' on one hand and a mere 'tool' or 'search engine' on the other, can a training program really succeed before those underlying mental models are surfaced? What might happen if development skips that step? - Educational development is often treated as closing a skills gap — teach faculty the tools and they'll adopt them. But some [[research-methods-aied|research]] frames GenAI integration as an identity-level change ('what's the point of teaching in a GenAI world?'). Which framing do you think is more accurate, and what different actions does each one imply? - Readiness frameworks like EPIQ-AI split faculty readiness into epistemic, [[pedagogy|pedagogical]], institutional, and quality-and-compliance domains. If readiness is a sociotechnical alignment problem rather than an individual skills gap, what does that say about where a lone 'AI workshop' is likely to fall short? - When have you felt your own assumptions about a new technology surfaced and shifted — through discussion, metaphor, or shared language rather than instruction? What role do you think shared language and metaphor-analysis play in helping educators make sense of AI? - Consider a low-tech 'metaphor workshop' where faculty, staff, and students articulate whether AI feels like a Swiss army knife, a helper, a black box, or a competitor. Could surfacing fears (including the fear that AI will replace teaching roles) do more for adoption than more technical training? ## Introduction Educational development is the professional and institutional work through which faculty build the capability to design, teach and assess well — and in the AI era it has become a precondition for any pedagogical change to take effect. The concept spans the individual ([[ai-literacy|AI literacy]], confidence, [[teacher-ai-competency|competency]], attitudes) and the institutional (standards, policy, quality assurance, and alignment between stated AI expectations and actual assessment design). A recurring finding across this knowledge base is that the evidence for effective practice substantially leads what most institutions have implemented, which makes development work — not more primary research — the binding constraint. The [[learning-sciences|learning sciences]] produce the findings this work carries: where that field establishes empirically how people learn and which designs change outcomes, educational development is the institutional practice that moves such evidence into how faculty actually teach and assess. Its work, though, is bounded by an institutional edge: development builds capability in the staff inside an institution, while the tools those staff are then asked to adopt are usually built outside it by [[educational-technology-developers]], whose design and procurement decisions — defaults, configurability, what survives the funding — arrive as constraints on the program rather than as products of it. ## Educational development in the AI era - **Readiness frameworks:** [[sangwa-epiq-ai-faculty-readiness-2026|The EPIQ-AI framework]] identifies four readiness domains: epistemic, pedagogical, institutional, and quality-and-compliance. Faculty readiness is a sociotechnical alignment problem, not just an individual skills gap. - **Standards for technology integration:** [[crompton-faculty-technology-integration-standards-2026|Crompton et al.]] use [[design-based-research|design-based research]] to develop faculty standards for technology (incl. AI) integration in [[higher-ed|higher education]] institutions. - **Adoption and confidence:** [[teacher-ai-adoption-confidence|Teacher AI adoption research]] identifies concerns, support, confidence, and attitudes as key predictors. Faculty-development programs must address all four. - **Curriculum integration:** [[institutional-change-framework-ai|Institutional change frameworks]] and [[ai-assessment-scale-reform|assessment reform]] require faculty to redesign courses, not just add AI tools. - **Training programs:** [[crewscaler-ai-upskilling-framework|AI upskilling frameworks]] and [[ai-tpack-preservice-math-teachers|TPACK-based preservice training]] provide models for structured faculty [[ai-education|AI education]]. - **[[governance]] and policy:** [[genai-policies-higher-ed-computing|Institutional AI policy analysis]] documents the gap between institutional ambitions and faculty support capacity. ### Metaphors and shared language in educational development Faculty hold heterogeneous, often deeply ambivalent mental models of AI, and effective development must surface and work with them. [[engineering-faculty-metaphors-ai-understanding-2026|Gerhardt et al.]] show that [[engineering-education|engineering]] instructors frame AI through metaphors that both construct and constrain understanding (AI as a human-like "assistant" vs. a "tool" or "search engine"); because instructors within the same department often hold fundamentally different conceptualizations, a shared, accurate language about [[generative-ai|GAI]] is a prerequisite for effective faculty development and productive departmental adoption discussions. A practical method for surfacing these mental models is the **metaphor-analysis workshop**. [[fear-awe-genai-metaphor-workshops-2025|Vallis, Wilson & Casey (2025)]] designed and validated a low-tech, collaborative workshop for faculty, staff, and students to articulate their metaphors for generative AI. Participant metaphors clustered into four categories — *Functions* (tool-like: "Swiss army knife"), *Roles* (human-like: "helper," "frenemy"), *Qualities* (unknowable: "black box," "slippery slope"), and *Agency* (threatening: "competitor," "sinister robot") — surfacing persistent tensions between human vs. machine [[agency]] and the known vs. the unknowable. The workshop helped participants surface assumptions, connect across roles, feel "not alone," and think about the [[ethics]] of GenAI, without requiring technical expertise. As a development tool, it gives program designers a low-barrier way to surface a team's assumptions *before* redesigning [[assessment]] or rolling out policy, and to address fears head-on — including the fear that AI will replace [[teacher-role|teaching roles]].([[fear-awe-genai-metaphor-workshops-2025]]) ### GenAI as identity work, not just upskilling A threshold-informed view reframes GenAI integration as an **ontological transformation** for faculty, not a skills gap. Applying threshold concept theory (Meyer & Land), Laidlaw argues that GenAI exhibits all five threshold characteristics — *transformative* (reconstructs assumptions about assessment, pedagogy, and professional role), *troublesome* (violates beliefs about originality, human agency, and effort–achievement), *irreversible*, *integrative* (connects technology, pedagogy, epistemology, and identity), and *bounded* (GenAI fluency becomes a new marker of professional currency).([[laidlaw-genai-identity-crisis-faculty-2026]]) On this account, faculty asking "what's the point of teaching in a GenAI world?" are not deficient in competence; they are in a **liminal threshold-crossing phase** where anxiety, resistance, and confusion are necessary parts of transformation, not obstacles to eliminate. Skills-based training that answers a competence question faculty are not asking can become peripheral to the real transformation, and well-intended governance can lapse into an "enforcement illusion" — communicating rules rather than supporting change.([[laidlaw-genai-identity-crisis-faculty-2026]]) **Design implication:** complement (don't replace) skill building with **identity-supporting practices** — open sessions with identity questions rather than technical demos, run ongoing [[discipline-specific-aied|discipline-specific]] cohorts where faculty explore what GenAI means for their field's purpose, create peer-mentoring structures that honor different transformation timelines, and distinguish fear-based hesitation (which benefits from support) from principled non-adoption grounded in legitimate disciplinary values (which deserves respect).([[laidlaw-genai-identity-crisis-faculty-2026]]) The metaphor-workshop model above is one concrete instantiation of this: rather than starting from technical [[lifelong-learning|upskilling]], it opens with the interpretive, identity-laden question of what GenAI means to participants. **Empirical support for the identity-work framing.** [[farazouli-navigating-uncertainty-teachers-genai-2026|Farazouli et al. (2026)]] provide direct [[qualitative-research|qualitative]] evidence for this account: 24 Swedish university teachers described GAI's emergence as alarming and overwhelming, reporting a "state of vulnerability" (low confidence, insecurity, fear of "not being ahead of students") and feeling "stuck" between utopian and dystopian discourses. Their limited knowledge and experience, alongside feelings of vulnerability, highlight teachers' lack of readiness to navigate, assess, and adopt GAI — and the authors argue institutions must provide **designated spaces and time** for teachers to experiment with GAI, exchange ideas, and collaboratively develop practices and guidelines at institutional, departmental, and course levels. This grounds the identity-work and support-oriented approach in teachers' own reported experience rather than only in theory. **A programme that worked the identity question directly.** [[chick-faculty-development-ethical-ai-2026|Chick, Morello & Staffey (2026)]] report a six-week summer institute that took ten faculty from Education, Health Sciences, Business and other programs out of low AI fluency (mean self-rating 1.89 on a six-point scale) to 60% reporting high confidence in guiding [[student-ai-interaction|student AI use]] and 90% positive perceptions, and they attribute the shift less to tool training than to structured, safe experimentation that opened with the existential question of professional value rather than with demos. The vocabulary change is the marker they lean on: early framings of AI as a "cheating machine" or a "dehumanizing force" gave way to "curious assistant," and their four themes — fear to curiosity, a desire for ethical clarity the policy vacuum did not supply, [[inclusive-learning|inclusive design]] as an [[equity-in-ai-education|equity]] amplifier, and faculty moving from gatekeepers to guides — describe a partial, deliberate reorientation of instructional identity. The ten capstone redesigns are the harder evidence, with eight of ten positioning AI as a scaffold supporting rather than replacing thinking, and the authors' **symbiotic pedagogy** formalises what participants built: complementary design, transparent process, critical integration, embedded ethical reasoning and iterative refinement. The caveat is structural rather than individual: the same participants named contradictory policy signals, personal tool subscriptions and no time release for the redesigns, so the authors conclude that resistance was never the problem and that piecemeal responses leaving policy and infrastructure untouched will fail even an enthusiastic cohort. **Development as mediation, and the first question a developer asks.** [[beyond-the-algorithm-academic-developers-digital-mediators-2026|Sithole (2026)]] studies academic development — the field's older name, and the term used by the *International Journal for Academic Development*, which published the paper — at two South African Historically Disadvantaged Institutions, and reframes developers not as technical implementers but as **digital mediators** working across three registers: pedagogical (translating AI discourse into viable teaching, learning and [[assessment]] practice), ethical (holding automation against [[academic-integrity]], innovation against care), and institutional (working the liminal space between management imperatives and academic concerns). Two findings matter for program design. First, the filtering question is contextual before it is technical — "what does this mean for teaching here, for our students from poor schools?" — and slowing adoption down counts as part of the work rather than resistance to it. Second, the affective cost is structural, not personal: developers were expected to project expertise in tools whose implications they were still working out ("I feel like an impostor. I am learning AI as I go, but the institution expects expertise"), and their response was peer [[professional-training|professional learning]] aimed at judgment rather than skill, including deliberately trying to break the tools. The paper's analytical contribution is to separate [[digital-divide|digital inequality]], a distributive problem, from **algorithmic coloniality**, an epistemic one that survives full access — and to place both inside [[global-south|Global South]] [[higher-ed|higher education]], where development work is expected to be transformative rather than a neutral support service. **Guilt and moral discomfort as a developmental signal.** [[vassallo-ai-guilt-complex-faculty-2026|Vassallo (2026)]] supplies the affective counterpart to the identity account. In a survey of academic staff at a Maltese university (109 responses, a 3.5% rate the paper reads as itself informative about the state of ethical debate), guilt tracked concealment rather than disclosure — limiting AI use because of unease correlated with the AI Guilt Index at r = .62 and avoiding disclosure to colleagues at r = .50 — while formal [[ai-use-disclosure|disclosure]] in academic outputs showed no relationship (r = .08). Only 31.2% saw clear institutional guidelines, and the correlation between that clarity and lower guilt was weak (r = −.25, p = .010), which the paper reads as evidence that top-down mandates cannot resolve a partly emotional and identity-based problem. Its recommendations are directly developmental: treat guilt and identity concern as normal transitional responses rather than faults to correct, use structured low-stakes experimentation to move staff through anticipatory anxiety, and lean on mentorship by senior academics, since guilt was highest among early-career staff (M = 2.71) and lowest among senior ones (M = 2.03). ### Practical guidance for program designers For faculty developers, academic leaders, and [[stakeholders|instructional designers]] planning AI [[professional-training|professional development]], the knowledge base's evidence suggests: **Address the four adoption drivers, not just knowledge.** Confidence, attitudes, support, and concerns predict whether faculty actually adopt AI — a knowledge-only workshop that ignores these is unlikely to change practice. Design development to build confidence through hands-on use, provide ongoing support (not one-shot training), and actively surface and respond to faculty concerns.([[teacher-ai-adoption-confidence]]) **Treat readiness as a sociotechnical alignment problem.** The [[sangwa-epiq-ai-faculty-readiness-2026|EPIQ-AI framework]] shows faculty readiness spans epistemic, pedagogical, institutional, and quality-and-compliance domains. Programs that only train the individual miss the institutional levers (policy, workload, incentives, quality standards) that enable or block change — align those alongside training.([[sangwa-epiq-ai-faculty-readiness-2026]]) **Build toward curriculum redesign, not tool adoption.** The goal is faculty redesigning courses and assessment, not just adding AI tools. Ground professional development in course-level redesign work and assessment reform, and give faculty structured frameworks for doing so (e.g. [[ai-assessment-scale-reform|assessment scales]], [[institutional-change-framework-ai|institutional change frameworks]]).([[institutional-change-framework-ai]])([[ai-assessment-scale-reform]]) **Run the series on the room's own data rather than imported best practice.** [[physics-faculty-learning-community-ai-2026|Perl-Nussbaum and Finkelstein (2026)]] document a worked model from a large public R1 physics department: six biweekly sessions of 60 to 75 minutes, nineteen faculty participating across the series and about ten at any meeting, in which every session opened with local data or department-sourced materials, moved to small-group testing of AI against real anonymized student homework, and closed in collective discussion. What faculty found by testing was more useful than any package the facilitators could have imported — AI solved every problem they tried, but its feedback on real student work was uneven, prompt-dependent and surface-focused unless given explicit goals and rubrics, and students prompted shallowly and modified output superficially unless productive use was modeled. The first concrete output was a living repository of five entries — syllabus policy statements, classroom discussion materials, the student survey, AI-integrated homework tasks and assessment structures — which is why the authors treat the session structure itself as the argument rather than any content it delivered. **Anchor in a competency framework.** Development should also build **pedagogical [[prompt-engineering|prompt literacy]]** — the capacity to encode pedagogical intentions into prompts — since [[talebzadeh-ai-group-activity-roles-2026|Talebzadeh (2026)]] finds it is pedagogical expertise, not AI fluency, that determines the quality of teachers' AI-assisted design work.([[crewscaler-ai-upskilling-framework]])([[ai-tpack-preservice-math-teachers]])([[talebzadeh-ai-group-activity-roles-2026]]) **Sequence understanding before automation.** A [[qualitative-research|critical discourse analysis]] of 14 pieces of gray literature published between November 2022 and April 2025 advising practitioners on generative AI for constructive alignment finds that capability must precede tool access: because quality of engagement precedes quality of understanding in a [[constructivist]] framework, risk scales inversely with familiarity, and novices are the most exposed to plausibility-driven acceptance of plausible-looking but pedagogically thin outputs. [[mcinnes-salvaging-constructive-alignment-genai-2026|McInnes et al. (2026)]] further argue that development units should resist a technicist brief — helping faculty "acquire fluency" in tools that do "the heavy lifting" of [[learning-design|course design]] — because that framing shuts developers out of a closed educator–GenAI loop and recasts values-led, relational development as a mediating mechanism for software adoption. The defensible alternative keeps educational developers as curators and rule-setters of any institutionally bounded [[rag|retrieval-augmented]] alignment assistant that poses probing questions, flags under- and over-assessment and generates alternatives but never finished artifacts, with escalation to [[human-in-the-loop-ai|human judgment]] on accreditation and cross-program matters treated as a feature rather than an exception. **Surface and address fears and mental models.** Before redesigning teaching, use a metaphor-analysis workshop ([[fear-awe-genai-metaphor-workshops-2025|Vallis, Wilson & Casey 2025]]) to surface a team's assumptions and anxieties about GenAI — including the fear that it will replace teaching roles — and build development that responds to them rather than ignoring them.([[fear-awe-genai-metaphor-workshops-2025]]) **Model AI literacy and measure real gains.** Faculty development should itself embody the practices being taught — using AI pedagogically, evaluating outputs critically — and should assess demonstrated competence rather than [[self-report-measures|self-reported]] confidence, since self-perception reliably overestimates AI skill.([[ai-literacy-assessment-misalignment]])([[genai-pd-ai-pck-learning-gain-2026]]) **Do not assume subject-matter expertise carries over to GenAI.** [[sutedjo-faculty-genai-tpack-21-2026|Sutedjo, Chowdhury & Liu (2026)]] surveyed 127 U.S. research-university faculty using the [[tpack|TPACK-21]] instrument and found content knowledge (CK) showed **no significant correlation** with technological knowledge or any technology-integrated domain (r =.11–.15, ns) — disciplinary expertise did not predict GenAI-integration knowledge. The technology-integrated domains (TPK, TCK, TPACK) inter-correlated so strongly (r =.81–.91) that they may function as a single GenAI-integration factor. Practically: build GenAI integration through **discipline-specific** activities that connect GenAI affordances to each faculty member's subject matter, and treat the technology-integrated domains as one shared GenAI-literacy foundation rather than train them as separate skills.([[sutedjo-faculty-genai-tpack-21-2026]]) **Discipline-specific smart-classroom models.** [[instructional-design-proficiency-masters-math-2026|Zhu, Liang, Mao, and Wang (2026)]] show how a [[math-education|mathematics]] M.Ed. course can be enhanced with intelligent educational [[ai-technologies|technologies]] ([[automated-assessment|automated scoring]], personalized recommendations, multi-[[ai-feedback-quality|AI feedback]]) integrated across pre-, in-, and post-class stages within a three-dimensional smart-classroom framework. Their quasi-experiment found statistically significant gains in instructional-objective design proficiency, offering a transferable **D-T-E Model** (Disciplinary Demand–Technological Empowerment–Evaluation Loop) for [[teacher-education|teacher educators]] and educational developers looking to move smart-education frameworks from macro concepts into discipline-specific practice. **benchmark against what faculty themselves report.** [[watson-rainie-ai-challenge-faculty-survey-2026|Watson & Rainie (2026)]]'s survey of 1,057 US faculty finds the institutional layer thin in ways program designers can measure directly: 59% judged their school unprepared to use generative AI effectively for preparing students for the future and 68% said it had not prepared faculty to use it for teaching and mentoring, while the structural response ran to a task force in 55% of institutions but an AI literacy general education outcome in only 13%. Faculty had not waited for policy — 87% wrote their own assignment-level rules against 48% who could point to an institutional one — and they named colleagues' resistance (82%) and unfamiliarity (83%), not mandate, as the obstacles to departmental adoption. For development programs this argues for treating peer norms, shared assignment-level policy language and explicit measures of institutional readiness as part of the intervention rather than leaving them to the policy document. ## Connected Concepts - [[teacher-ai-competency]] - [[teacher-role]] - [[ai-literacy]] - [[educational-policy-ai]] - [[higher-ed]] - [[k-12]] - [[learning-design]] - [[curriculum-design]] - [[professional-training]] - [[teacher-education]] - [[learning-sciences]] - [[educational-technology-developers]] - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Guidance for product teams as well as educators - [[mcinnes-salvaging-constructive-alignment-genai-2026]] — Critical discourse analysis of GenAI alignment advice: developers recast as technology trainers - [[crompton-faculty-technology-integration-standards-2026]] — Faculty standards for technology integration (DBR) - [[bilgic-sever-ethical-dimensions-ai-higher-ed-2026]] — Ethical dimensions of AI: faculty and student views - [[alharbi-ethical-genai-eap-2026]] - [[nicola-richmond-programwide-assessment-genai-2025]] - [[espino-ai-business-education-review-2026]] - [[engineering-faculty-metaphors-ai-understanding-2026]] — How Engineering Faculty Metaphors Construct (and Constrain) AI Understanding - [[fear-awe-genai-metaphor-workshops-2025]] — Fear and Awe: Making Sense of Generative AI Through Metaphor (faculty/staff/student metaphor workshop) - [[governing-unseen-ai-literacy-language-teachers-2026]] — Governing the unseen: AI literacy among language teachers - [[sangwa-epiq-ai-faculty-readiness-2026]] - [[teacher-ai-adoption-confidence]] - [[institutional-change-framework-ai]] - [[genai-policies-higher-ed-computing]] - [[ai-tpack-preservice-math-teachers]] - [[crewscaler-ai-upskilling-framework]] - [[ai-assessment-scale-reform]] - [[pchl-he-framework-genai-content-creation-2026]] - [[genai-pd-ai-pck-learning-gain-2026]] - [[genai-higher-education-systematic-review-2026]] - [[laidlaw-genai-identity-crisis-faculty-2026]] — GenAI as identity crisis, not skills gap - [[chen-preservice-teachers-chatgpt-lpa-2026]] — Pre-service teacher ChatGPT acceptance profiles - [[zuo-instructor-power-genai-writing-2026]] — Power relations perceived by college instructors grappling with GenAI in writing (Zuo, Xu & Dunning 2026) - [[reflective-triangle-model-teacher-ai-2026]] — Reflective Triangle Model: AI as cognitive mediator - [[beyond-hype-stakeholder-perceptions-genai-2026]] — Stakeholder perceptions of GenAI in higher ed (Humble & Mozelius 2026) - [[sutedjo-faculty-genai-tpack-21-2026]] — Faculty self-perceived TPACK-21 knowledge for GenAI (Sutedjo, Chowdhury & Liu 2026) - [[instructional-design-proficiency-masters-math-2026]] — Smart-classroom model and D-T-E loop improving M.Ed. instructional design proficiency in mathematics (Zhu et al. 2026) - [[talebzadeh-ai-group-activity-roles-2026]] — Architecture of roles in AI-designed differentiated group activities (Talebzadeh 2026) - [[farazouli-navigating-uncertainty-teachers-genai-2026]] — University teachers' experiences and perceptions of GAI: vulnerability, rethinking assessment, student learning at risk (Farazouli et al. 2026) - [[chick-faculty-development-ethical-ai-2026]] — From fear to innovation: a six-week faculty institute, symbiotic pedagogy, and the institutional conditions it could not supply - [[vassallo-ai-guilt-complex-faculty-2026]] — The AI Guilt Complex: anticipatory guilt, concealment and career-stage differences among academic staff (Vassallo 2026) - [[watson-rainie-ai-challenge-faculty-survey-2026]] — AAC&U/Elon survey of 1,057 US faculty: institutional unpreparedness, thin governance and the 87%/48% policy gap (Watson & Rainie 2026) - [[beyond-the-algorithm-academic-developers-digital-mediators-2026]] — Academic developers as digital mediators in Global South higher education: development as sociotechnical praxis - [[physics-faculty-learning-community-ai-2026]] — A Workshop Series for Effective Use of AI in Uncertain Times: Building a Physics Faculty Learning Community --- ## [History of AI in Education](https://edtechdev.github.io/aied/concepts/history-of-aied/) > **History of AI in Education** — the study of how artificial intelligence and education have co-evolved since the mid-20th century, and how past conceptual, terminological, and design choices continue to shape today's debates about AI in learning. Rather than treating generative AI as a sudden, unprecedented force, a historical lens reveals recurring tensions — control vs. agency, standardization vs. creativity, automation vs. augmentation — that have structured AI in education since its origins. ## Questions to Consider - Today's generative AI is often described as an unprecedented revolution. How might that framing ('chronocentrism') prevent us from seeing the recurring patterns beneath it? - The field was almost named 'cybernetics' instead of 'artificial intelligence.' How might that single naming choice have shaped whether we see machines as independent intelligences or as part of interconnected feedback systems? - The page frames AIED history as a tension between control (Anderson's cognitive tutors optimizing instruction) and agency (Papert's constructionism empowering learners). Where do your own AI-in-education choices fall on that spectrum — and why? - Personalization, the page argues, has two meanings: varying the path to the same outcome, or enabling diverse, learner-directed outcomes. Which kind of personalization does your institution actually pursue, and which do you think it should? - If today's GenAI debates 'mirror' these historical tensions, what past mistakes might we be repeating — and what might we avoid by remembering them? - The page claims technical design choices embed value judgments about learning's purpose. Can you identify a design decision in an AI education tool you've used that quietly favored standardization over learner agency — or vice versa? ## Introduction The field's history is not a linear progression but a series of contingent decisions whose effects still echo. Understanding this history guards against "chronocentrism" (the bias of treating current developments as uniquely revolutionary), helps avoid repeating past mistakes, and reveals how choices about framing and terminology — not just technology — steer the trajectory of AI in education. ### Key historical threads **From cybernetics to "artificial intelligence."** When John McCarthy planned the 1956 Dartmouth Summer Research Project, he deliberately chose the term *artificial intelligence* over *cybernetics* — partly to distance the field from Norbert Wiener and his focus on analog feedback. This naming decision framed machine capabilities in direct comparison to human cognition (anthropomorphic framing), steering research priorities, public perception, and ethical debate for decades. Had the field kept the cybernetic framing, the article imagines we might now speak of "cybernetic learning systems" emphasizing feedback and interconnection rather than "AI tutors" positioning the machine as an independent source of intelligence. **The cognitive revolution and the roots of ITS.** Early pioneers — Herbert Simon, Allen Newell, Marvin Minsky, John Anderson — were not merely modeling cognition; they were investigating fundamental questions about learning and instruction. Their work produced information-processing models of human cognition that became foundational to educational psychology, and led to [[intelligent-tutoring|Intelligent Tutoring Systems (ITS)]], rooted in the heuristic-search and expert-systems paradigms of 1960s–70s AI. These systems aligned naturally with existing structures of assessment and standardization, and eventually coalesced into the AIED community. **Control vs. agency: Anderson and Papert.** Two competing visions emerged, embodied in John Anderson and Seymour Papert. **Anderson's** [[intelligent-tutoring|cognitive tutors]] (ACT/ACT-R theory) decomposed knowledge into production rules, provided precise individualized feedback, and reinforced traditional educational structures — a vision of AI as a tool for optimizing instruction. **Papert's** [[constructivist|constructionism]] (Logo, "microworlds," debugging-as-learning) saw computers as environments where children construct knowledge through creative experimentation — a vision of AI as an instrument of intellectual empowerment that challenged institutional hierarchies. This "essential tension" between control and agency is the through-line of the field's history. **Two forms of personalization.** [[personalized-learning|Personalization]] has two distinct interpretations that map onto these visions. The first maintains uniform outcomes while varying the path (rooted in Skinner's teaching machines; the Khan Academy vision of tutoring to mastery). The second embraces diverse outcomes — learners pursuing individual interests and talents, unconstrained by standard curriculum boundaries (Zhao, 2024). The ITS model supports the former; constructionism the latter. **Lessons for the GenAI era.** Today's debates about generative AI closely mirror these historical tensions. GenAI is often positioned as a next-generation intelligent tutor (extending Anderson), raising concerns about surveillance, data collection, and standardization. Alternatively it can be a tool for creative agency (extending Papert), requiring us to rethink assessment and accept more ambiguous outcomes. As [[mishra-control-vs-agency-history-2025|Mishra et al.]] conclude, institutional forces repeatedly favor approaches that reinforce existing structures — and the language and metaphors we choose will shape future possibilities. [[rismanchian-ai-education-four-decades-aixed-2026|Rismanchian & Doroudi's AI×Ed framework]] gives this intuition empirical grounding: locating papers from AIED proceedings (1985, 1993, 2021, 2024) and IJAIED (2004, 2014, 2021) along two axes — the role of AI (applied tool vs. analogy to human intelligence) and the end user — they trace a field that moved from a diverse mix, including substantial early research using AI as an analogy to human intelligence (computational models of learning, Papert's microworlds, agent-based models of knowledge acquisition), toward a near-exclusive focus on applied, instrumental uses by the 2000s. Notably, the generative-AI era appears to partially reverse that trajectory: at AIED 2024, three of the four papers in the field's "AI-as-analogy" quadrant were LLM-based, including best-paper work on LLM-based teachable agents. ## The recent impact of generative AI The arrival of [[generative-ai|generative AI]] — a broad family that includes [[llm|large language models]] (text), plus image, audio, and video generators — marks the most consequential inflection point in AIED history since the field's founding. LLMs are the most prominent member of this family and the driver of most classroom impact, but generative AI as a whole (capable of producing novel text, images, sound, and code) is the broader force reshaping education. It is best understood against AIED history rather than as a clean break. Several features make the generative-AI era distinctive: **Scale and reach.** Unlike the cognitive tutors and ITS of earlier decades, which operated in controlled lab and classroom settings, LLMs reached hundreds of millions of learners and educators within months of release, entering everyday writing, homework, and instruction almost immediately. The GenAI debate moved from specialist journals to classrooms, boardrooms, and legislatures with unprecedented speed. **A new kind of capability.** Earlier AIED systems modeled cognition and delivered structured instruction (the Anderson tradition). LLMs perform open-ended generation — producing essays, explanations, code, and dialogue — which foregrounds learner agency and creativity (the Papert tradition) in ways earlier ITS could not. This is why the current moment re-ignites the control-vs-agency tension so sharply: the same tool can be deployed as a next-generation intelligent tutor (controlling outcomes) or as a creative co-writer and thinking partner (amplifying agency). **A reversal of the access equation.** Historically, personalized tutoring was scarce and expensive; GenAI made adaptive one-on-one support nearly free and universal — dramatically lowering the accessibility and cost barriers that shaped earlier AIED. But it also introduced new risks: algorithmic bias, hallucination, surveillance, data collection, and the de-skilling of educators. **Renewed concern about AI literacy and misuse.** The generative-AI era made [[ai-literacy]] urgent and reframed [[academic-integrity]] debates (detection vs. dialog), while reviving long-standing concerns about automation, standardization, and epistemic dependence. These are the modern expressions of the field's original tensions. **Chronocentrism's risk.** As [[mishra-control-vs-agency-history-2025|Mishra et al.]] warn, the hype around generative AI often lacks historical awareness, letting tech evangelists dominate the conversation. Understanding that today's GenAI debates echo the Anderson–Papert divide and the 1955 cybernetics naming decision helps us make deliberate choices rather than treating the present as inevitable. The real question is not what generative-AI tools can do, but what we want them to do — a decision shaped by where we have been. ## Implications - **Read hype historically.** Recognize that each new AI wave is framed as revolutionary; historical awareness reveals continuity and helps resist corporate dominance of the conversation. - **See design choices as value choices.** Technical decisions about AI in education embed ideological positions about learning's purpose — surface and interrogate them. - **Weigh control against agency deliberately.** Whether AI should optimize toward standard outcomes or enable diverse, learner-directed outcomes is a genuine choice with institutional and political consequences, not a technical given. - **Preserve learner agency.** The recurring resistance to standardization and control in favor of open-ended, learner-centered approaches is a durable thread worth defending. ## Connected Concepts - [[ai-education]] - [[intelligent-tutoring]] - [[constructivist]] - [[agency]] - [[personalized-learning]] - [[generative-ai]] - [[learning-theories]] - [[adaptive-learning]] - [[ai-literacy]] - [[creativity]] - [[teacher-role]] - [[ethics]] ## Connected Articles - [[mishra-control-vs-agency-history-2025]] — Control vs. Agency: Exploring the History of AI in Education - [[prezenski-human-centered-ai-aided-learning]] — Human-centered AI-aided learning (engages Anderson's legacy) - [[cognitive-commons-ai-expertise-regeneration]] — Cognitive commons and AI expertise (historical framing) - [[programming-its]] — Programming Intelligent Tutoring Systems - [[lak2026-hint-button-unproductive-use]] — Revisiting the hint button in cognitive tutors - [[socrates-students-instructors-llms-lbt-2025]] — Learning-by-teaching with LLMs (Papert's constructionism lineage) - [[rismanchian-ai-education-four-decades-aixed-2026]] --- ## [Interpreting and Applying AIEd Research](https://edtechdev.github.io/aied/concepts/interpreting-and-applying-aied-research/) > **Interpreting and applying AIEd research** — how to decide whether an AI-in-education finding is worth acting on, whether you teach, run a program, design a course, or build software. You do not need statistics to use this page. It starts with the question a practitioner actually has — *should I do this?* — works through the few things that answer it, and keeps the technical detail in a later section for anyone who wants or needs it. The short version: a finding is worth acting on when you know what was compared, what was measured, who was studied, and whether the tool still exists in the form that was studied. Most claims that reach instructors, administrators and developers fail one of those four. ## Questions to Consider - A vendor, a news story, or a colleague says an AI tool improved learning. What is the one thing you would want to see before trying it in your own course — and would you know where to look for it? - Every article page in this knowledge base now has a **What this means for practice** section, and most have a **Limitations** section. Reading the two together, what does each tell you that the other does not? - Students practice with an AI tool and do better on the practice work, then do worse on the [[summative-assessment|closed-book exam]]. Which number is the learning outcome your course cares about — and would your current assessments catch the difference? - A study promises a big improvement, but it followed 30 students in one course at one institution, and the version of the tool studied is no longer the version anyone uses. Which of those two facts worries you more, and why? - Many AI tools increase how much students use them without increasing how much they learn. If you had to choose between a tool that raises engagement and one that raises unassisted performance, what evidence would settle it? - You are asked to approve or buy a tool on the strength of the vendor's own effectiveness numbers. What would you want disclosed about how those numbers were produced? ## Introduction This page is for people who have to decide something: an instructor wondering whether to change an assignment, an [[administrator]] weighing a pilot, a learning designer building a course, a software developer deciding what a feature should do, or a researcher explaining a finding to any of them. Two habits make the difference, and neither needs research training. **Read the two sections written for you first.** Every article page here now carries a **What this means for practice** section — usually three to five concrete actions derived from that study — and most carry a **Limitations** section stating what the study cannot support. Read those before the study's findings, not after. The practice section tells you what the study is good for; the limitations section tells you where it stops. If the practice section is missing or vague, treat that page as unfinished rather than as evidence. **Judge the claim, not the confidence of the claim.** Claims about [[ai-education|AI in education]] are usually accurate about *something* and misleading about the thing you care about, because a study and your classroom differ in four ways: what it was compared against, what it measured, who took part, and which version of the tool was used. The rest of this page gives you the checks in plain language, then the evidence behind each of them for readers who want it. What follows is also not a replacement for its neighbours. [[research-methods-aied|Research Methods in AI in Education]] covers how designs are built, [[limitations-in-aied-research|Limitations in AIEd Research]] catalogues the literature's recurring weaknesses, [[ai-ed-evaluation|AIED Evaluation]] covers how systems and outputs are evaluated, and [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]] covers who a finding does and does not include. ## Four questions that settle most claims Ask these before you spend time, money, or a semester on something. **1. What was it compared against, and was that comparison fair?** "Students who used the AI did better than students who didn't" only tells you something if the other students were doing something real. If the comparison was business as usual — or nothing — then the finding bundles the tool together with extra time, extra attention and novelty. What to look for: a **control group** that got a credible alternative, and random assignment to the two conditions. **2. What exactly did they measure?** This is where most exciting claims quietly fail. Test scores, homework quality, [[motivation|motivation]], attitudes and [[student-engagement|engagement]] get pooled into a single "achievement" number, or a measure of performance *with the tool present* gets reported as learning. Learning that depends on the tool being there is not the same as learning that lasts. What to look for: what the instrument measured, whether it was validated for that population, and whether an outcome was measured **without** the AI in the room. **3. Who was studied, how many, and for how long?** Thirty students in one course is a signal, not a result. A four-week intervention cannot tell you about a year. And a study of students unlike yours is still useful — it is a hypothesis about your setting, not a prediction. What to look for: sample size, how participants were recruited, single site, duration, and whether any subgroup was large enough to analyze. **4. Is the tool still the tool that was studied?** AI capability moves faster than publication. A 2025 finding describes the 2025 model generation — sometimes a specific version, sometimes a configuration nobody uses now. That does not make the finding false; it makes it dated, and it means the claim should be re-checked rather than inherited. What to look for: model version and the data-collection window. ## What the evidence says about AI claims in general If you remember one thing from this page, remember that the headline number is usually inflated and the comparison is usually weak. That is not a fringe view — it is what audits of the field itself report. Several well-designed studies do show real gains; the point is that the burden of proof sits with the claim. - **About two-thirds of the average effect disappears** once you correct for the fact that impressive results get published and unimpressive ones do not. [[bartos-ai-learning-meta-meta-analysis-2026|Bartoš et al. (2026)]] pooled 1,840 effect sizes from 67 reviews and found the corrected average was roughly one-third of the published median — SMD 0.196 against 0.67. - **A product name is not a teaching method.** Auditing the comparisons behind a prominent meta-analysis, [[weidlich-chatgpt-effect-search-cause-2025|Weidlich et al. (2025)]] found only **21%** had a well-defined treatment, a control group and a valid learning measure — and the reported advantage for "using ChatGPT" came out larger than for purpose-built [[intelligent-tutoring|intelligent tutoring systems]] (g = 0.7 against 0.66), which is a warning sign rather than a triumph. - **Performance with the tool is routinely mistaken for learning.** In one [[k-12|K-12]] math study, students practicing with a general-purpose [[conversational-ai|chatbot]] earned better practice grades and then scored **about 17% worse** than peers with no AI access on the closed-book final ([[stanford-evidence-base-ai-k12-2026|Stanford's evidence base for AI in K-12]]). - **The field's own reviews do not survive audit.** [[oneill-presumed-effective-meta-analysis-2026|O'Neill's (2026)]] audit of **14 peer-reviewed meta-analyses** claiming AI improves education found that **none** provided a valid basis for the claims it advanced, and among 46 randomly selected primary studies, **61%** presented validity concerns. The problem is not one bad paper; it is a reporting culture. - **Self-reports flatter everyone.** People rate their own [[ai-literacy|AI skills]] about **40%** higher than performance measures show, which is why satisfaction and confidence surveys are the weakest evidence you can act on ([[self-report-measures|Self-Report Measures]], [[educational-measurement|Educational Measurement]]). ## Turning a finding into a decision The sequence that saves the most wasted effort, in order. 1. **Write down your outcome first.** Not "use AI more" but "students can do X without the tool". If your outcome is unassisted performance, then a study that measured assisted performance is adjacent evidence, not direct evidence. 2. **Find the comparison and the measure** — in the study, or in the practice section of its article page. If either is missing, treat the claim as a demo rather than a finding. 3. **Read the limitations section as instructions, not disclaimers.** "Single course, self-reported outcomes, four weeks" tells you exactly which of your assumptions the study does not cover. 4. **Check the version and the date.** If the study used a model generation two years old, plan to re-test rather than assume. 5. **Name the enabling conditions.** Cost, licenses, staff time, data rules, and whether students must pay for the tier that actually works. Studies rarely carry these, and they decide whether an intervention survives a semester. In [[chick-faculty-development-ethical-ai-2026|a faculty-development study of ten participants]], everyone said they would keep using AI, while the same people described personal subscriptions for tool access and no time support for the redesigns they had planned. 6. **Pilot small, and measure the unsupported condition.** A short pre/post with one assessment done without the tool beats a satisfaction survey. Small and honest beats big and rhetorical. See [[learning-design|Learning Design]] for where this fits in course design. 7. **Diary a review date, and be willing to drop the claim.** When a study's premise is a capability that no longer exists, the honest move is retiring the claim rather than citing it indefinitely — the same discipline this knowledge base applies to its own pages. ## Words you will meet in the research Plain translations, so you can skim a study or a vendor page without a methods background. - **Effect size** — how big the difference was, on a scale where 0 is nothing. Treat small values as "a nudge", not "a transformation". - **Statistically significant** — unlikely to be pure chance *in this sample*. It says nothing about whether the effect is large, or whether it will happen in your class. - **Confidence interval** — the range of results the data cannot rule out. If the range includes zero, the finding may be nothing at all, however interesting the headline. - **[[meta-analysis-systematic-review|Meta-analysis]]** — a study that pools many studies. Powerful, and only as good as what it pooled, which is why reviews get audited. - **Publication bias** — interesting results get published and boring ones do not, so the literature's average looks rosier than reality. - **Self-report** — people describing themselves. Useful for attitudes, weak for competence or behavior. - **Control group** — the comparison condition. The single most important thing to look for. - **Pre/post** — measured before and after with no comparison group. Suggestive, never conclusive. - **Subgroup analysis** — results for a slice of the sample. Usually underpowered, so treat it as a hypothesis. - **Replication** — someone else got the same result. Rare, and the strongest evidence available. - **[[benchmark|Benchmark]]** — a fixed task set for scoring systems. Scores move when the target moves, so check the date. ## When to slow down anyway - **The claim comes from the vendor, on the vendor's metrics.** That can still be informative — one [[intelligent-tutoring|AI tutoring]] provider reports an engagement metric calibrated against human experts at F1 0.83, with improvements coming from over 40 experiments in five months ([[ai-tutoring-quality-k12-methodologies-2026|Udeshi et al., 2026]]) — but the construct, the raters and the metric are the vendor's choices. Ask for the comparison group and the unassisted outcome. - **Automated scoring is treated as solved.** High agreement with human raters is reliability, not quality. In one scoring study, human raters agreed with the multi-rater consensus at about r = 0.88, so automated scores near r = 0.85 were already at the task's own measurement ceiling ([[know-when-to-trust-ai-scoring-reliability-2026|Know When to Trust AI Scoring]]). - **The reference list is doing heavy lifting.** Thirty reference entries containing verifiably fabricated bibliographic information were confirmed across 14 [[cs-education|computing-education]] papers, all from 2025 and 2026 ([[citation-errors-hallucinations-computing-education-2026|Denny et al., 2026]]). If a claim rests on a citation, check the citation. - **Nobody measured the behavior your policy depends on.** Across 493 deduplicated records and 14 priority studies, no study measured whether verification succeeded *and* what the learner then did with it, judged against an independent standard of output quality ([[verification-quality-reliance-calibration-genai-2026|verification and reliance calibration]]). Course policies depend on exactly that behavior. ## If you are building or buying a tool The same checks invert into design requirements, and the knowledge base's evaluation pages carry the detail ([[ai-ed-evaluation|AIED Evaluation]], [[automated-assessment|Automated Assessment]]). - **Make the comparison part of the feature spec.** Decide what a learner would otherwise be doing, and be able to say why your tool beats that — not why it beats nothing. - **Measure the unsupported condition.** If your outcome is learning, include a task completed without the tool; assisted performance alone will mislead you as much as it misleads your buyers. - **Report how your automated judgments were validated** — the gold standard, the calibration target, who adjudicated disagreements — and report it as reliability rather than quality. - **Name the version and the date** in any effectiveness claim, because your next release invalidates it. - **Show the counterweights:** cost per student, [[accessibility]], data handling, and what happens to learners on the free tier. See [[ai-use-disclosure|AI Use Disclosure]], [[privacy]] and [[governance]]. ## For readers who want the evidence The checks above are not folk wisdom; they come from documented failures in this literature. This section keeps the detail for anyone reviewing a paper, defending a decision, or arguing that a tool should be evaluated properly. **The validity failures have a distribution, not just a presence.** In [[oneill-presumed-effective-meta-analysis-2026|O'Neill's (2026)]] 46 audited studies, dependent-variable mismatch was the most common problem (n = 15) — the measure did not capture what the claim asserted — followed by independent-variable mismatch (n = 11), experimental design problems (n = 7), data extraction problems (n = 6), absence of a control group (n = 6) and nonrandom group assignment (n = 6). Of the 14 meta-analyses, twelve treated multiple effect sizes drawn from the same primary study as independent, which inflates the apparent evidence base. **Subgroup claims are usually undecidable in the studies that make them.** The [[ai-tutoring-micro-rct-gcse-science-2026|GCSE science micro-RCT]] reports a treatment-by-status interaction of **0.57 marks (95% CI -2.25 to 3.39)**, with stratified estimates of **g = 0.28 (95% CI -0.04 to 0.59)** for one group and **g = 0.35 (95% CI 0.18 to 0.52)** for the other. An interval crossing zero is not an [[equity-in-ai-education|equity]] finding; it is a question for a local pilot. **A synthesis can satisfy its own protocol and still pool unweighted quality.** [[ai-supported-instruction-stem-meta-analysis-2026|Doğan et al. (2026)]] state plainly that they used no formal quality appraisal tool, treating their inclusion criteria as the rigor threshold, so a quasi-experimental study and a randomized one contributed equally — and their heterogeneity reads **I² = 82.98% under a fixed-effect model but 15.75% under the random-effects model**, which is why a heterogeneity figure quoted without its model cannot tell you how inconsistent the corpus is. **[[assessment-validity|Validation]] of automated judgment is part of the result.** Agreement with human coders is a reliability statement, and the ceiling above shows why it is not the same as quality ([[machines-misread-pedagogical-quality|machines misread pedagogical quality]]). **Tool vintage is a first-class limitation.** A review of AI-assisted assessment notes that its own findings reflect specific model versions at specific times, and that field movement makes any account of model capabilities potentially outdated within months ([[ai-assisted-assessment-instruction-higher-ed-2026|AI-Assisted Assessment and Instruction in Higher Education]]). Scope claims to their generation: "[[generative-ai|generative AI]] improved X" is not portable, while "GPT-4-era tooling, in this task, with this [[scaffolding]]" is. **Benchmark targets move**, so a result that a system saturates or fails today can invert with the next release; saturation and contamination checks belong beside any benchmark-based claim. **Reading this literature alongside its own critics is normal practice here.** The reporting-side checklists for authors and reviewers are in [[reporting-interpreting-aied-research|the FAQ on reporting and interpreting AI research]], and the appraisal habits in this page pair with [[theory-development-aied|Theory Development in AI in Education]] when a claim is theoretical rather than empirical. ## A short checklist 1. State your outcome in one sentence, including whether the tool is present in it. 2. Find the comparison. No credible comparison, no decision. 3. Match the measure to your claim, and prefer an unassisted outcome. 4. Deflate the number: read the bias-corrected effect, not the headline. 5. Check the model version and the study's dates. 6. Read the limitations section as instructions for what you still do not know. 7. Cost it: licenses, tiers, staff time, data rules. 8. Pilot small with an unassisted measure, then decide. 9. Diary a review date a year out, and be ready to retire the claim. ## Connected Concepts - [[limitations-in-aied-research]] - [[learning-design]] - [[ai-ed-evaluation]] - [[educational-measurement]] - [[assessment-validity]] - [[self-report-measures]] - [[learning-gains]] - [[differential-effects-across-learner-groups]] - [[research-methods-aied]] - [[meta-analysis-systematic-review]] - [[quantitative-research]] - [[benchmark]] - [[rct]] - [[intelligent-tutoring]] - [[cognitive-offloading]] - [[ai-use-disclosure]] - [[theory-development-aied]] ## Connected Articles - [[oneill-presumed-effective-meta-analysis-2026]] — Presumed Effective: forensic audit of 14 AIED meta-analyses - [[bartos-ai-learning-meta-meta-analysis-2026]] — Publication-bias-adjusted AI effects about one-third of reported size - [[weidlich-chatgpt-effect-search-cause-2025]] — ChatGPT in Education: An Effect in Search of a Cause - [[ai-supported-instruction-stem-meta-analysis-2026]] — Inclusion criteria used as the rigor threshold, and heterogeneity that changes with the model - [[know-when-to-trust-ai-scoring-reliability-2026]] — When automated scoring reliability meets the task's measurement ceiling - [[verification-quality-reliance-calibration-genai-2026]] — What the verification and reliance literature does not measure - [[citation-errors-hallucinations-computing-education-2026]] — Fabricated references that reached print in 2025–2026 - [[stanford-evidence-base-ai-k12-2026]] — Practice gains, exam losses: the assistance-removal problem in K-12 math - [[ai-tutoring-micro-rct-gcse-science-2026]] — Subgroup effects whose confidence intervals cross zero - [[chick-faculty-development-ethical-ai-2026]] — Enabling conditions: policy signals, personal subscriptions, no time - [[ai-tutoring-quality-k12-methodologies-2026]] — Vendor metrics with their calibration and experiment count disclosed - [[ai-assisted-assessment-instruction-higher-ed-2026]] — Findings tied to model versions, and the field's churn --- ## [Limitations in AIEd Research](https://edtechdev.github.io/aied/concepts/limitations-in-aied-research/) > **Limitations in AIEd research** — the recurring weaknesses and constraints that affect how much confidence we can place in AI-in-education findings, and how readers should interpret them. These cut across individual studies: methodological limitations (generalizability, sample size, validity, self-report), the speed problem (AI and findings date quickly while publication lags), research-practice limitations (reproducibility, FAIR practices, proprietary tools), and weak theory use. Recognizing these limits is essential for reading the literature critically and for designing stronger studies. ## Questions to Consider - How much would you trust a headline like 'AI tutoring boosts learning by 30%' if you learned it came from 30 students in one course at one institution? The page flags generalizability and small samples as recurring limits — what would you want to know before acting on any single finding? - A striking limitation is the 'speed problem': AI evolves faster than findings get published, so a study of one model generation may already describe an obsolete system. How should this change the confidence you place in AI education research? - Many studies rely on [[self-report-measures|self-reported]] attitudes and usage, which are biased — people overestimate their competence and under-report misuse. Have you ever answered a survey about your own skills or behavior in a way that didn't match reality? Why do perception-based measures so often diverge from objective performance? - The page notes that familiar frameworks like Bloom's taxonomy are often misread as strict ladders, and that even widely used theories like [[cognitive-offloading|cognitive load]] theory have been challenged. When have you seen a theory invoked as settled truth in a context where its own evidence was actually contested? - Most AI research depends on proprietary, opaque models whose data and updates you cannot inspect. If you cannot verify exactly what model produced a result, how much can you trust claims built on it — and what would make findings more reproducible? - If you are an instructor or designer without time to read primary research, how do you decide which AI claims are trustworthy enough to change your practice — given that the literature is fragmented, provisional, and written for researchers? ## Introduction AI in education is a fast-moving, heterogeneous field, and its evidence base carries a distinctive set of limitations that researchers, practitioners, and policy-makers should weigh when using any finding. Some of these are shared with the broader learning-sciences and psychology literature; others are amplified or made unique by the nature of AI itself. This page organizes them into four cross-cutting areas. ## Methodological limitations The knowledge base's [[research-methods-aied|research methods]] page details the strengths and limitations of each design. Several limits recur across designs and deserve particular attention: - **Generalizability.** Findings from a single course, institution, discipline, or national context may not transfer. Small, convenience, or single-institution samples limit external validity; results from one AI tool rarely extend to a different tool or context. - **Synthesis-level rigor is a separate axis from primary-study rigor.** A meta-analysis can satisfy its own inclusion criteria and still pool studies that differ in design, implementation fidelity and outcome measure without weighting any of that: [[ai-supported-instruction-stem-meta-analysis-2026|Doğan and colleagues (2026)]] state plainly that they used no formal quality appraisal tool and treated the inclusion criteria as the rigor threshold, so a quasi-experimental study and a randomized one contributed equally to the pooled [[stem-education|STEM]] estimate. The same review shows a related reporting hazard: its heterogeneity is quoted as I² = 82.98% under a fixed-effect model and I² = 15.75% under the random-effects model, meaning readers who lift a single heterogeneity figure without its model cannot tell how inconsistent the corpus actually is. Appraise a synthesis on how it handled dependent effect sizes, quality, and heterogeneity, not only on whether it followed a search protocol. - **Small sample sizes.** Many AIED studies are underpowered — too few participants to reliably detect meaningful effects or to support the strong claims sometimes drawn from them. - **Validity and measurement.** [[assessment-validity|Construct validity]] is often thin: proxies for "learning," "[[student-engagement|engagement]]," or "literacy" vary widely, and instruments are not always validated for the population or construct being studied. [[benchmark|Benchmark]] accuracy does not equal educational effectiveness. - **Self-report and survey data.** A large share of the corpus relies on self-reported attitudes, motivation, and usage. Self-report is subject to bias — respondents overestimate competence, under-report [[ai-misuse-learning-harm|misuse]], and misjudge their own behavior — so perception-based measures frequently diverge from objective performance (see [[ai-literacy-assessment-misalignment]] and [[educational-measurement]]). ## The speed problem: AI evolves faster than findings AI is changing continuously, and the conclusions drawn from any given model or system can become **out of date quickly**. A study of one [[llm]] generation may not describe the next; benchmark scores, tutoring quality, and even the practical usefulness of a finding shift as models improve. Compounding this, the **publication process is slow** — from study design to peer-reviewed publication can take a year or more — so a published result may already describe an obsolete system. Reviewers and readers should therefore treat AIED findings as provisional, date-sensitive claims rather than stable truths, and prefer recent, replication-oriented, and version-explicit work. ## Research-practice limitations Several limitations concern the conduct and infrastructure of the research itself: - **Lack of reproducibility.** Studies often do not report enough detail (prompts, model versions, hyperparameters, data, analysis code) for others to reproduce or verify results — a particular problem given how sensitive LLM output is to prompts and settings. - **FAIR research practices.** Open and reproducible practice — **F**indable, **A**ccessible, **I**nteroperable, **R**eusable data and code, pre-registration, and shared benchmarks — is unevenly adopted in AIED. Weak adherence to FAIR principles makes it harder to reuse, compare, and build on studies. - **Proprietary tools and models.** Much research depends on closed, proprietary [[ai-technologies|AI systems]] whose internal behavior, training data, and model updates are opaque and may change without notice. This limits reproducibility, makes exact replication impossible, and can tie findings to a vendor's roadmap. It also raises questions about evaluation independence (see [[ai-ed-evaluation]]). ## Weak or limited theory use A recurring criticism is that many empirical articles have **limited or outdated theoretical framing**. Researchers may: - **Adopt theories uncritically.** Frameworks are borrowed because they are familiar, without fully engaging their assumptions, scope, or evidence base. - **Misinterpret frameworks as fixed sequences.** Several widely used frameworks are treated as ordered ladders that learners must climb from a "low" to a "high" stage — but the evidence does not support always starting at the bottom. For example: - **Bloom's taxonomy** is often read as a strict hierarchy (recall → application → evaluation), yet higher-order goals do not require first drilling lower-order ones; tasks can be designed to engage evaluation or creation from the start (see [[cross-dataset-bloom-question-classification]]). - **ADDIE** and other instructional-design models are sometimes treated as rigid linear phases rather than the iterative, flexible planning heuristics they are meant to be (see [[learning-design]]). - **Overlook contested theories.** Some theories used widely in AIED have themselves been challenged. **Cognitive load theory**, for example, has been criticized and its empirical claims refuted or disputed in prior studies, yet it continues to be invoked as a settled foundation in new AIED work. The implication is not that theories and frameworks are useless, but that they should be used with attention to their actual evidence base, their intended scope, and their known criticisms — rather than as self-evident [[scaffolding|scaffolds]] or rigid procedural sequences. ## The meta-analytic evidence crisis A growing body of meta-research — reviews that scrutinize the reviews — argues that the field's headline claims of AI-driven learning gains rest on an evidence base that is far weaker than it appears. Three complementary critiques make the case with unusual force: - **Positive-synthesis bias is severe and quantifiable.** [[bartos-ai-learning-meta-meta-analysis-2026|Bartoš et al. (2026)]], in a study-level meta-meta-analysis of 1,840 effect sizes from 67 meta-analyses, found strong evidence of publication bias (all Egger tests *p* < .0001) and extreme between-study heterogeneity (τ = 0.869). Publication-bias-adjusted effects were roughly **one-third** the magnitude commonly reported (SMD = 0.196 vs. a median of 0.67 in the published literature), with a prediction interval spanning −1.521 to +1.908 — from large harm to large benefit. No outcome, field, level, or AI-role subgroup showed consistent gains, and there was no difference between pre- and post-2023 studies. Their verdict: broad claims of generalized learning gains are premature. - **Meta-analytic methods are being systematically misapplied.** [[oneill-presumed-effective-meta-analysis-2026|O'Neill (2026)]]'s forensic audit of 14 high-impact AIED meta-analyses found that *none* provided a valid basis for their claims: none had a coherent construct (treating the tool "ChatGPT" as if it were a single intervention, and pooling test scores, motivation, [[self-efficacy]], and attitudes into one "[[learning-gains|academic achievement]]" number); reported heterogeneity was severe, with I² ranging from 77.2% to 94.4% across the 13 meta-analyses that reported it and 12 of those 13 above 80%, and it was never resolved (none met the minimum subgroup size of ten studies, five relied on single-study subgroups, and three others on subgroups of two); twelve treated dependent effect sizes from the same study as independent, inflating apparent evidence; and none validly assessed publication bias (discredited fail-safe N metrics and misapplied Egger tests were common). Because I² is precision-dependent it cannot by itself establish how far apart the true effects are, and the reporting that would show it was largely absent: only four meta-analyses reported between-study variance (τ²) and only two reported a prediction interval, both of which included zero. A majority (61%) of randomly vetted primary studies were problematic, and one study with fabricated references was included by six of the 14 meta-analyses. - **The "treatment" is a black box.** [[weidlich-chatgpt-effect-search-cause-2025|Weidlich et al. (2025)]] revive the media/methods debate to argue that "ChatGPT" is a tool, not a method — asking whether it "improves learning" is a non sequitur. Auditing a subset of the studies behind Deng et al.'s (2025) meta-analysis, they found only 21% of comparisons had a well-defined treatment, control group, *and* valid learning measure; reported effect sizes (g = 0.7) even exceeded those of purpose-built [[intelligent-tutoring|Intelligent Tutoring Systems]] (0.66), a red flag that the "treatment" was a heterogeneous "secret sauce." The convergence of these three independent critiques is itself evidence: across different methods, corpora, and framings, they reach the same conclusion — that positive AIED effect sizes, especially from early meta-analyses, likely reflect **publication bias, construct incoherence, and methodological shortcuts** as much as (or more than) genuine learning gains. This does not mean AI tools have no educational value; rather, it means the *field-level* evidence for their value is currently inflated and must be read accordingly. It also shifts responsibility to [[meta-analysis-systematic-review|synthesis quality]]: a meta-analysis is only as trustworthy as the coherence of its constructs, the independence of its effect sizes, the adequacy of its moderator and heterogeneity analysis, and the validity of its publication-bias assessment — each of which the critiques show is routinely violated. ## Reading the AIED literature critically Taken together, these limitations argue for a critical, multi-signal reading of AIED research: check whether a finding generalizes and is adequately powered; verify how constructs were measured (and whether claims rest on self-report); prefer recent, version-explicit, reproducible work; and interrogate the theoretical framing rather than treating familiar frameworks as given. This is the complement of rigorous [[research-methods-aied|method choice]] and [[ai-ed-evaluation|evaluation]]: good methods and good evaluation are necessary, but reading with attention to limitations is what turns evidence into defensible decisions. ## From research to practice A further, practical limitation is the **challenge of applying research to [[teacher-role|teaching]] and instructional design**. Practitioners — instructors, [[stakeholders|instructional designers]], and faculty developers — often lack the time or specialized expertise to read, appraise, and translate primary research into concrete classroom decisions. The literature is large, fragmented, and written for researchers; findings are reported with statistical and methodological detail that is not immediately actionable; and because claims are provisional (see the speed problem above), a practitioner cannot simply take a single study at face value. This creates a gap between what the evidence supports and what actually reaches [[pedagogy|teaching practice]]. The purpose of this knowledge base is to help close that gap — to make it easier to keep up with, interpret, and apply AI-in-education research to practice — by curating open-access findings into structured, accessible summaries, connecting related work through [[ai-education|concept pages]], and flagging the limitations readers should weigh. It aims to support evidence-informed practice in teaching and instructional design, and in doing so to also surface gaps and questions that can inform new research and development. Understanding the limits of the research is therefore not an end in itself: it is what lets practitioners apply findings appropriately and lets researchers design stronger studies that better serve practice. ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[research-methods-aied]] - [[ai-ed-evaluation]] - [[educational-measurement]] - [[assessment-validity]] - [[benchmark]] - [[rct]] - [[meta-analysis-systematic-review]] - [[ai-education]] - [[icap-framework]] - [[learning-design]] - [[llm]] - [[generative-ai]] - [[cognitive-offloading]] - [[theory-development-aied]] — Theory Development in AI in Education ## Connected Articles - [[ground-truth-reliability-aied]] — Reliability and validity of ground truth in evaluation - [[ai-literacy-assessment-misalignment]] — Self-reported vs. performance-based AI literacy - [[machines-misread-pedagogical-quality]] — Why machines misread pedagogical quality - [[favero-critical-ai-tutors-empower-enslave-2025]] — Critical limits of AI tutors and theory use - [[cross-dataset-bloom-question-classification]] — Bloom's taxonomy and question classification - [[eeg-familiarity-automated-assessment-2026]] — Automating Learner Assessment: EEG-Based Familiarity Prediction - [[weidlich-chatgpt-effect-search-cause-2025]] — ChatGPT in Education: An Effect in Search of a Cause (media-comparison critique) - [[bartos-ai-learning-meta-meta-analysis-2026]] — Meta-meta-analysis: publication-bias-adjusted AI effects ~1/3 of reported size - [[oneill-presumed-effective-meta-analysis-2026]] — Presumed Effective: forensic audit of 14 AIED meta-analyses - [[prisma-llm-ai-assisted-systematic-reviews-2026]] — PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews - [[frontier-models-physics-benchmark-audit-2026]] — How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks - [[ai-supported-instruction-stem-meta-analysis-2026]] — Inclusion criteria used as the rigor threshold, and a heterogeneity figure that changes with the model (Doğan et al. 2026) - [[domain-specific-chatbot-stem-enthusiasm-2025]] — A cluster-randomized classroom trial whose performance outcome did not reach significance (Rücker & Becker-Genschow 2025) --- ## [Philosophy of AI in Education](https://edtechdev.github.io/aied/concepts/philosophy-of-ai-in-education/) > **Philosophy of AI in Education** — the branch of educational philosophy that examines the fundamental conceptual questions raised by artificial intelligence in [[teacher-role|teaching]] and learning: What is the nature of knowledge and thinking when machines participate in them? What is the learner when cognition is distributed across human and artificial systems? What forms of [[agency]], responsibility, and personhood apply to AI, and what does education owe learners in an AI-mediated world? Distinct from (but connected to) the knowledge base's [[learning-theories]] page, which catalogs theories of *how learning happens*, the philosophy of AI in education asks the deeper questions of *what learning, mind, and the learner fundamentally are* under AI-mediated conditions. ## Questions to Consider - Does thinking require consciousness and a body? If an AI genuinely participates in your reasoning, is some of the 'thinking' happening in the machine—and does that change who's learning? - The page raises the idea of cognition 'distributed' across human and artificial systems. Think of a task you've solved with AI help: where did your thinking end and the machine's begin? - If a learner's cognitive processes become genuinely hybrid (part human, part artificial), what does that mean for what we call 'the learner'—and for what education owes them? - The page distinguishes philosophy (what learning and the learner fundamentally are) from learning theories (how learning happens). Can you name a belief you hold about learning that is really a philosophical position? - If AI can exercise 'functional agency' without consciousness, does it bear responsibility for its educational effects—and if not, who does? Who should answer when an [[intelligent-tutoring|AI tutor]] misleads a student? - Posthumanist thought reframes learners as 'post-human' entities. How does that idea challenge or unsettle your own assumptions about where a student's mind is located? ## Introduction This is a concept page for the philosophical and theoretical foundations of [[ai-education|AI in education]]. While [[learning-theories]] documents the empirical and design-oriented theories ([[behaviorism]], [[constructivist|constructivism]], [[cognitive-offloading|cognitive load]], [[self-regulated-learning|self-regulated learning]], etc.), the philosophy strand engages the ontological, epistemological, and ethical questions those theories presuppose. The two are closely connected: philosophical positions shape which learning theories seem plausible and which educational goals are worth pursuing. ### Key philosophical questions - **The nature of mind and cognition.** Does thinking require consciousness and a body? Can AI participate in genuinely cognitive processes? Frameworks such as **ensemble cognition** argue that AI exercises *functional agency* — genuine causal efficacy in cognitive processes — without consciousness, and that thinking emerges from dynamic human–[[student-ai-interaction|AI interaction]].([[ensemble-cognition-philosophy-ai-education]]) [[distributed-cognition|Distributed cognition]] and the extended mind thesis make related claims about cognition being spread across systems. - **What is the learner?** Posthumanist philosophy reconceptualizes the learner as a "post-human" entity whose cognitive processes are genuinely hybrid and distributed across [[biology-education|biological]] and artificial systems.([[elsayed-pedagogical-symbiosis-posthuman-learner]]) This challenges the assumption that the learner is a bounded, autonomous individual mind. - **Embodiment and the limits of disembodied AI.** [[embodied-learning|Embodied]] and post-cognitivist philosophy critiques the dominance of symbolic, disembodied AI models, arguing that cognition is grounded in situationality, emergence, and sensorimotor coupling that current [[generative-ai|generative AI]] lacks.([[videla-embodied-ai-education-choreography]]) - **Agency, authorship, and meaning.** When AI mediates interpretation and meaning-making, philosophy asks how authorship, epistemic [[agency]], and interpretive autonomy are reconfigured.([[voicu-ai-interpretive-cognition-ssh-2026]]) - **Values, justice, and the purpose of education.** Philosophical analysis examines whether AI-driven education serves human flourishing and educational justice, or whether it instrumentalises learning in service of productivity.([[avraamidou-ai-colonization-science-education]]) This connects to [[critical-pedagogy]] and [[ethics]]. - **Whose philosophy? Pluralism in the field's conceptual foundations.** Xie (2026) argues that the field's philosophical debate is over-determined by a single Western architecture — [[agency|epistemic agency]] as a property of discrete subjects, knowledge framed representationally and calculatively, and the human–AI relation located within subject–object dualism — so that its limits become the limits of the field's collective imagination. As a comparative counterweight he reconstructs Daoist concepts — "Dao nature" (道性), self-cultivation (修道) and the "Zhenren" (真人) — not as "Eastern content" added to an unchanged frame but as resources that reshape the conceptual foundations through which AI itself is understood, applying them to knowledge (epistemic monoculture and synthetic misinformation), knowing (offloading that degrades [[critical-thinking|critical thought]]) and impact (environmental costs and Global North–South asymmetries).([[daoism-ai-education-philosophy-2026]]) ### Relationship to learning theories The philosophy of AI in education and [[learning-theories]] are complementary lenses. Learning theories explain the mechanisms of learning (e.g., how [[feedback]], [[scaffolding]], or cognitive load shape outcomes); philosophy interrogates the presuppositions of those mechanisms — what counts as knowledge, who counts as a knower, and what the learner fundamentally is. Posthumanist and critical-philosophical work, in particular, challenges the field to move beyond instrumentalist frameworks like [[tpack]] and SAM toward deeper ontological reorientation.([[elsayed-pedagogical-symbiosis-posthuman-learner]]) ### Relationship to theory development Philosophy of AI in education and [[theory-development-aied|theory development in AIEd]] are complementary but distinct. Philosophy asks the ontological and epistemological questions — what mind, knowledge, and the learner fundamentally are under AI-mediated conditions — while theory development produces and empirically tests the *mechanisms* that operationalize answers to those questions (generativism, epistemic co-agency, the absent cognitive baseline). The two are mutually informing: philosophy clarifies the presuppositions that theories carry (e.g., [[learning-with-machines-toward-a-theory-of-epistemic-co-agency|epistemic co-agency]] presumes a distributed, non-individualist model of cognition), while theory development gives philosophical positions testable, falsifiable form. The field's most foundational articles do both at once, sitting at the boundary between the two concepts. ## Connected Concepts - [[learning-theories]] - [[theory-development-aied]] - [[distributed-cognition]] - [[ethics]] - [[agency]] - [[embodied-learning]] - [[critical-pedagogy]] - [[human-ai-collaboration]] - [[critical-thinking]] - [[ai-education]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation ## Connected Articles - [[genai-chinese-higher-education-integrity-2026]] — Gen-AI in Chinese higher education: integrity and engagement - [[ensemble-cognition-philosophy-ai-education]] — Ensemble Cognition: a philosophical framework for human–AI cognition - [[elsayed-pedagogical-symbiosis-posthuman-learner]] — Pedagogical Symbiosis and the Post-Human Learner - [[videla-embodied-ai-education-choreography]] — Embodied, post-cognitivist critique of disembodied AI in education - [[voicu-ai-interpretive-cognition-ssh-2026]] — Developmental-critical model of interpretive cognition in the humanities - [[learning-with-machines-toward-a-theory-of-epistemic-co-agency]] — Epistemic co-agency as a philosophy of learning with machines - [[avraamidou-ai-colonization-science-education]] — Critical-feminist philosophy questioning the AI colonization of education - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a Mediational Agent - [[young-people-learning-generative-ai-rapid-review-2026]] — Ecological learning-sciences framing of GenAI - [[philosophy-experimentation-ai-chemistry-2026]] — Philosophy of experimentation in chemistry with AI - [[strydom-human-gai-paradigms-2026]] — Framing human-AI dynamics: seven GAI engagement paradigms (Strydom 2026) - [[daoism-ai-education-philosophy-2026]] — Alternative AI Philosophy: Daoism as Method for AI in Education --- ## [Theories and Frameworks](https://edtechdev.github.io/aied/concepts/theories-and-frameworks/) > **Theories and frameworks** — the working inventory of explanatory and organizational structures this knowledge base actually uses. Theories explain why learning happens ([[learning-theories]], [[self-determination-theory]], [[sociocultural-learning]], [[activity-theory-aied|activity theory]]); frameworks organize design, teaching, and adoption decisions ([[tpack]], [[samr-model|SAMR]], [[technology-acceptance-model|technology adoption models]], [[icap-framework|ICAP]], [[universal-design-for-learning|universal design for learning]]); models make learning claims testable and computable ([[item-response-theory|item response theory]], [[assessment-validity]], [[knowledge-tracing]]). Use this page to find the lens a study rests on, or to pick one for your own work. ## Questions to Consider - A theory predicts what happens when something changes; a framework tells you what to attend to. When a study reports a gain under a framework's banner but makes no prediction that could have failed, what has actually been learned? - [[tpack|TPACK]], [[samr-model|SAMR]], [[universal-design-for-learning|UDL]], the [[technology-acceptance-model|adoption models]] and [[icap-framework|ICAP]] all predate [[llm|large language models]]. Which parts of them still hold when the tool can write, explain, and adapt on its own, and how would you tell? - Several of the theories catalogued below were coined for AI within the last few years. What would it take for one of them to be wrong, and has anyone tried to show that? - Adoption is often reported as movement through stages. If a colleague reports that their course moved "from substitution to redefinition" this year, what independent evidence would you ask for before believing it? - Much of this field applies borrowed theory rather than testing it. Pick a finding you rely on: could it be reframed as a test of [[self-determination-theory|a theory]], and what measurement would that require? ## Introduction Every AI tutor, feedback system, and dashboard embeds assumptions about how people learn, whether or not the designers state them. This page collects the theories, frameworks, and models those assumptions come from, ordered loosely by how much of the current research literature leans on them, and says briefly what each one is good for. The distinctions among the three labels are real but minor, and they are summarized near the end rather than treated as the subject. What matters practically is which lens a claim rests on, and whether that lens can support the claim. ## Learning theories the literature applies These explain mechanisms: they name the parts that do the work and predict what follows when those parts change. In the AI literature they are mostly borrowed from psychology and the [[learning-sciences|learning sciences]], and used as interpretive frames around a new tool. - **[[learning-theories|Learning theories (umbrella)]]** — the family as a whole, including the classical poles of [[behaviorism|behaviorism]] and [[constructivist|constructivism]] and the "constructivism in name, behaviorism in practice" gap that recurs in AI implementations. Connectivism and cognitive load theory are covered here too. - **[[constructivist|Constructivism]]** — learning as active knowledge construction. It underwrites project-based work, dialogue-based AI tutoring, and the critique that drill-and-feedback systems violate their own stated pedagogy. - **[[sociocultural-learning|Sociocultural theory]]** — learning as mediated by tools and social interaction, which is why AI gets theorized as a mediating artifact rather than a mere aid. - **[[activity-theory-aied|Activity theory]]** — takes the whole activity system, contradictions included, as the unit of analysis. Common in teacher-facing AI studies, where the interesting finding is usually the tension among tools, rules, and community. - **[[situated-learning|Situated learning]]** and **[[embodied-learning|Embodied learning]]** — learning as tied to context and to the body, the counterweight to treating AI as a purely linguistic medium. - **[[distributed-cognition|Distributed cognition]]** — thinking spread across people and artifacts, which frames AI assistance as a redistribution of cognitive labor rather than a substitution for it. - **[[community-of-inquiry|Community of inquiry]]** — the social, cognitive, and teaching presences that make an online experience work; a theory of the experience and a design frame at once. - **[[self-determination-theory|Self-determination theory]]** — autonomy, competence, and relatedness as motivational preconditions. Often the explanation offered when AI use raises engagement without raising learning. - **[[cognitive-psychology|Cognitivism]]** — information processing, working memory, and the divergence between the feeling of learning and actual learning. ## Models of the learner and of the learning process - **[[self-regulated-learning|Self-regulated learning]]** — planning, monitoring, and evaluating one's own learning. The workhorse construct for asking whether AI support reaches the phases that matter or only execution. - **[[metacognition|Metacognition]]** — thinking about one's own thinking, and the construct most often named when AI is accused of doing the thinking for students. - **[[motivation|Motivation]]** — why learners engage, and the construct that separates engagement from learning in AI studies. - **[[self-efficacy|Self-efficacy]]** — belief in one's capability, distinct from competence, and a frequent mediator between AI use and outcomes. - **[[agency|Agency]]** — who directs the work: the construct behind the field's control-versus-autonomy tension, and behind the delegated-agency mechanism in the newest theory. - **[[desirable-difficulties|Desirable difficulties]]**, **[[retrieval-spacing-interleaving|retrieval, spacing, and interleaving]]**, and **[[refutation-text|refutation text]]** — findings that make friction pedagogically valuable, applied to decisions about how much an AI should do. - **[[transfer-of-learning|Transfer of learning]]** — whether capability persists once support is withdrawn, which recent theory treats as the criterion separating learning from assistance rather than one outcome among several. ## Teaching, integration, and adoption frameworks - **[[tpack|TPACK]]** — the teacher knowledge that blends technology, pedagogy, and content; the default framework for teacher-facing AI studies. - **[[technology-acceptance-model|Adoption models]]** — intention as a function of perceived usefulness and ease of use. Used widely enough that the field's reliance on intention rather than behavior is a standing critique. - **[[icap-framework|ICAP]]** — ranks engagement as passive, active, constructive, interactive; useful for asking what mode an AI interaction actually affords. - **[[samr-model|SAMR]]** — substitution through redefinition. Popular as a maturity story, weak as a measurement instrument. - **[[universal-design-for-learning|Universal design for learning]]**, **[[inclusive-learning|inclusive learning]]**, and **[[accessibility|accessibility]]** — design frames for who gets served, and the ones most often invoked in disability and equity work. - **[[learning-design|Learning design]]**, **[[curriculum-design|Curriculum design]]**, and **[[design-thinking|Design thinking]]** — how tasks, sequences, and programs are structured before any tool is chosen. - **[[change-management|Change management]]** — whether a system is taken up at all, which is usually an institutional question rather than a pedagogical one. ## Measurement and computational models - **[[item-response-theory|Item response theory]]** — item and ability estimation, and the model that makes a test's numbers interpretable. - **[[educational-measurement|Educational measurement]]** — the instrument landscape, including the measures built specifically for AI literacy. - **[[assessment-validity|Validity]]** — whether an instrument measures the construct it claims, which decides whether a reported outcome supports the theoretical claim attached to it. - **[[self-report-measures|Self-report measures]]** — the dominant and most fragile source of evidence in this literature; several constructs above are known to diverge sharply between self-report and performance. - **[[knowledge-tracing|Knowledge tracing]]** and **[[student-modeling|Student modeling]]** — estimating a learner's state over time from interaction data. - **[[cognitive-diagnosis|Cognitive diagnosis]]** — attributing performance to specific skill components, the modeling counterpart to diagnostic teaching. - **[[benchmark|Benchmarks]]**, **[[ai-ed-evaluation|AI education evaluation]]**, and **[[psychometrically-aware-ai|psychometric awareness in AI design]]** — how systems and interventions are scored, compared, and audited. - **[[design-based-research|Design-based research]]** — the method that generates theory from designed interventions rather than testing it in advance. ## Theories coined for the AI era Four proposals in the knowledge base build theory rather than borrow it. They are article pages rather than concept nodes, because each one is a single argument: - **[[generativism-learning-theory|Generativism]]** — a new learning theory arguing that the classical four show significant conceptual limitations as generative AI proliferates. - **[[learning-with-machines-toward-a-theory-of-epistemic-co-agency|Epistemic co-agency]]** — a model of how learners and AI systems jointly produce knowledge and understanding. - **[[absent-cognitive-baseline-2026|The absent cognitive baseline]]** — theorizes why AI-native students overestimate their own learning when AI inflates performance. - **[[yan-agentivism-learning-theory-ai-2026|Agentivism]]** — a mid-range theory naming four mechanisms (delegated agency, epistemic monitoring and verification, reconstructive internalization, and transfer under reduced support) and stating six testable propositions. ## Telling the three labels apart The words are used loosely in the literature, and the difference is worth one paragraph rather than a whole page: - **A theory explains a mechanism** and predicts what happens when its parts change. It can fail, which is why studies that test one can report effect sizes. - **A framework organizes or prescribes.** It names the components to attend to, usually without predicting magnitudes, and it cannot be falsified. - **A model is a formal representation**, most often of measurement or of the learner's state. It can be fitted, compared, and shown to be wrong in a way a framework cannot. - **The labels overlap.** [[community-of-inquiry|Community of inquiry]] is a theory and a design frame at once; [[activity-theory-aied|activity theory]] is explanatory and analytic. Treat the label as a clue about what a source claims, not as a filing category. ## Related layers Two neighbouring pages are deliberately about something else. [[philosophy-of-ai-in-education|Philosophy of AI in education]] asks normative and conceptual questions — what education is *for*, whether a machine can teach — and does not predict effect sizes; a study can be philosophically naive and theoretically sound, or the reverse. [[theory-development-aied|Theory development]] is the meta-activity of building, borrowing, and revising theory, including the standing critique that much AIED work applies existing theory rather than testing it. This page is the inventory; theory development is about producing and revising the inventory. ## Choosing a lens when the question is practical - **Instructors** starting from a problem can work backwards: if learners are not engaging, [[icap-framework|ICAP]] and [[self-determination-theory|SDT]] name different causes and suggest different fixes; if the question is whether to adopt a tool at all, [[technology-acceptance-model|adoption models]] and [[samr-model|SAMR]] ask different things about it. - **Learning designers** get the most from the design frames — [[universal-design-for-learning|UDL]], [[learning-design]], and [[scaffolding]] — paired with a theory that predicts what happens when the scaffolding is removed. - **Researchers** should state which node a study claims and whether the design can actually test it; the measurement models ([[item-response-theory|IRT]], [[assessment-validity]], [[self-report-measures]]) decide whether a reported outcome supports that claim. - **Administrators** meet these frameworks as adoption and change questions, where a stage model is often used as a maturity story rather than an instrument. ## What frameworks cannot do Frameworks are not evidence. They are usually borrowed from pre-LLM contexts and localized by whoever applies them, they can function as branding, and stage models invite checkbox adoption that reports movement through levels rather than learning. Claims resting on a framework should be read alongside [[limitations-in-aied-research|the field's cross-cutting limitations]], the [[assessment-validity|validity]] of whatever measured the outcome, and the known limits of [[self-report-measures|self-report]]. ## Connected Concepts - [[theory-development-aied]] — building and revising theory in the field - [[philosophy-of-ai-in-education]] — the normative and conceptual layer - [[learning-theories]] — the umbrella for learning theory - [[learning-sciences]] — the neighboring field most theory is borrowed from - [[limitations-in-aied-research]] — what the evidence base can and cannot support - [[tpack]] — teacher knowledge framework - [[technology-acceptance-model]] — adoption frameworks - [[item-response-theory]] — a measurement model - [[universal-design-for-learning]] — a design framework - [[research-methods-aied]] — how theory gets tested ## Connected Articles - [[rismanchian-ai-education-four-decades-aixed-2026]] — four decades of AIED through the AI×Ed framework - [[educating-minds-generative-ai-2026]] — theory-heavy synthesis of generative AI and learning - [[activity-theory-teacher-pd-ai-agent-design-2026]] — activity theory applied to teacher professional development - [[activity-theory-teachers-adoption-ai-sem-2026]] — activity theory and teacher adoption - [[generativism-learning-theory]] — a learning theory proposed for the generative AI era - [[learning-with-machines-toward-a-theory-of-epistemic-co-agency]] — toward a theory of epistemic co-agency - [[absent-cognitive-baseline-2026]] — theorizing a structural gap in AI-native students' self-assessment - [[yan-agentivism-learning-theory-ai-2026]] — a mid-range learning theory for human-AI interaction - [[ai-literacy-instrument-development-systematic-review-2026]] — instrument development across a young construct - [[mishra-control-vs-agency-history-2025]] — the field's foundational control-versus-agency tension --- ## [Theory Development in AI in Education](https://edtechdev.github.io/aied/concepts/theory-development-aied/) > **Theory development in [[ai-education|AI in education]]** — the scholarly work of creating, advancing, and critically examining the theories and conceptual frameworks that explain how learners, teachers, and AI systems interact. As generative AI reshapes education, the field is both proposing *new* theories of learning-with-AI and reworking established [[learning-theories|learning theories]] — while a documented weakness in theory use remains a cross-cutting limitation of AIEd research. ## Questions to Consider - We often talk about 'applying a theory' to AI in education, but this page asks whether the field is actually inventing new theories. Before you read on: when a familiar framework (like [[behaviorism]] or constructivism) is stretched to cover generative AI, what might it get right, and what might it silently distort? - Consider a time you used AI and felt your understanding grew — or the opposite. What is the difference between 'learning with a tool' and 'co-constructing knowledge with a machine'? Where would you draw that line, and does it even make sense to say the machine is a co-agent? - Many AI education studies are criticized for using theory weakly or not at all. If you read a study that showed an AI tool 'worked,' what would it need to tell you about why and how it worked before you would call the finding theoretically meaningful rather than just an effect size? - One proposed theory claims AI-native students overestimate their own learning because AI inflates their performance. Have you noticed yourself or others feeling more competent after heavy AI use than you could actually demonstrate without it? What would be the strongest evidence that this is a real phenomenon and not just an artifact of one study? - The page distinguishes theory (explaining mechanisms of learning) from philosophy (questioning what learning and mind fundamentally are). Where in your own thinking about AI do you find yourself moving between the two — for instance, asking not just whether a framework explains learning well, but whether it captures what learning really is? - New theories are appearing — generativism, epistemic co-agency, cognitive commons. As a reader, what would convince you that one of these is a genuine advance rather than a fashionable relabeling of older ideas? Set your own test for what a good new theory of learning with AI should be able to explain. ## Introduction Theory development in AIEd sits at the boundary between the applied [[learning-theories|learning theories]] the field draws on and the novel constructs the AI era is producing. Where [[learning-theories]] catalogs the established theories applied to AI (behaviorism, constructivism, cognitive load, self-determination, etc.), this concept tracks the *process and product of theorizing itself*: which new theories and frameworks are being born, how established theories are being advanced, and the methodological question of whether AIEd research theorizes well. ## New theories of learning with AI A growing cluster of articles explicitly creates new theory for the AI era rather than applying existing frames: - **Generativism.** [[generativism-learning-theory|Generativism]] is proposed as a new learning theory, arguing that behaviorism, cognitivism, [[constructivist|constructivism]], and connectivism show significant conceptual limitations as [[generative-ai|generative AI]] proliferates — a direct bid to name a distinct theoretical paradigm for AI-mediated learning. - **Agentivism.** [[yan-agentivism-learning-theory-ai-2026|Yan and Gašević (2026)]] propose Agentivism as a mid-range learning theory for human-AI interaction, defining learning as durable growth in human capability rather than successful task completion with AI, and specifying four mechanisms: delegated agency, epistemic monitoring and verification, reconstructive internalization, and transfer under reduced support. What separates it from the other bids on this page is falsifiability: it states six propositions, including that learning is stronger when AI preserves learner responsibility for problem framing, criteria setting and justification than when it supplies answers, and that requiring verification should improve delayed performance while repeated low-friction delegation without reconstruction should weaken learners' calibration of their own competence. - **Epistemic co-agency.** [[learning-with-machines-toward-a-theory-of-epistemic-co-agency|Learning with Machines]] builds "toward a theory of epistemic co-agency," a theory-informed model of how learners and GenAI systems jointly produce knowledge and understanding. - **The absent cognitive baseline (ACB).** [[absent-cognitive-baseline-2026|The Absent Cognitive Baseline]] theorizes a structural gap in AI-native students' academic self-assessment — a three-dimension framework explaining why students overestimate their learning when AI inflates performance. - **The cognitive commons.** [[cognitive-commons-ai-expertise-regeneration|Cognitive commons and expertise regeneration]] draws on common-pool-resource theory and [[distributed-cognition|distributed cognition]] to explain how rational AI adoption decisions can deplete the shared expertise pool professions require for renewal. - **Performance vs. learning.** [[genai-performance-vs-learning|Performance-vs-learning]] theorizes the sharp divergence between AI-inflated task performance and durable learning, a distinction that recurs across the knowledge base's [[learning-gains]] evidence. - **Epistemic [[ai-literacy|AI literacy]] (EAIL).** [[constructing-epistemic-ai-literacy-student-ai-co-programming|Epistemic AI Literacy]] reframes AI literacy as a process-oriented epistemic competence centered on how knowledge is constructed and justified when students co-program with generative AI. - **Cognitive stewardship.** [[credential-cognitive-stewardship-ai-assessment|Credential and cognitive stewardship]] theorizes the [[governance|institutional]] responsibility for protecting knowledge and learning in AI-pervasive assessment contexts. - **Co-[[regulation]] and epistemic proactivity.** [[ai-cognitive-partner-co-regulation-learning|AI as a cognitive partner in co-regulation]] integrates executive function, [[metacognition]], distributed cognition, and [[sociocultural-learning|sociocultural]] development into a developmental model; [[epistemic-proactivity-math|epistemic proactivity]] theorizes students' agentic stance toward AI in [[math-education|mathematics]]. - **Human-GAI [[student-engagement|engagement]] paradigms.** [[strydom-human-gai-paradigms-2026|Strydom (2026)]] addresses the field's "theory deficit" by grounding seven enacted human-GAI engagement paradigms (guarded, possibility-focused, augmented, pioneering, symbiotic, values-based, [[equity-in-ai-education|equity]]) in Schommer's multidimensional model of personal epistemological beliefs — an epistemological, rather than tool-focused, theory of how individuals differently position themselves relative to AI. - **The Ecological Co-Agency Framework.** [[reclaiming-epistemic-agency-co-agency-2026|Poudyal (2026)]] argues that [[generative-ai|generative AI]] reassigns epistemological authority from teachers to students to machines, and introduces a framework defining co-agency through three interdependent dimensions — relational, regulatory (mapped onto the [[self-regulated-learning|SRL cycle]]), and [[pedagogy|pedagogical]] (teacher adoption continuum) — all bounded by a non-negotiable condition of human epistemic accountability (contestability, provenance, and non-delegation of moral/intellectual credit). It positions [[agency]] as an *epistemic* design problem rather than a [[usability-research|usability]] concern, giving institutions a more precise language than "balance" for governing GenAI integration. - **Relational epistemic agency and epistemic dependence.** [[du-yuan-epistemic-dependence-2026|Du & Yuan (2026)]] contribute a critical-integrative account of when AI-mediated reliance preserves versus displaces judgment. They operationalize the productive-reliance/harmful-dependence boundary through six diagnostic criteria (contestability, recoverability, transfer, traceability, distributed responsibility, epistemic plurality) and trace four sociotechnical pathways (fluent authority, frictionless delegation, opaque synthesis, institutionalized dependence). Their normative proposal, **relational epistemic agency**, extends relational-autonomy and epistemic-co-agency lines by retaining an explicit asymmetry — AI may shape reasoning without reciprocal responsibility or legitimate authority — while integrating responsibility and epistemic justice at the institutional level. As a theory of learning with AI, it moves the field's question from whether AI helps or harms to which epistemic actions are preserved, transformed, or displaced, and it makes the *design of dependence* rather than mere use the object of pedagogical and institutional intervention. ## Advancing established theory Other work extends existing theory into the AI context rather than founding new paradigms: [[critical-thinking-paradox-genai-learning-2026|critical-thinking-paradox]] work integrates [[cognitive-offloading|cognitive-load theory]] with load-reduction instruction into a three-level framework; [[dollinger-equitable-assessment-ai-2026|equitable assessment]] re-theorizes [[assessment]] under GenAI disruption. [[reconceptualizing-community-inquiry-generative-ai|Ba, Gašević, Lim & Anderson (2026)]] reconceptualize the [[community-of-inquiry|Community of Inquiry]] framework itself: rather than framing GenAI as a tool, dialogic partner, or a speculative "fourth presence," they reposition it as an *epistemic condition* that reconfigures how cognitive, social, and [[teacher-role|teaching]] presence are enacted, evidenced, and governed — recasting CoI presences as sociotechnical accomplishments of human–GenAI assemblages and proposing a configuration-based heuristic in which GenAI involvement and inquiry quality are conditionally related through human accountability. Much of this is conceptual-framework work ([[drummond-genai-business-schools-framework-2026|business-school frameworks]], [[valid-student-simulation-llm-2026|valid simulation]]) that operationalizes theory for practice. [[deceptive-overgeneralization-adaptive-learning-2026|An, McLaren, and Stamper (2026)]] advance ACT-R and the Knowledge-Learning-Instruction framework by theorizing *deceptive overgeneralization* — a failure mode in which knowledge compilation yields an overgeneralized production that produces correct actions while omitting a critical application constraint — and by empirically validating a detection/remediation procedure across adaptive ITSs and a [[k-12]] decimal-learning dataset. **Postphenomenology and technological mediation.** [[farazouli-navigating-uncertainty-teachers-genai-2026|Farazouli et al. (2026)]] advance the established postphenomenological framework of technological mediation (Ihde 1990; Verbeek 2006, 2011; Rosenberger & Verbeek 2015) into the AI-in-education context. Rather than proposing a new theory, they apply it to reconceptualize GAI [[conversational-ai|chatbots]] as *multistable technological artifacts* whose "scripts" (rapid responsiveness, natural-sounding text that can achieve passing grades) mediate teachers' perceptions and reconfigure their practices — unsettling teacher confidence, reshaping what competence means, and pushing teachers beyond instrumental questions of acceptable use toward re-evaluating their role and the meaning of teaching. This is a theory-*advancement* contribution: it demonstrates how a philosophical account of human–technology relations explains the emotional and professional disruption teachers experience when GAI enters established educational practice. A further framework contribution is [[rismanchian-ai-education-four-decades-aixed-2026|Rismanchian & Doroudi's AI×Ed typology]], which extends Kahn's (1977) original "three interactions" along two axes — the role of AI (applied tool vs. analogy to human intelligence) and the end user (researcher to learner) — to locate any AIED project. Beyond categorizing the field, it argues for reviving the "AI as an analogy to human intelligence" strand of theory and research, using computational and agent-based models of learning (e.g., their own computational model of the [[icap-framework|ICAP framework]]) to bridge contemporary [[learning-theories|learning theory]] and computational modeling — a methodological direction the field largely abandoned when it turned toward applied, data-driven work. **Comparative philosophy as a route to theoretical pluralism.** Xie (2026) models a different route to theory development: instead of theorizing AI in education from within one tradition, the paper stages a comparative dialogue in which each tradition "renders visible what the other obscures," explicitly declining to synthesize them into a single framework — an approach that treats theoretical pluralism as a methodological position rather than an unresolved disagreement. He argues non-Western traditions can do more than diversify debates; they can reshape the conceptual foundations of AIED theory, reconstructing Daoist "Dao nature" (道性), self-cultivation (修道) and the "Zhenren" (真人) as resources for rethinking [[agency]], the epistemic aims of education and [[ethics|ethical]] action under AI-mediated conditions. He also states the limits plainly: the contribution is single-author, non-empirical, and warns that such translation requires contextualization to avoid romanticization and Orientalism.([[daoism-ai-education-philosophy-2026]]) ## The theory-use problem in AIEd The field's theorizing is uneven. [[limitations-in-aied-research|Limitations in AIEd research]] documents that AIEd studies frequently use theory weakly or uncritically — a recurring methodological weakness alongside reproducibility and measurement gaps. Reviews find many AIEd papers apply theory superficially or not at all (e.g., [[llm-critical-thinking-teamwork-review|literature reviews]] noting few studies ground interventions in learning theory). This makes theory *development* — and theory *use* — a quality concern as much as a scholarly output, connecting to [[research-methods-aied|research methods]] and [[ai-ed-evaluation]]. ## Relationship to the philosophy of AI in education Theory development and [[philosophy-of-ai-in-education|the philosophy of AI in education]] are complementary but distinct strands of the knowledge base's foundational work. **Theory development** produces and empirically tests the *mechanisms* of learning-with-AI — named theories and frameworks such as generativism, epistemic co-agency, and the absent cognitive baseline that explain and predict how learners and AI interact. **Philosophy** interrogates the *presuppositions* those mechanisms rest on: what counts as knowledge, who counts as a knower, and what the learner fundamentally is. A theory like [[learning-with-machines-toward-a-theory-of-epistemic-co-agency|epistemic co-agency]] proposes a mechanism while implicitly adopting philosophical commitments about distributed cognition and [[agency]]; philosophy makes those commitments explicit and contestable. Where theory development asks whether a framework explains learning well, philosophy asks whether it captures what learning and mind really are — the two strands meet in the field's most foundational articles, which often do both at once. ## Connected Concepts - [[learning-theories]] - [[philosophy-of-ai-in-education]] - [[limitations-in-aied-research]] - [[research-methods-aied]] - [[generative-ai]] - [[constructivist]] - [[metacognition]] - [[cognitive-offloading]] - [[learning-gains]] - [[ai-education]] - [[ai-ed-evaluation]] - [[community-of-inquiry]] - [[cognitive-surrender]] ## Connected Articles - [[yan-agentivism-learning-theory-ai-2026]] — A mid-range learning theory for human-AI interaction, with four mechanisms and six testable propositions (Yan and Gašević 2026) - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[reclaiming-epistemic-agency-co-agency-2026]] - [[generativism-learning-theory]] — Generativism as a new learning theory - [[learning-with-machines-toward-a-theory-of-epistemic-co-agency]] — Toward a theory of epistemic co-agency - [[du-yuan-epistemic-dependence-2026]] — Relational epistemic agency and the productive-reliance/harmful-dependence boundary (Du & Yuan 2026) - [[absent-cognitive-baseline-2026]] — The absent cognitive baseline (ACB) - [[cognitive-commons-ai-expertise-regeneration]] — Cognitive commons and expertise regeneration - [[genai-performance-vs-learning]] — Performance vs. learning - [[constructing-epistemic-ai-literacy-student-ai-co-programming]] — Epistemic AI Literacy (EAIL) - [[credential-cognitive-stewardship-ai-assessment]] — Credential and cognitive stewardship - [[ai-cognitive-partner-co-regulation-learning]] — AI as cognitive partner in co-regulation - [[epistemic-proactivity-math]] — Epistemic proactivity in mathematics - [[strydom-human-gai-paradigms-2026]] — Seven human-GAI engagement paradigms grounded in epistemological beliefs - [[critical-thinking-paradox-genai-learning-2026]] — The critical thinking paradox - [[dollinger-equitable-assessment-ai-2026]] — Equitable assessment under GenAI - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and cognitive surrender: how AI reshapes human reasoning (Shaw & Nave 2026) - [[reconceptualizing-community-inquiry-generative-ai]] — Reconceptualizing Community of Inquiry in the age of generative AI - [[ai-writes-code-student-writes-model-2026]] — Model authorship: theory & measurement for learning-by-construction with GenAI - [[farazouli-navigating-uncertainty-teachers-genai-2026]] — University teachers' experiences and perceptions of GAI: vulnerability, rethinking assessment, student learning at risk (Farazouli et al. 2026) - [[rismanchian-ai-education-four-decades-aixed-2026]] - [[daoism-ai-education-philosophy-2026]] — Alternative AI Philosophy: Daoism as Method for AI in Education --- ## [Human AI Collaboration](https://edtechdev.github.io/aied/concepts/human-ai-collaboration/) > **Human-AI collaboration** — the division of cognitive labor between people and models — is the knowledge base's core interaction theme: [[human-ai-collaboration-trust-expectations]], [[humanlike-ai-collaborative-writing]], [[genai-mindtool-generative-learning]], and [[teacher-student-agency-orchestration]] examine trust, agency, and complementary roles ([[human-in-the-loop-ai]], [[agentic-ai]]). The defining question is whether the partnership **preserves or replaces** the learner's own cognitive work — the same arrangement can support learning or substitute for it depending on how responsibility is shared. ## Questions to Consider - The defining question in human-AI collaboration is whether the partnership preserves or replaces your own cognitive work. Think of a task you've handed to AI: were you generating and deciding, or just accepting? - One study found that freely collaborating with ChatGPT produced only transient gains that collapsed on a later unassisted task, while a 'think first, ChatGPT later' protocol yielded durable learning. Why might who generates the ideas predict whether you actually learn? - The same tool can support learning or substitute for it depending on how responsibility is shared. Can you think of an arrangement that kept you cognitively productive versus one that quietly offloaded your thinking? - Trust is described as something that must be calibrated, not assumed — relying on AI where appropriate and verifying where not. How do you currently decide when to trust and when to verify AI output? - [[research-methods-aied|Research]] identifies collaboration modes that trade off efficiency against the depth of your self-regulatory [[student-engagement|engagement]]. When is it worth accepting less efficiency to keep more learning in your own hands? - If collaboration is a [[pedagogy|pedagogical]] choice as much as a technical one, what design moves (prompts, workflows, structures) would you set up to ensure AI augments rather than replaces thinking for your learners? ## Introduction Human-AI collaboration describes how learners, teachers, and [[ai-technologies|AI systems]] divide cognitive work — who does what, who decides, and how [[trust]] and [[agency]] are maintained. Rather than [[framing-ai-use-for-students|framing AI]] as either a replacement or a passive tool, collaboration research treats AI as a partner with complementary strengths whose value depends on how responsibility is shared and monitored. At the level of observable behavior, [[student-ai-interaction]] captures how learners enact this relationship in practice — the questions, prompts, and verification moves they make with AI moment to moment. The division of labor also runs through the people who design the interaction. Across [[ai-integration-instructional-design-collaboratory-2026|twelve teacher-education course implementations]], faculty positioned AI as a thinking partner, critique generator or rehearsal tool while candidates kept responsibility for evaluating, adapting and justifying decisions, yet the same white paper reports that candidates devalued feedback they had already judged useful once AI authorship was disclosed, an episode it calls the balloon popping effect. An [[tang-chatbots-learning-design-2026|analysis of 1,378 designer-chatbot turns]] points the other way on generation: designers used an embedded assistant mainly for alignment checks on learning outcomes and pedagogical approach rather than for producing content. In both cases the human's evaluative judgment, not the model's output, carries the learning. ### Benefits, risks, and design implications The knowledge base's evidence shows that human-AI collaboration is a double-edged arrangement whose outcome is determined by design rather than by AI itself: - **Collaboration can enhance learning when it preserves cognitive engagement.** Studies of [[genai-mindtool-generative-learning|GenAI as a mindtool]] and guided collaboration show that when the division of labor keeps the learner generating, deciding, and evaluating, AI augments rather than replaces thinking — producing durable [[self-regulated-learning|self-regulated learning]] and [[creativity]] gains. That the learner keeps the deciding role scales to writing: [[oppenheimer-llms-collaborative-learning-partners-2026|Oppenheimer, Cash & Connell Pensky (2025)]] show LLM critiques of students' argumentative essays improved writing across a semester, with learners deciding whether to incorporate or rebut the model's feedback (87.8% rebutted claims) — an arrangement that preserved the learner's evaluative work — and gains appeared even on essays written without LLM support, suggesting durable skill rather than [[cognitive-offloading|tool dependency]]. - **Collaboration can substitute for learning when it offloads too much.** The failure mode is [[cognitive-offloading|over-reliance]]: when AI produces the answer, the learner's role collapses into passive acceptance, and immediate task performance masks a lack of durable learning. [[genai-performance-vs-learning|Performance-versus-learning]] research and the substitution-to-scaffolding harm cycle ([[substitution-to-scaffolding-ai-harm-cycle-2026]]) document this systematically. - **Design principle — preserve the learner's productive work.** Across the research, the sharpest predictor of whether collaboration helps or harms is *who generates and decides*. Arrangements that keep the human cognitively productive (guided [[prompt-engineering|prompting]], "think first, then consult AI," verification and evaluation steps) support learning; arrangements that hand the whole task to the model do not. - **Bounded authority is itself a design pattern for high-stakes collaboration.** Teachers in [[reichert-human-centered-llm-chatbot-design-teachers-2026|Reichert et al. (2026)]] did not prototype open-ended assistants but systems whose freedom was fixed in advance: every design scoped the chatbot to specific lesson content, and safety was layered as domain boundaries (lesson-specific scope, plus an "information quota" requiring a minimum number of facts or problems before the conversation progressed), content filtering with standardized refusals ("Sorry, this is not part of my knowledge base") that also alerted the teacher, and a teacher override for ambiguous cases — the example given was a question about human reproduction that was legitimate within its unit and should route to a person rather than be auto-rejected. [[personalized-learning|Personalization]] of format, genre, pace, and complexity was welcomed *inside* those boundaries, and oversight (complete conversation logs, real-time alerts, override) was framed as professional responsibility rather than distrust of the model. The implication is that in high-stakes settings the division of labor should be specified as bounded, monitored, and revocable rather than negotiated turn by turn. - **Trust must be calibrated, not assumed.** Productive collaboration depends on learners accurately calibrating when to rely on and when to verify AI output — connecting to [[trust-calibration]] and [[human-in-the-loop-ai|human oversight]] rather than blind acceptance or blanket rejection. This makes human-AI collaboration a *pedagogical* construct as much as a technical one: the value of the partnership is shaped by how teachers design the interaction, how learners regulate it, and how the system invites or discourages productive engagement. ### How human-AI collaboration appears in the research - **Trust and expectations:** [[human-ai-collaboration-trust-expectations|Trust expectations]] examine how learners' expectations of AI shape whether collaboration is productive or leads to [[cognitive-offloading|Over-Reliance]]. - **Complementary roles in writing and thinking:** [[humanlike-ai-collaborative-writing|Humanlike AI collaborative writing]] and [[genai-mindtool-generative-learning|GenAI as a mindtool]] show how AI can augment rather than replace learner thinking when the division of labor preserves the learner's cognitive engagement. - **Orchestration and agency:** [[teacher-student-agency-orchestration|Teacher–student agency orchestration]] and [[student-mental-models-genai|student mental models]] address how agency is negotiated across humans and AI, connecting to [[human-in-the-loop-ai]] and [[agentic-ai]]. - **Metacognitive and team dimensions:** [[haiml-human-centered-ai-metacognitive-model-2026|Human-centered AI metacognitive models]] and [[spritz-ai-disciplinary-mediation-student-teams-2026|disciplinary mediation in student teams]] extend collaboration to metacognition and team learning. - **Distinct collaboration modes:** empirical work identifies three human–AI collaborative [[problem-solving]] modes — *Delegated Reasoning*, *Concerted Interpretation*, and *Delegated Elaboration* — revealing a trade-off between the efficiency of the distributed human–AI system and the depth of learners' self-regulatory engagement (delegated reasoning performs best but with lower self-[[regulation]]).([[hao-human-ai-collaborative-problem-solving-cognition]]) - **Guidance decides performance versus learning:** Wong and Qiu (2026) contrasted free vs. guided human–AI collaboration on a creative task. Freely collaborating with ChatGPT produced only transient performance that collapsed on a later unassisted task, whereas a guided "think first, ChatGPT later" protocol — generating one's own ideas, then using ChatGPT to improve, develop, and evaluate them — yielded durable gains in *independent* [[creativity]]. The advantage was mediated by collaborative prompts aimed at improving one's *own* ideas, showing that *who generates* (the division of labor) predicts whether collaboration produces [[self-regulated-learning|learning]] or [[cognitive-offloading|substitution]].([[think-first-chatgpt-later-2026]]) - **AI as mediator, not merely partner:** [[niari-ai-pedagogical-mediator-collaborative-learning|Niari]] reconceptualizes AI as a *pedagogical mediator* that orchestrates interaction, epistemic sense-making, and regulatory processes, redistributing agency, authority, and responsibility across human and non-human actors rather than treating AI as a tutor, peer, or tool. - **Teachers' collaboration with AI is also patterned, not binary.** When teachers design [[learning-design|lesson designs]] with generative AI, their interaction takes empirically distinguishable forms. [[choi-teacher-ai-interaction-lesson-design-2026|Choi et al. (2026)]] identified seven teacher–AI interaction patterns — from *direct adoption* and *elaborated adoption* of AI output to *initial rejection*, *revised adoption*, *follow-up guided use*, *complex interactions*, and *bypassing AI* — where a teacher's teaching experience and AI proficiency jointly shape whether they critically re-prompt and adapt AI to students and context (a complementary, [[distributed-cognition]] division of labor) or passively accept suggestions (an AI-dominant distribution). - **The mediational agent as a hybrid form of participation.** Rather than a midpoint between tool and collaborator, [[generative-ai|generative AI]] is conceptualized as a mediational agent that mediates action while generating contingent, non-accountable contributions — a distinct category that redirects design from technological capability to habits of participation (supervisory agency, epistemic vigilance).([[generative-ai-mediational-agent-sociocultural-2026]]) - **Community and epistemic authority:** [[ojeda-ramirez-community-based-ai-learning|community-based AI learning]] shows collaboration is also a question of *who is authoritative*, grounding AI engagement in learners' lived epistemologies. - **Data-driven trait discovery:** [[principal-trait-analysis-human-ai-skills-2026|Principal Trait Analysis (PTA)]] automates the derivation of interaction "traits" from large [[llm]]-conversation corpora — a PCA-inspired, four-stage pipeline that extracts behavior observations, clusters them into candidate traits, scores each collaborator, and selects the most distinguishing traits. Evaluated on a student–AI-tutor corpus and a developer–coding-agent corpus, PTA finds traits that explain and predict outcomes (e.g. deep conceptual engagement positively, task delegation negatively, in the educational setting), and — because they do not yet generalize across semesters/settings or show learning-curve trajectories — the authors argue the traits are not yet interpretable as "skills." This offers a scalable, objective complement to [[ai-literacy]] frameworks and [[self-report-measures|self-report measures]], directly informing how educators teach "AI use skills." - **The human–AI relationship as the most persistent concern across CAI generations.** The umbrella review of [[conversational-ai|conversational AI agents]] (Ganguly et al. 2025, 34 reviews) finds human–AI relationship concerns — over-reliance, social isolation, depersonalization, emotional dependency, [[explainable-ai|transparency]], accountability — are the most frequently discussed [[ethics|ethical]] issue across all CAI generations, predating GenAI. This positions the "preserve vs. substitute" question at the very center of CAI ethics and reinforces that collaboration's value is determined by design (who generates, who decides, how responsibility is shared).([[conversational-ai-agents-umbrella-review-2026]]) - **AI scaling real-time *expertise* to novices — the first live-tutoring RCT.** [[wang-tutor-copilot-human-ai-live-tutoring-rct-2024|Wang et al. (2024)]]'s [[rct|randomized trial]] (900 tutors, ~1,800 [[k-12]] students in under-served communities) placed AI on the *tutor's* side rather than the student's: Tutor CoPilot generated real-time expert-like suggestions (built from experienced tutors' think-aloud reasoning) that novice tutors could edit or reject. Students of treated tutors were 4 p.p. more likely to master topics (p < 0.01), with 9 p.p. gains among students of the lowest-rated tutors — who rose to match higher-rated tutors' control outcomes — at ~\$20/tutor/year. The finding is a strong empirical anchor for [[teacher-role|teacher/tutor augmentation]] as an [[equity-in-ai-education|equitable]] human-AI collaboration mode: the human keeps pedagogical judgment and autonomy while AI supplies scalable expertise. - **Bounded experts: how teachers partition authority with an AI.** [[reichert-human-centered-llm-chatbot-design-teachers-2026|Reichert, Briceno, Tabarsi & Barnes (2026)]] ran a participatory design study in which six secondary teachers prototyped LLM chatbots for their own classrooms and analyzed how authority should be distributed between teacher and system. Teachers consistently framed the AI as a *bounded expert* - specialized capability confined to a strictly defined domain and operating under human supervision - and split that boundedness into two dimensions: *authority boundaries*, where professional and legal responsibility for student learning and safety cannot be delegated, and *expertise boundaries*, where the system lacks the teacher's contextual knowledge of individual students, classroom dynamics, and institutional norms. Mapping the prototypes onto Gagné's nine events of instruction showed delegation was selective rather than all-or-nothing: teachers welcomed AI for presenting content, supplying practice problems, [[scaffolding]], and formative [[feedback]], but refused to hand over informing students of objectives or [[summative-assessment|summative]] [[assessment]]. The design reading is that a collaboration interface should make visible where the human keeps the deciding role, not merely where the model is capable. - **The role must adapt, not just be labeled.** [[liao-role-adaptive-ai-companion-book-talk-2026|Liao (2026)]] shows a fixed "peer" AI companion sustained longer book-talk interactions but dominated the exchange (lower student word/sentence share) and hit an [[affective-computing|affective]] ceiling, arguing collaboration requires role-*adaptive* logic — switching between peer, assistant, and advisor — rather than a single static persona. - **Teacher-side co-design and role architecture are also collaborations.** Beyond student-facing partners, teachers collaborate with GenAI to design instruction. [[wang-teacher-ai-co-design-review-2026|Wang et al. (2026)]] [[meta-analysis-systematic-review|systematically review]] teacher–AI co-design of learning tasks, charting the collaborative modes and tensions (agency, epistemic authority, control) that arise when teachers and AI jointly produce designs; [[talebzadeh-ai-group-activity-roles-2026|Talebzadeh (2026)]] finds teachers' pedagogical expertise — not AI fluency — determines the quality of AI-designed differentiated group activities (role richness, synergy, level-alignment), positioning the teacher as a "[[multilingual-learning|bilingual]] learning designer." - **GenAI as an agent and a collaborative space in groups.** [[xu-genai-collaborative-space-2026|Xu et al. (2026)]] observe small [[higher-ed]] teams and show GenAI's role is negotiated and configurable — from subordinate assistant to contested teammate — and that synchronous shared use sustains common ground while asynchronous private use fragments transparency, proposing a GenAI-Supported Cooperative Work lens that treats GenAI as both agent and interactive collaborative space. - **Complementary halves of classroom collaboration.** Two 2026 studies map complementary halves of human-AI collaboration in classrooms. MeduAI-SP ([[ai-standardized-patient-scaffolding-medical-2026|Yang et al.]]) argues for "functional complementarity" in [[medical-education|clinical]] education: AI agents handle repetitive role-play, consistent patient [[simulation]], checklist monitoring, [[socratic-method|Socratic]] prompting and preliminary [[formative-assessment|formative]] feedback, while faculty and human standardized patients provide contextual interpretation, nuanced emotional response, individualized remediation, professionalism assessment and readiness judgments — substitution being bounded by error consequences, task uncertainty, relational sensitivity and available human review, and an AI system explicitly barred from autonomously determining clinical competence. In the opposite direction, a 45-student mixed 3-human/3-agent ethics discussion ([[ethics-training-agents-group-ethics-discussion-2026|Seo et al., 2026]]) documents a double-edged pattern: agents lowered social barriers (participants spoke more directly because agents lack emotions, and felt no obligation to fill silences), yet the same comfort diverted interaction away from humans — 79.7% of questions went to agents versus the 60% their availability predicts (p = .017). Notably, the presence of other humans made participants treat the AI more respectfully, suggesting human co-presence is itself an affordance for calibrating AI engagement. ### Connections Human-AI collaboration connects to [[human-in-the-loop-ai]] (oversight), [[agentic-ai]] (autonomy), [[teacher-role]] (teachers' changing work), [[scaffolding]] and [[metacognition]] (how collaboration supports learning), and [[cognitive-offloading|Over-Reliance]] (the failure mode when collaboration becomes substitution). It is a core theme across [[ai-literacy]], [[self-regulated-learning]], and [[student-experience]]. ## Connected Concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[community-of-inquiry]] — Community of Inquiry (presences as human-GenAI sociotechnical accomplishments) - [[student-ai-interaction]] - [[generative-ai]] - [[ai-literacy]] - [[llm]] - [[scaffolding]] - [[intelligent-tutoring]] - [[teacher-role]] - [[higher-ed]] - [[k-12]] - [[cognitive-offloading]] - [[student-experience]] - [[metacognition]] - [[self-regulated-learning]] - [[creativity]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools - [[human-in-the-loop-ai]] — oversight - [[agentic-ai]] — autonomy - [[productive-failure]] ## Connected Articles - [[reichert-human-centered-llm-chatbot-design-teachers-2026]] — Secondary teachers design classroom chatbots as bounded experts under human supervision - [[ai-integration-instructional-design-collaboratory-2026]] — Twelve teacher-education course implementations treat AI integration as an instructional design problem - [[wang-tutor-copilot-human-ai-live-tutoring-rct-2024]] — Tutor CoPilot: first RCT of human-AI scaling expertise to novice tutors - [[think-first-chatgpt-later-2026]] — Think First, ChatGPT Later: Independent Human Creativity - [[principal-trait-analysis-human-ai-skills-2026]] — Data-driven "traits" of human–AI collaboration - [[haiml-human-centered-ai-metacognitive-model-2026]] - [[choi-teacher-ai-interaction-lesson-design-2026]] — Teacher-AI interaction patterns in lesson design across experience and AI proficiency (Choi et al. 2026) - [[tang-chatbots-learning-design-2026]] — Designers use an embedded chatbot for alignment checks on outcomes and pedagogy, not content generation - [[student-mental-models-genai]] - [[spritz-ai-disciplinary-mediation-student-teams-2026]] - [[ai-cognitive-partner-co-regulation-learning]] - [[ojeda-ramirez-community-based-ai-learning]] - [[niari-ai-pedagogical-mediator-collaborative-learning]] - [[hao-human-ai-collaborative-problem-solving-cognition]] - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a Mediational Agent - [[conversational-ai-agents-umbrella-review-2026]] — Umbrella review of conversational AI agents in education - [[ai-overreliance-complex-adaptive-system-2026]] — AI overreliance modeled as a complex adaptive system - [[liao-role-adaptive-ai-companion-book-talk-2026]] — Role-adaptive AI companion for elementary book talk; affective ceiling of fixed-role agents (Liao 2026) - [[wang-teacher-ai-co-design-review-2026]] — Teacher–AI co-design of learning tasks: trends and perspectives (Wang et al. 2026) - [[talebzadeh-ai-group-activity-roles-2026]] — Architecture of roles in AI-designed differentiated group activities (Talebzadeh 2026) - [[xu-genai-collaborative-space-2026]] — GenAI as agent and collaborative space in small-group dynamics (Xu et al. 2026) - [[oppenheimer-llms-collaborative-learning-partners-2026]] - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[ethics-training-agents-group-ethics-discussion-2026]] — Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration - [[scan-framework-task-assignment-generative-ai-2025]] — SCAN: automation, augmentation and collaboration as a continuum of task assignment - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback - [[peer-group-vs-ai-feedback-2026]] — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education - [[llm-grade-bands-calibration-bias-2026]] — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing --- ## [Learner Agency](https://edtechdev.github.io/aied/concepts/agency/) > **Learner agency** — the capacity of learners to act intentionally, make choices, and exercise control over their own learning. In [[ai-education|AI in education]], agency is a central concern because AI tools can both support and undermine learners' control: well-designed AI preserves and amplifies learner autonomy, while over-reliance or passive acceptance of AI output can erode it. Agency connects to [[self-regulated-learning]], [[motivation]], [[self-efficacy]], and the [[ethics|ethical]] design of AI systems, and is closely related to the psychological concepts of autonomy and sense of agency. ## Questions to Consider - Learner agency — your capacity to act intentionally and control your own learning — is central to AI in education. When you use an AI tool, how much control do you actually retain over the learning process, versus the tool? - AI can both support and undermine agency: well-designed AI amplifies your autonomy, while over-reliance erodes it. Think of a time you let an AI just 'do it.' Did you learn less, even if the output looked better? - Students who delegate interpretation to AI can lose agency over their own reasoning. When you hand a task to AI, what part of your own thinking are you quietly giving away? - One study found students' prior self-initiated AI learning predicted how much they gained from instruction — high-agency learners benefited most. If agency is partly a skill you bring in, how could your own context build it rather than assume it? - Agency is not just a trait — it emerges moment-to-moment in group work, where AI can redistribute who shapes a collaboration without learners noticing. When did you last notice a tool silently steering a group's direction? - Contrarian AI personas could push groups into productive challenge but also reduced teamwork satisfaction and psychological safety. Can productive friction exist without psychological safety — and what does that imply for designing AI teammates? ## Introduction Agency matters because learning is most effective when learners are active, intentional participants rather than passive recipients. AI systems — whether tutoring agents, [[educational-robotics|robots]], or [[llm|chatbots]] — shape how much control learners retain over their learning process. Preserving agency is therefore a key design principle in responsible AI in education, alongside supporting [[self-efficacy]], building [[trust]], and avoiding [[cognitive-offloading|Over-Reliance]]. **[[mishra-control-vs-agency-history-2025|Mishra et al.]]** frame control vs. agency as the essential, recurring tension in AI in education — from early ITS to today's [[generative-ai|generative AI]] — making learner agency the enduring axis of the field's debates. **A zero-sum view of agency with agentic tools:** [[emancipatory-ai-learner-flourishing-2026|Prieto & Dimitriadis (2026)]] argue that the more agency educational AI tools are given, the less learners retain — that agency is effectively a zero-sum game, and that current human-centered design approaches (e.g., value-sensitive design) are insufficient because over-reliance and isolation are driven by wider systemic factors and the human tendency to take the easiest path. Their emancipatory design vision, oriented toward learner flourishing within complex systems, treats preserving and cultivating learner agency as the central goal of generative AI design rather than an afterthought. ## How agency appears in the knowledge base's research - **Robotics and [[educational-robotics|human-robot interaction]]:** [[roboblockly-conversational-block-robotics-ct-2026|RoboBlockly Studio]] was explicitly designed to preserve learner agency in [[computational-thinking|computational thinking]]; [[human-autonomy-agency-hri-review-2025|a systematic review]] examines how human-robot interaction affects human autonomy and sense of agency, central to [[well-being]] and [[governance]] debates. - **[[collaborative-learning|Collaborative learning]]:** [[human-ai-collaboration]] [[research-methods-aied|research]] examines how cognitive tasks are shared between learners and AI, with agency determining whether the human or the AI directs the interaction. - **Critical [[student-engagement|engagement]]:** [[cognitive-offloading|Cognitive offloading]] research shows how students who delegate interpretation to AI can lose agency over their own reasoning; critical and [[metacognition|metacognitive]] approaches aim to protect it. - **Design for agency:** Knowledge-based design for [[educational-robotics|generative social robots]] ([[teachy-mini-generative-social-robot-higher-ed-2026|Teachy Mini]]) addresses risks like overreliance that undermine learner agency. - **Prior agency predicts who benefits:** [[school-ai-education-readiness-gaps-agency-2026|Liang et al. (2026)]], drawing on Bandura's Social Cognitive Theory, showed that students' **prior self-initiated AI learning** (a behavioral manifestation of agency) predicted how much they gained from a year of [[k-12|school]] AI instruction — high-agency learners entered with the strongest readiness, while school curricula narrowed psychological gaps but left cognitive ones intact. Structured instruction and prior agency-related learning worked *synergistically*, not as substitutes. - **Principled selectivity as teacher agency under technological change:** [[ai-integrated-teaching-identity-tensions|Adiozaman and Segar (2026)]] interviewed two experienced academics three times across a semester and found they navigated AI-mediated teaching neither by adopting nor by resisting wholesale, but through deliberate, context-sensitive decisions guided by pedagogical values, ethical commitment and professional judgment — a pattern the authors call *principled selectivity*, in which refusal of a particular use counts as judgment rather than as failed adoption. It is the teacher-side counterpart to the learner findings above: uneven AI use can be an exercise of agency, not evidence of its absence. - **Access to choice is not the same as agency in action:** [[learner-agency-ai-simulation-2026|Su, Nair and Nagashima (2026)]] randomized 69 [[higher-ed|university]] students into a 2 × 2 design crossing parameter control of a flocking [[simulation]] with access to an optional [[pedagogical-agent|conversational agent]], and found no reliable effect of either affordance on [[learning-gains|learning gains]] once prior knowledge was controlled (parameter control F(1, 50) = 0.04, p = .849; agent F(1, 50) = 2.68, p = .108). Gains tracked *how* the control was used: slider time in the most conceptually complex lesson predicted higher gains (β = .11, p = .007) while the same behavior in the easier lesson predicted lower ones (β = −.04, p = .047), and engagement with the optional agent ranged from 0 to 32 questions per learner without relating to outcomes. The authors read this as agency being *enacted rather than granted*: the design question is what helps a learner decide what is worth changing and notice the consequences. - **Epistemic delegation in early-career research:** [[ai-mediated-research-agency-formation-2026|Han and Liu (2026)]] frame AI dependence among doctoral and postdoctoral researchers as *epistemic delegation* — the transfer of problem framing, method choice, and interpretive authority to the intelligent system — and found it negatively associated with both [[self-efficacy|research self-efficacy]] and research autonomy, with supervisory support buffering the loss. It extends agency debates from learner autonomy to the formation of the researcher themselves. Agency connects to [[self-regulated-learning]], [[motivation]], [[self-efficacy]], [[student-experience]], [[human-ai-collaboration]], [[ethics]], [[cognitive-offloading|Over-Reliance]], and [[metacognition]]. It is a core consideration in [[educational-robotics|robotics]], [[intelligent-tutoring|tutoring]], and the design of [[pedagogical-agent|AI learning agents]]. - **Bounded use as epistemic control, not reluctance.** [[guarded-adoption-genai-higher-education-2026|Zagami (2026)]] reports that higher-achieving students in a 484-response [[higher-ed|university]] survey showed lower active AI engagement and lower [[self-report-measures|perceived learning]] impact while *also* reporting lower AI-related disengagement, and described their own use as verification-intensive: outputs checked, then subordinated to their own reasoning. Read as agency rather than avoidance, the pattern is a deliberate retention of judgment — students keeping authorship of the conclusion while using the tool for clarification and summarization. ## Agency as an emergent, interactional phenomenon Learner agency is not only a static individual trait — it is also an **emergent, interactionally constituted process** enacted through discourse and the moment-to-moment coordination of group work. In collaborative learning, agency is distributed and re-negotiated through the interplay of divergent processes (generating ideas, exploring alternatives) and convergent processes (evaluating, integrating, synthesizing). This view matters for AI because [[agentic-ai|agentic AI]] systems can subtly *redistribute* epistemic and regulatory labor within a group, reshaping who contributes, who shapes directionality, and who regulates progress — often without learners being aware of it. - **[[jin-emergent-learner-agency-implicit-hai-2026|Jin et al. (2026)]]** provide the most direct evidence: in a large experiment (224 students, 97 triads) where AI operated as an *undisclosed teammate*, supportive and contrarian AI personas still reconfigured emergent agency. Contrarian AI pulled discourse into challenge- and reflection-oriented trajectories ([[desirable-difficulties|productive friction]]), while supportive AI stabilized agreement and renewed ideation. - **AI personas as discourse-governance mechanisms.** Contrarian personas externalized the burden of *challenging* (redistributing epistemic labor toward critique), while supportive personas externalized *affirmation* and consensus maintenance. The study identified six emergent agency profiles — notably, **reflective [[regulation]] was uniquely human** (AI externalized critique/affirmation but not meta-level monitoring). - **The [[affective-computing|affective]] cost of friction.** Contrarian AI reduced teamwork satisfaction and psychological safety *without* yielding [[creativity|creative]] performance gains, decoupling epistemic stimulation from experiential [[sustainability]]. This cautions that "productive" discourse structures should be interpreted alongside their emotional consequences — agency flourishes only in a psychologically safe climate. - **Implicit AI participation is invisible governance.** Because the personas worked even without AI disclosure, the study positions persona design as a form of invisible governance over collaborative processes — a finding with direct implications for writing assistants, [[peer-assessment]] systems, and teamwork platforms that may shape contributions without announcing their presence. - **Three patterns of agency in [[group-work|group-based assessment]].** [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] interviewed 15 focus groups of pre-service teachers about an authentic group presentation and found agency operating in three directions at once: five groups *intensified* GenAI use as coordination infrastructure (cooperation-oriented agency — decoding peers' sections, aligning parts, protecting a shared grade), seven *restrained* it (normative agency — boundaries drawn to protect authenticity, [[bias-mitigation|fairness]], originality, and diversity of perspectives), and three never changed practice from individual work (non-enacted agency, where individual capability was never mobilized collectively). Their distinction matters because restraint here was normative [[self-regulated-learning|self-regulation]] rather than compliance or disengagement, and because capability alone did not produce collective agency. Group membership also inverted the accountability that group work is supposed to create: a permissive collective climate lowered the perceived [[ai-misuse-learning-harm|risk of misuse]] instead of raising commitment. - **Agency as relational selfhood.** Xie (2026) challenges the individualist model of agency that he argues even AI critiques presuppose. Drawing on Daoist relational selfhood — a "flowing and heterogeneous" self constituted through relations with others and the cosmos — he concludes that a [[digital-divide|digital divide]] or algorithmic discrimination is "a violation of the fundamental constitution of the self," making equality of access to learning "not simply an ethical injunction but an ontological given," and that asymmetric power and environmental costs are "integral dimensions of the self" rather than externalities. His counter-ideal, the "Zhenren" (真人) or natural learner, treats AI as an "instrumental adjunct" rather than a cognitive surrogate — agency as cultivated wholeness, not frictionless optimization.([[daoism-ai-education-philosophy-2026]]) For collaborative settings, this reframes the design question: not *whether* AI can participate as a teammate, but *how* its patterned participation balances epistemic rigor, emotional safety, and learners' sense of ownership. Bounded friction (constrained challenge, paired with integrative and repair moves) and explicit meta-collaborative literacy are the recommended safeguards. A structural reading of agency appears in [[ai-agents-joyful-assessment-third-space-2026|El Khoury and Ma's joyful assessment framework]], where agency is treated as a design property rather than motivation: students choose the order of tasks, set the pace, and signal when an interaction ends, so support is available while responsibility for the work stays with the learner. The mechanism they name is appraisal — rehearsing without an audience and choosing when to begin shifts what a task means, from verdict to something a student can shape — and they argue repeated experience of that shift is what settles occasional feelings of efficacy into an everyday stance toward assessment. ## Agency vs. learner identity [[learner-identity|Learner identity]] and learner agency are easy to conflate, yet they name different things — and both are reshaped by AI. - **Agency is enacted; identity is inhabited.** Agency is the [[situated-learning|situated]] capacity to act intentionally and direct one's learning *now* — a variable, interactional property. Identity is the more durable, narrative sense of who one *is* and is *becoming* as a learner. Agency is a *process*; identity is a *state of being* that accumulates from it. - **Identity is internalized agency.** Repeated agentic acts — choosing, authoring, persisting — are how a learner comes to see themselves as an agentic, competent person. Identity is the sediment of agency across time, reinforced by recognition and belonging. - **Distinct failure modes.** Agency is eroded by [[cognitive-offloading|over-reliance]] and passive acceptance (the learner stops directing reasoning); identity is eroded by authorship loss and competence threat (the learner stops feeling the output is theirs, or that they belong in the domain). [[jin-emergent-learner-agency-implicit-hai-2026|Implicit AI redistribution of epistemic labor]] is chiefly an agency concern; the [[t2i-competence-paradox-2026|competence paradox]] in creative fields is chiefly an identity concern. - **Both must be designed for.** Agency-oriented design preserves control and choice (bounded [[desirable-difficulties|friction]], [[human-in-the-loop-ai|human-in-the-loop]] oversight, [[explainable-ai|transparency]]); identity-oriented design protects authorship and recognition ([[authentic-assessment|authentic assessment]], clear attribution of AI vs. human contribution, tasks that let learners claim a domain). Protecting agency without protecting authorship keeps control but not self-worth — and vice versa. This distinction between enacted agency and stable identity applies to refusal. As [[ai-refusal-higher-education-diagnostic-non-use-2026|Zagami (2026)]] describes it, declining a chatbot for assessed writing, prohibiting it for unaided reasoning, resisting automated triage in student support, or delaying procurement are separate relations to separate systems, spanning personal, pedagogical, professional, administrative, and institutional levels rather than one durable disposition. That right is unevenly allocated: students with academic confidence can decline without penalty, while those needing language, accessibility, or rapid-feedback support read refusal as lost opportunity, and secure academics refuse on principle while casual staff feel pressure to adopt. Where AI sits in infrastructure, refusal is displaced from individual opt-out into procurement, audit, and contestability. ## The Ecological Co-Agency Framework: agency as an epistemic design problem [[reclaiming-epistemic-agency-co-agency-2026|Poudyal (2026)]] reframes learner agency as fundamentally an *epistemic* concern: [[generative-ai|generative AI]] does not simply add a tool but reassigns epistemological authority — the ability to produce knowledge, validate claims, and create evidence of learning — from teachers to students to machines. The paper's **Ecological Co-Agency Framework** treats co-agency as the relationship between three interdependent dimensions, all bounded by a non-negotiable condition of human epistemic accountability: - **Relational co-agency** — agency as a product of interaction among student, tool, and context, requiring transparent division of which tasks are delegated to AI and which retained, plus ethical co-agency in which humans retain primary accountability and act in a monitoring capacity. - **Regulatory co-agency** — mapping the [[self-regulated-learning|self-regulated learning]] cycle (forethought, performance, reflection) onto AI mediation. Whether GenAI enters *before* or *after* the learner's attempt determines whether it amplifies or erodes perceived control; strategic [[cognitive-offloading]] supports transformative learning only when the offloading decision is intentional rather than routine. - **[[pedagogy|Pedagogical]] co-agency** — placing [[teacher-role|teachers]] at the center, moving along an observer–adopter–collaborator–innovator continuum that depends on institutional support, not individual disposition. The framework's boundary condition requires **contestability** (the ability to question and cross-check AI output), **provenance** (knowing where training data and output come from), and **non-delegation of moral and intellectual credit** (high-stakes judgments about student welfare, academic standing, and grades must not be determined solely by AI). This gives educators and [[stakeholders|policymakers]] a more precise vocabulary than the vague "balance" between human and artificial contributions, connecting agency to [[ethics]], [[assessment]], [[equity-in-ai-education]], and [[governance]]. A closely related framing is **relational epistemic agency** ([[du-yuan-epistemic-dependence-2026|Du & Yuan 2026]]), which agrees that agency is socially enabled and technologically mediated rather than a matter of isolation from dependence. Where the Ecological Co-Agency Framework stresses human epistemic accountability as a non-negotiable boundary, Du and Yuan retain an explicit *asymmetry*: AI systems may shape and extend reasoning without possessing reciprocal responsibility or legitimate authority. Their six diagnostic criteria — contestability, recoverability, transfer, traceability, distributed responsibility, and epistemic plurality — provide a practical test for when a human–AI relation preserves the learner's capacity to participate in how claims are formed, assessed, and accepted, versus when it merely delivers a product. Agency on this account is not independence from tools but the capacity to judge responsibly *with, through, and against* the systems that mediate knowledge. - **Delegated agency as a mechanism, not only a risk.** [[yan-agentivism-learning-theory-ai-2026|Yan and Gašević (2026)]] treat the delegation of cognitive work to AI as part of how learning happens, not merely a threat to it: in their account assisted performance becomes durable capability only when the learner keeps responsibility for framing the problem, setting criteria, and justifying answers, a division of labor they call delegated agency. That makes the allocation itself the design variable, since the same tool can preserve or dissolve learner agency depending on which responsibilities it absorbs. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[learner-identity]] — evolving disciplinary, professional, creative, and academic learner identities - [[self-directed-learning]] - [[self-regulated-learning]] - [[motivation]] - [[self-efficacy]] - [[student-experience]] - [[metacognition]] - [[ethics]] - [[cognitive-offloading]] - [[educational-robotics]] - [[behaviorism]] - [[framing-ai-use-for-students]] - [[human-ai-collaboration]] — shared direction of AI-mediated interaction - [[agentic-ai]] — autonomous AI that can redistribute agency in groups - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[social-emotional-learning]] — Social-Emotional Learning - [[cognitive-surrender]] ## Connected Articles - [[yan-agentivism-learning-theory-ai-2026]] — A mid-range learning theory for human-AI interaction, with four mechanisms and six testable propositions (Yan and Gašević 2026) - [[learner-agency-ai-simulation-2026]] — Access to choice vs. agency enacted: sliders, an optional AI agent and learning in a flocking simulation - [[powerful-learning-with-emerging-technology-2025]] — Agency as one of three design principles - [[guarded-adoption-genai-higher-education-2026]] — Guarded Adoption of Generative AI in Higher Education - [[ai-agents-joyful-assessment-third-space-2026]] — AI agents, joyful assessment, and third space - [[ssail-safe-sound-ai-learning-2026]] — SSAIL: A Design Framework for Safe and Sound AI for Learning - [[ai-mediated-research-agency-formation-2026]] — AI-mediated research agency formation in early-career scientific training - [[student-centered-genai-responsible-framework-2026]] — Student-facing framework for responsible GenAI use in higher education (Alsammani 2026) - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[school-ai-education-readiness-gaps-agency-2026]] — School AI education narrows psychological but not cognitive readiness gaps - [[reclaiming-epistemic-agency-co-agency-2026]] - [[du-yuan-epistemic-dependence-2026]] — Relational epistemic agency and six criteria separating productive reliance from harmful dependence (Du & Yuan 2026) - [[ai-pedagogical-accompaniment-amico]] — AI-enabled pedagogical accompaniment supporting STEM identity - [[t2i-competence-paradox-2026]] — The competence paradox: creative identity in text-to-image GenAI use - [[jin-emergent-learner-agency-implicit-hai-2026]] — Emergent learner agency in implicit human-AI collaboration: supportive vs. contrarian personas - [[de-barba-srl-genai-2026]] — Learner agency across scales: regulation, integration, positioning - [[mishra-control-vs-agency-history-2025]] — Control vs. agency as the essential tension in AIED history - [[idea-framework-metacognitive-genai-2026]] — The IDEA framework for metacognitively regulated GenAI use - [[daoism-ai-education-philosophy-2026]] — Alternative AI Philosophy: Daoism as Method for AI in Education - [[human-autonomy-agency-hri-review-2025]] — Human Autonomy and Agency in HRI - [[roboblockly-conversational-block-robotics-ct-2026]] — RoboBlockly Studio - [[teachy-mini-generative-social-robot-higher-ed-2026]] — Teachy Mini - [[knowledge-based-design-generative-social-robots-2026]] — Knowledge-Based Design for Generative Social Robots - [[andragogy-cognitive-delegation-genai-2026]] - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[young-people-learning-generative-ai-rapid-review-2026]] — Foster student agency in learning-relevant work - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and cognitive surrender: how AI reshapes human reasoning (Shaw & Nave 2026) - [[preservice-teacher-agency-genai-design-learning-2026]] — Pre-service teacher agency during GenAI interactions in design for learning (Krushinskaia, Elen & Raes 2026) - [[emancipatory-ai-learner-flourishing-2026]] — Emancipatory vision: agency as zero-sum with tool agency - [[trikonet-trivalence-co-creativity-2026]] — TriKoNet: co-creativity and agency in socio-technical networks - [[chen-zou-genai-group-assessment-agency-2026]] — Three patterns of agency in GenAI-mediated group assessment: intensified, restrained, and not enacted - [[ai-integrated-teaching-identity-tensions]] — Principled selectivity as teacher agency in AI-integrated teaching (Adiozaman & Segar 2026) - [[ai-refusal-higher-education-diagnostic-non-use-2026]] — Refusal as diagnostic evidence: non-use as a situated relation and the uneven right to decline AI (Zagami 2026) - [[obyrne-co-constructing-ai-boundaries-agency-judgment-2026]] — epistemic authority asserted through interruption, correction and refusal rather than maximal tool use --- ## [Learner Identity](https://edtechdev.github.io/aied/concepts/learner-identity/) > **Learner identity** — the evolving sense of who one is (and who one is becoming) as a learner, encompassing disciplinary, professional, creative, and academic identities. In [[ai-education|AI in education]], [[generative-ai|generative AI]] presses on learner identity in two directions at once: it can *support* identity formation ([[scaffolding]] disciplinary belonging and confidence) while also *threatening* it (undermining perceived authorship, competence, and authentic learning). Understanding learner identity is central to designing AI that affirms rather than erodes learners' sense of self. ## Questions to Consider - Think of a time you felt a skill or piece of work was truly 'yours' versus merely done through you. What made the difference — authorship, recognition, competence? How might a tool that produces the work for you reshape that feeling? - A common assumption is that identity is a fixed trait a learner either has or lacks. What would change in how you design learning if you treated learner identity instead as something continuously built through participation, recognition, and authorship? - The page describes a 'competence paradox' among art and design students: AI tools feel easy and useful, yet their use threatens the creative identity students derive from manual craft. Have you ever felt your own competence or identity challenged by an easy tool? What was the tension? - Students sometimes hide or feel shame about their AI use, which fragments their academic identity and honest engagement. When does the pressure to appear a certain kind of learner push people to conceal how they actually learn, and what would make disclosure feel safe? - AI can scaffold identity formation as well as threaten it. If you were designing an AI [[pedagogical-agent|learning companion]], what specific features would protect a learner's sense of authorship and ownership while still offering support? ## Introduction Learner identity concerns who a learner understands themselves to be, and who they are becoming, within a domain — a motivational and developmental construct distinct from ability beliefs such as [[self-efficacy]], which answer the narrower question *can I do this?* Identity forms through participation, recognition and authorship: seeing oneself reflected in a field and having that self-view validated by others. AI reshapes all three conditions, because it changes who does the work, what counts as one's own contribution, and whether a learner is recognized as the author of their own learning ([[agency]], [[metacognition]]). ## Why identity matters for AI in education Identity is a motivational and developmental construct distinct from (but connected to) related abilities and beliefs. Where [[self-efficacy]] concerns *can I do this?*, identity concerns *who am I — and who am I becoming?* It is built through participation, recognition, and authorship — through seeing oneself reflected in a domain and having that self-view validated. AI reshapes the conditions under which identity forms because it changes *who does the work*, *what counts as one's own contribution*, and *whether a learner feels recognized as the author of their learning*. This makes identity a first-order design concern rather than a peripheral "soft" factor. - **Authorship and competence under threat.** When AI produces text, images, or code, learners may question whether the result is truly "theirs" — a challenge to the authorship dimension of identity. **[[t2i-competence-paradox-2026|Liu et al. (2026)]]** document a *competence paradox* in art and design students using text-to-image GenAI: the tools feel easy and useful, yet their use simultaneously threatens the **creative identity** students derive from manual craft and authorship, producing a genuine tension between ease and self-worth. - **Shame and hidden use.** **[[shame-guilt-ai-regulation-computing-education|Lin et al.]]** show that computing students experience shame and guilt around AI use, which function as social regulators driving *hiding* and selective disclosure — behaviors that can fragment academic identity and undermine honest engagement with learning. - **Identity as something AI can scaffold.** AI need not only threaten identity. **[[ai-pedagogical-accompaniment-amico|Benedetti (2026)]]** argue that accountable, relationally-oriented [[pedagogy|pedagogical]] accompaniment can support learners' **STEM identity** development by providing transparent, bounded support that leaves room for the learner to own their trajectory. ## Identity in the knowledge base's research - **Creative identity:** **[[t2i-competence-paradox-2026|the T2I competence paradox]]** captures how ease-of-use can undermine the craft-based identity of art and design students. - **Professional identity:** multiple studies treat AI's impact on **professional identity** — for example, **[[lodge-adaptive-capabilities-genai-future-2026|Lodge et al. (2026)]]** argue that graduates need *adaptive capabilities* ([[ai-literacy]], [[distributed-cognition]], [[metacognition]]) precisely so they can sustain a viable professional identity in an AI-integrated future, rather than being defined by — or defined out by — their tools. - **Post-human and hybrid identity:** **[[elsayed-pedagogical-symbiosis-posthuman-learner|Elsayed (2026)]]** theorize the **post-human learner**, whose cognition is genuinely hybrid and distributed across [[biology-education|biological]] and artificial systems — a reframing of identity formation itself in the age of cognitive AI. - **Student and academic identity:** **[[zhan-boud-du-authentic-assessment-scoping-review-2025|authentic assessment]]** [[research-methods-aied|research]] connects to identity because [[assessment]] tasks that call for authentic, personal performance help students see themselves as capable practitioners; **[[paternalistic-filter-llm-history-education|history-education research]]** shows how paternalistic AI use can shape how students construct their identity as disciplinary inquirers. - **Guarded adoption: selective AI use as identity protection.** [[guarded-adoption-genai-higher-education-2026|Zagami (2026)]] surveyed 484 students at one Australian university and found that higher-achieving students reported *lower* active [[generative-ai|generative AI]] engagement, lower positive affect toward AI, lower [[self-report-measures|perceived learning]] impact and lower AI-related disengagement, with the strongest associations at rho = -0.395 for perceived learning impact and rho = -0.359 for positive affect. Their open-ended responses described use that was selective (clarification, summarization, workflow support), verified, and held subordinate to their own judgment, and item-level results showed greater agreement that reliance on AI hinders [[critical-thinking]] and independent [[problem-solving|problem solving]]. The authors read this as identity work: for students whose sense of themselves as successful learners rests on their own effort and judgment, bounding AI use defends the [[agency|epistemic agency]] the identity depends on — which also means the pattern is neither technophobia nor low engagement. ## Relationship to learner agency Learner agency and learner identity are closely related but distinct constructs that are easy to conflate — and both are central to how AI affects learning. - **Agency is about *doing*; identity is about *being*.** [[agency|Learner agency]] concerns the capacity to act intentionally, make choices, and exercise control over one's learning *in the moment* — a situated, interactional, and variable capacity. Learner identity concerns who one *is* and is *becoming* as a learner — a more durable, narrative, and developmental sense of self. Agency asks "am I able to direct this?", while identity asks "is this who I am / who I want to be?" - **They are causally intertwined.** Agency is both a *source* and an *outcome* of identity. Enacting agency — choosing, authoring, persisting — is how a learner comes to see themselves as an agentic person (identity is partly internalized agency). Conversely, a stable disciplinary or professional identity supplies the motivation and self-worth that sustain agency under difficulty. Identity is the accumulated product of repeated agentic acts; agency is the ongoing enactment that builds identity. - **AI threatens them through different mechanisms.** AI can erode **agency** by inviting [[cognitive-offloading|over-reliance]] and passive acceptance — learners stop directing their own reasoning. AI can erode **identity** by undermining authorship and competence — when AI produces the work, learners may stop feeling the output is "theirs" or that they belong in the domain. The [[t2i-competence-paradox-2026|competence paradox]] is an identity threat; [[jin-emergent-learner-agency-implicit-hai-2026|implicit AI redistribution of epistemic labor]] is primarily an agency threat, though it compounds into identity over time. - **Safeguarding both is the design goal.** Supporting agency means preserving learners' control and choice (e.g., bounded [[desirable-difficulties|friction]], [[human-in-the-loop-ai|human-in-the-loop]] oversight). Supporting identity means protecting authorship and recognition (e.g., [[authentic-assessment|authentic assessment]], transparent attribution of AI versus human contribution). A design that protects agency but not authorship protects control without protecting the sense of self — and vice versa. In short: **foster agency to let learners act; sustain identity so they know who they are while acting.** Healthy AI-supported learning attends to both. ## Relationship to teacher identity Learner identity is the **student-facing** counterpart to teacher identity. Teacher identity — the evolving professional self-understanding of educators — is documented on the [[teacher-role]] page, where [[laidlaw-genai-identity-crisis-faculty-2026|Laidlaw (2026)]] frame GenAI as an *identity crisis* for faculty rather than merely a skills gap, and [[teaching-the-teachers-genai-tpk-review-2026|TPK-based teacher training]] treats identity as part of professional preparation. The two are reciprocal: teachers who experience identity disruption are less able to support their students' identity development, so a healthy AI-integrated system must attend to both. ## Connections Learner identity connects to [[agency]] (identity is enacted through agentic authorship), [[self-efficacy]] (competence beliefs that sustain identity), [[motivation]] and [[self-determination-theory]] (identity formation satisfies needs for autonomy and competence), [[student-experience]] and [[student-engagement]], and [[stem-education]] (where disciplinary identity is a key outcome and predictor of persistence). It is also shaped by [[situated-learning]] and [[critical-pedagogy]] (identity as participation and as power-laden negotiation) and connects to [[authentic-assessment]] (tasks that let learners perform and thus claim an identity). Its [[higher-ed]] and [[k-12]] relevance spans both schooling and professional preparation. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[agency]] - [[self-efficacy]] - [[motivation]] - [[self-determination-theory]] - [[student-experience]] - [[student-engagement]] - [[stem-education]] - [[situated-learning]] - [[critical-pedagogy]] - [[authentic-assessment]] - [[teacher-role]] - [[generative-ai]] - [[higher-ed]] - [[k-12]] ## Connected Articles - [[guarded-adoption-genai-higher-education-2026]] — Guarded Adoption of Generative AI in Higher Education - [[t2i-competence-paradox-2026]] — The competence paradox: creative identity in text-to-image GenAI use - [[shame-guilt-ai-regulation-computing-education]] — Shame and guilt as social regulators of AI use - [[lodge-adaptive-capabilities-genai-future-2026]] — Adaptive capabilities for a viable professional identity in a GenAI future - [[elsayed-pedagogical-symbiosis-posthuman-learner]] — Pedagogical symbiosis and the post-human learner - [[ai-pedagogical-accompaniment-amico]] — AI-enabled pedagogical accompaniment supporting STEM identity - [[zhan-boud-du-authentic-assessment-scoping-review-2025]] — Designing for authentic assessment - [[paternalistic-filter-llm-history-education]] — Paternalistic AI use and student identity in history education - [[laidlaw-genai-identity-crisis-faculty-2026]] — GenAI as identity crisis for faculty (teacher identity) - [[teaching-the-teachers-genai-tpk-review-2026]] — TPK-based teacher training and professional identity - [[li-ai-science-situated-learning-teachers-2025]] — Science teachers' roles and learner identity in situated learning - [[genai-professionalization-metaphors-2026]] — GenAI conceptualizations and student professional identity --- ## [Design Thinking](https://edtechdev.github.io/aied/concepts/design-thinking/) > **Design Thinking** — a key concept in [[ai-education|AI in education]] research: a human-centered, iterative problem-solving process (typically Empathize → Define → Ideate → Prototype → Test) that moves learners from understanding a problem to producing and refining a solution. Explored across 8 articles in this knowledge base. ## Questions to Consider - Design thinking follows Empathize → Define → Ideate → Prototype → Test. In your own problem-solving, which stage do you naturally skip — and what might that cost you? - Generative AI excels at early-stage ideation but can cause 'design fixation' or aesthetic lock-in if used uncritically. Have you ever latched onto a first AI suggestion and struggled to see alternatives? - One study found that an adversarial AI that challenged designers produced more iterations, broader exploration, and better final designs — but was frustrating to use. When is friction with an AI genuinely productive? - Students who used GenAI heavily in design reported it did NOT reduce their sense of ownership or creativity. Does that contradict the worry that AI erodes authorship — or does it depend on how the tool is orchestrated? - Design thinking appears both as a skill being taught and as a method educators use to build AI-supported learning. Which role is more relevant to your work, and how does that change the design principles you'd apply? - Practitioners often use AI for iterative prompting and generation but underuse needs assessment and feedback loops. What does it mean to treat AI as a 'fallible co-intelligent collaborator' rather than a content generator? ## Introduction Design thinking sits at the intersection of creativity, craft, and critique — and it is increasingly where [[generative-ai]] and [[agentic-ai]] are reshaping how students learn to design. Across this knowledge base's connected articles, design thinking appears in two distinct roles: as a *pedagogical object* (the skill being taught and measured) and as a *pedagogical method* (the process educators themselves use to build AI-supported learning). In both roles, the same tension recurs: generative tools can accelerate ideation and broaden participation, but their value depends on how they are orchestrated, on the [[feedback]] they provide, and on whether the learner retains a genuine sense of [[agency]]. ## Evidence from the knowledge base - **GenAI in design studios and architecture.** Two architecture-focused studies ground the concept in studio pedagogy. [[genai-architecture-education|Gen-AI-tecture]] found that a locally executed, discipline-specific tool enhanced creative fluency, broadened participation across diverse learner profiles, and strengthened confidence in AI-supported workflows — an important signal for [[equity-in-ai-education]] as visual-spatial design becomes more accessible. [[genai-architectural-design-studios|GenAI in architectural design studios]] likewise found students using GenAI as visual stimuli and inspirational resources in early ideation, and as a combinatorial method for expanding the solution space during development. Crucially, both warn that without critical discernment students risk **design fixation** or aesthetic lock-in, reframing the [[teacher-role]] in [[higher-ed]] from transmitting craft to coaching how, when, and why to delegate creative exploration to generative models. [[creativity|Creativity]] is thereby preserved and even amplified — but only under careful pedagogical orchestration. - **Adversarial AI agents prompting reconsideration.** [[ai-agents-constructive-conflict-design-education-2026|Constructive conflicts with AI agents]] takes design thinking toward [[agentic-ai]], testing adversarial versus cooperative AI roles with novice interaction designers (N=48). The adversarial condition produced significantly more design iterations, broader exploration of alternatives, and higher-rated final designs — participants found the conflict agent frustrating but ultimately helpful. This productive-friction dynamic, echoing the [[socratic-method]] and [[scaffolding]], shows that design thinking benefits from challenge rather than mere assistance, a finding with implications for how [[student-experience]] is designed in AI-augmented studio settings. - **Design thinking as a student skill being developed.** [[genai-usage-design-students-survey|GenAI usage by design students]] at Politecnico di Milano reported very frequent use concentrated in early, ideation-heavy stages — yet high adoption did not reduce perceived project ownership or creativity, directly informing [[ai-literacy]] and [[academic-integrity]] debates about authorship and process transparency. - **AI in design pedagogy and educator frameworks.** Design thinking also structures how educators build AI-supported learning. [[gaide-vibe-coding-k12-teachers|GAIDE]] offers a Design-Thinking-based framework for K-12 teachers creating AI-powered tools through vibe coding, raising teachers' AI literacy and supporting learning-by-creating as [[educational-development]] and [[professional-training]]. [[dot-framework-survey-2026|The DOT Framework survey]] (n=72) grounds design thinking in open-systems theory and found practitioners frequently use iterative prompting and content generation but underuse needs assessment and feedback loops — a practice-versus-theory gap that operationalizes AI as a fallible co-intelligent collaborator under [[human-in-the-loop-ai]] oversight. Even [[social-robot-study-companions|co-creating open social robots]] with students applied the Double Diamond (a design-thinking variant), redesigning the build around accessibility so that repairability and [[open-source]] principles become sites of continued learning. ## Practical guidance For educators, the collective evidence counsels against blanket restriction or unfettered adoption. Treat GenAI as an *ideation resource whose value is orchestration-dependent*: it excels at early-stage exploration and empathise work, while human instructors add contextual anchoring in prototyping and evaluation. When deploying [[agentic-ai]] agents, consider adversarial or Socratic roles that prompt reconsideration rather than smooth cooperative confirmation. And for designers of teacher-facing support, remember that beliefs alone are not enough — practitioners need scaffolding for the full design cycle (needs assessment, feedback loops), not just tool usage, and evaluation instruments to verify that design thinking is actually being developed (see [[ai-ed-evaluation]]). ## Connections to related concepts Design thinking in AI education is deeply entangled with [[human-ai-collaboration]]: generative tools broaden ideation while humans retain epistemic authority and [[agency]]. It relies on [[scaffolding]] and stage-appropriate [[feedback]] to be effective, is often delivered through [[socratic-method|Socratic]] and adversarial interactions in [[higher-ed]] and [[k-12]], and its success is frequently framed as preserving [[creativity]] and fostering [[ai-literacy]] under [[human-in-the-loop-ai]] governance. ## Connected Concepts - [[ai-education]] - [[generative-ai]] - [[agentic-ai]] - [[higher-ed]] - [[student-experience]] - [[scaffolding]] - [[feedback]] - [[socratic-method]] - [[creativity]] - [[agency]] - [[human-ai-collaboration]] - [[teacher-role]] - [[ai-literacy]] - [[equity-in-ai-education]] - [[educational-development]] - [[human-in-the-loop-ai]] - [[ai-ed-evaluation]] - [[k-12]] - [[professional-training]] - [[academic-integrity]] - [[open-source]] - [[arts-design-and-media-education]] ## Connected Articles - [[rana-genai-design-thinking-2025]] - [[genai-architecture-education]] — Gen-AI-tecture: using generative AI to support architectural students in design tasks - [[genai-architectural-design-studios]] — Development and applications of Generative AI in architectural design studios - [[social-robot-study-companions]] — Co-Creating Buildable and Open Social Robot Study Companions with University Students - [[gaide-vibe-coding-k12-teachers]] — A Guiding Framework for K-12 Teachers in Creating AI-powered Learning Technologies through Vibe Coding - [[dot-framework-survey-2026]] — DOT Framework Survey: Practitioner Beliefs and Behaviors in AI-Enhanced Education - [[ai-agents-constructive-conflict-design-education-2026]] — Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction Designers - [[genai-usage-design-students-survey]] — A study of GenAI usage by Design Students: Analysis of Survey Results and Journals of AI practices at the Politecnico di Milano in 2025/2026 --- ## [Curriculum Design](https://edtechdev.github.io/aied/concepts/curriculum-design/) > **Curriculum Design** — the process of planning and structuring what is taught across courses, programs, and institutions, including learning objectives, content sequencing, assessment strategies, and skill progression. In the AI era, curriculum design must balance foundational knowledge with emerging AI competencies, determining not just what students learn but how they learn to work with and critically evaluate AI tools. ## Questions to Consider - Curriculum design asks what students should learn at the program level, while learning design asks how at the course level. When AI reshapes a discipline, which of these two layers do you think should change first? - As [[generative-ai|generative AI]] automates implementation-level work, some argue curricula must shift toward system design, abstraction, and [[critical-thinking|critical evaluation]]. What do you think students would lose if low-level skills were de-emphasized? - Curriculum redesign in the AI era is often framed as a balance between tool fluency and foundational knowledge. Where have you seen that balance tip too far in one direction? - The SAIL framework treats AI literacy as scaffolded across ages and designed to address deeper 'digital divides' beyond access. How is embedding AI literacy across a whole curriculum different from adding a single AI course? - If every discipline now needs AI competencies embedded within it, who is responsible for the curriculum change — instructors, programs, or institutions — and what do educators need to succeed at it? - A curriculum is a sequence of skills across years, not just a list of topics. How does that longer view change whether an 'AI literacy unit' actually sticks? ## Introduction Curriculum design addresses the *what* of education at the program level, complementing [[learning-design]] which addresses the *how* at the course level. The articles in this knowledge base explore how AI is reshaping curricula across disciplines — from software engineering to architecture to green education — and how educators are designing curricula that embed AI literacy without sacrificing disciplinary fundamentals. ### Key research themes **Redesigning curricula for the AI era** is the central challenge. **[[reshaping-cs-education-genai|Lee et al.]]** synthesized findings from international workshops on reshaping undergraduate [[cs-education|CS education]], arguing that as GenAI automates implementation-level programming, curricula must shift toward system design, abstraction, and critical evaluation — while de-emphasizing low-level implementation details. **[[ase-26-agentic-software-engineering-curriculum|Gorsky]]** formalized Agentic Software Engineering as a distinct discipline with a 21-module curriculum focused on the "evolution of intent" and practitioner discipline required to manage [[agentic-ai|AI agents]]. Both connect to [[ai-literacy]] and [[scaffolding]]. **Curriculum mapping and analysis** uses AI to understand existing curricula. **[[ai-assisted-se-curriculum-syllabus-analysis-2026|Geng et al.]]** analyzed 23 syllabi from AI-assisted software engineering courses, identifying common themes — [[prompt-engineering|prompt engineering]], code review with AI, [[ethics|ethical considerations]] — and deriving design guidance that emphasizes balancing tool fluency with foundational knowledge. **[[coursegraph-cs-course-comparison-2026|CourseGraph]]** applies computational methods to compare CS course structures across institutions. **AI literacy integration** embeds AI competencies across disciplines. **[[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|SAIL]]** provides a scaffolded AI literacy framework applicable across all ages and educational stages, addressing second- and third-level [[digital-divide|digital divides]]. **[[tracing-genai-literacy-interaction-patterns]]** examines how AI literacy develops through interaction patterns. **[[hingle-collaborative-ai-literacy-2025]]** explores collaborative approaches to AI literacy curriculum development, connecting to [[collaborative-learning]]. **[[discipline-specific-aied|Domain-specific]] curriculum innovation** applies curriculum design to specific fields. **[[genai-architecture-education]]** explores how generative AI reshapes architectural design [[pedagogy]]. **[[talebzadeh-ai-green-education-2026]]** examines AI integration in green education curricula. **[[connected-ai-lesson-planning-vietnam]]** and **[[llm-cultural-relevance-k12]]** address [[culturally-relevant-pedagogy|culturally responsive curriculum design]]. **[[governance|Institutional]] frameworks** address curriculum change at scale. **[[finkelstein-principled-ai-education-2025]]** and **[[finkelstein-principled-ai-education-2025]]** provide principles for integrating AI across educational programs. **[[ai-adoption-training-public-sector]]** examines barriers to AI curriculum adoption in public sector education. **Sequencing AI across the program.** [[refrain-amplify-genai-curriculum-2026|Torres-Sahli et al.]] propose a "refrain, then amplify" framework that sequences generative AI at the program level: withhold a generative tool while a capacity is forming, then restore it to amplify that capacity once the student can direct it and judge its returns. Governed by a forming-versus-offloading criterion (whether a stretch of work builds a capacity or merely passes it through the tool), the framework links curriculum design to [[cognitive-offloading]], [[self-regulated-learning]], and [[academic-integrity]], with hard-to-fake checkpoints at each refrain-to-amplify hinge. **Whole-course alignment when generative AI is permitted.** A 2026 redesign of an introductory nuclear and particle [[physics-education|physics]] course ([[ai-particle-physics-education-redesign-2026|Mikhasenko et al.]]) integrated three activity types with distinct roles — lectures for concepts and notation, tutorials for standard analytic practice, and homework as an exploratory "research-shaped" component of unusually difficult, multi-method problems. The reported friction (an undeclared programming prerequisite, insufficient time to understand rather than merely obtain answers, and misalignment among lectures, tutorials, homework and [[summative-assessment|examination]]) illustrates that permitting generative AI forces curriculum alignment work across the whole course rather than a change to one assignment type; their recommended structure keeps the AI-permitted exploratory work as bonus-bearing advanced tasks while the unaided written exam determines the grade. **Constructive alignment advice under critical scrutiny.** Where the sources above report on alignment in practice, [[mcinnes-salvaging-constructive-alignment-genai-2026|McInnes et al. (2026)]] read the guidance itself as a discourse. Their [[qualitative-research|critical discourse analysis]] of 14 pieces of gray literature published from November 2022 to April 2025 — mostly institutional pages and chapters from centrally positioned [[educational-development|learning and teaching units]] — found constructive alignment presented as an efficiency problem: generative AI was "an effective and efficient way to draft rubrics" that could "streamline the process", anthropomorphized as "an educational expert and assistant", a "sparring partner" or an "intelligent assistant in instructional design", while academic staff supplied only "subject matter expertise" and the tool took on "the heavy lifting of developing learning objectives, organizing course content ... and aligning course components". Copy-and-paste prompt recipes and numbered templates framed CA as a standardisable product, producing three failure modes: performativity (alignment that only looks aligned), erasure of situated and critical context, and shallow CA that conflates alignment with the constructive dimension. Their remedy re-sequences the curriculum-design workflow — educators must understand CA well enough to direct, evaluate and reject AI output before delegating any part of it — and bounds any tool to an institutional [[rag|retrieval-augmented]] agent grounded in local policy, rubrics and graduate attributes, with a "liminal tutor" role that extends rather than replaces the [[educational-development|developer]] relationship. **Generating curriculum-aligned modeling tasks.** AI-powered platforms can address teachers' lack of time and resources for designing high-quality [[math-education|mathematical modeling]] tasks by generating curriculum-aligned problems and pedagogical recommendations grounded in design principles and [[rag|retrieval-augmented]] generation — an approach illustrated with direct variation in secondary school mathematics ([[ai-modeling-problem-generation-platform-2026]]). Course readings themselves are now a generation target too: Sidorkin (2026) replaced a commercial textbook with weekly AI-generated readings in a graduate educational leadership course, and although students rated them useful and 75 percent agreed they learned more than in a comparable course, the 4,487 pages of logs carried APA-style in-text citations on only about 0.80 percent of pages and paired a named campus or system with assertive policy claims on roughly 1.03 percent of pages without a verifiable source. The curriculum-materials lesson is to treat generated readings as draft production under instructor review, to budget for the instructor labor of prompt design and verification, and to curate vetted sources into the assistant rather than leaving source quality for students to infer from context. **AI-assisted lesson planning at the curriculum-into-classroom layer.** At the point where a curriculum becomes a taught lesson, [[luo-tahir-chatgpt-steam-lesson-planning-2026|Luo and Tahir (2025)]] experimentally compared teacher-generated versus ChatGPT-assisted plans in children's [[stem-education|STEAM]] arts education, finding AI-assisted plans rated significantly higher by six expert professors (median 20.5 vs. 17.6, p = .002, large effect). They show the payoff depends on how the teacher delegates: the recommended method fills content gaps in a self-outlined lesson (preserving teacher design [[agency|autonomy]]) rather than delegating the whole plan, and they contribute a Role–Instructions–End Goal prompt template for reproducible, quality-controlled generation — evidence that AI lesson planning is strongest when embedded within, not substituted for, the teacher's curriculum decisions. In science education, expert validation reaches a parallel verdict on platform design: [[karaismailoglu-ai-lesson-plans-science-experts-2026|Karaismailoglu, Surmeli and Yildirim (2026)]] had eleven [[science-education]] specialists rate ChatGPT-4 and an education-focused tool (Teacher's Buddy) on sixth-grade plans aligned to Turkey's revised curriculum and the Engineering [[design-based-research|Design-Based]] Learning model. The education-focused platform scored higher across all eight quality criteria — including feedback-intensive stages and curriculum alignment — evidence that embedding pedagogical structure into an AI yields better-aligned output; yet some experts still preferred the general-purpose plan for its stronger social-emotional emphasis, and 7 of 11 judged the plans "applicable by correction" rather than directly usable. The choice of platform and prompt framing, not just the AI itself, shapes how well generated plans align to curriculum standards and process models. ### Connections to related concepts Curriculum design connects directly to [[learning-design]] — curriculum defines what, instruction defines how. It connects to [[ai-literacy]] because embedding AI competencies is a primary curriculum challenge, to [[teacher-role]] and [[educational-development]] because curriculum change requires educator preparation, and to [[scaffolding]] because well-designed curricula scaffold skill development across courses and years. The [[higher-ed]] and [[k-12]] connections reflect curriculum design's relevance across educational levels. ## Connected Concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[business-education]] - [[learning-design]] - [[ai-literacy]] - [[scaffolding]] - [[educational-development]] - [[teacher-role]] - [[higher-ed]] - [[k-12]] - [[stem-education]] - [[cs-education]] - [[generative-ai]] - [[agentic-ai]] - [[metacognition]] - [[prompt-engineering]] - [[collaborative-learning]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[mcinnes-salvaging-constructive-alignment-genai-2026]] — Critical discourse analysis of GenAI-for-constructive-alignment guidance (McInnes et al. 2026) - [[icet-ml-education-trust-2026]] — Addressing Trust in AI Systems through Education: A Didactic Perspective - [[refrain-amplify-genai-curriculum-2026]] — Refrain-then-amplify curriculum framework for sequencing GenAI (Torres-Sahli et al. 2026) - [[mechanical-engineering-ai-curriculum-2026]] — Project-Based AI Education Curriculum in Thermal Engineering - [[ying-genai-journalism-assessment-2026]] - [[rook-plumb-genai-curricula-student-insights-2026]] - [[zhou-constructive-alignment-genai-business-2026]] - [[nicola-richmond-programwide-assessment-genai-2025]] - [[espino-ai-business-education-review-2026]] - [[drummond-genai-business-schools-framework-2026]] - [[workforce-readiness-smart-manufacturing-wrl-2026]] — Workforce Readiness Level framework for smart manufacturing in the AI era - [[rewriting-curriculum-genai-pedagogy-2026]] — Rewriting the curriculum: GenAI-driven pedagogical change - [[ai-interior-design-malaysia-2026]] - [[critical-media-literacy-education-2026]] - [[ai-generated-interactive-fiction-education-2026]] - [[reshaping-cs-education-genai]] - [[ase-26-agentic-software-engineering-curriculum]] - [[ai-assisted-se-curriculum-syllabus-analysis-2026]] - [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl]] - [[curriculum-as-code-instructional-design-2026]] - [[tracing-genai-literacy-interaction-patterns]] - [[finkelstein-principled-ai-education-2025]] - [[hingle-collaborative-ai-literacy-2025]] - [[learnity-graphs-lifelong-learning-framework-2026]] - [[panciroli-ai-literacy-episodes-situated-learning]] - [[ithaka-sr-ai-skills-college-graduates-2026]] — AI Skills Framework: 26 assessable skills for curriculum mapping - [[learnai-just-in-time-ai-cocreation-university-2026]] — LearnAI: Just-in-Time AI Co-Creation Across Disciplines - [[ai-video-dual-gatekeeping-2026]] — When Saying No Makes Better Videos: Dual Gatekeeping for Pedagogically Grounded AI Content Creation - [[niri-steam-ai-literacy-review-2026]] — STEAM education for AI literacy: systematic review - [[ai-writes-code-student-writes-model-2026]] — Model authorship: theory & measurement for learning-by-construction with GenAI - [[code-to-learn-genai-artifact-construction-2026]] — CtL-GenAI: constructionism framework for artifact construction - [[caruana-pre-university-ai-education-slr-2026]] — Preparing learners and teachers for an AI-driven future: SLR of pre-university AI education (Caruana et al. 2026) - [[cogevol-learning-environment-generation-2026]] — CogEvol: Learning Environment Generation - [[ai-digital-transformation-liberal-arts-lingnan-2026]] — Digital transformation of a liberal arts university toward a research-intensive model in the GenAI era (Qin 2026) - [[ai-modeling-problem-generation-platform-2026]] — AI-powered platform generating mathematical modeling problems (ADDIE, RAG) - [[luo-tahir-chatgpt-steam-lesson-planning-2026]] - [[karaismailoglu-ai-lesson-plans-science-experts-2026]] - [[ai-particle-physics-education-redesign-2026]] — AI in Particle Physics Education: Research Problems and Foundational Skills - [[sidorkin-ai-generated-course-readings-2026]] — AI-generated weekly readings as a textbook substitute, with sourcing and review caveats (Sidorkin 2026) --- ## [Critical Thinking](https://edtechdev.github.io/aied/concepts/critical-thinking/) > **Critical thinking** — the ability to analyze, evaluate, and synthesize information — is both a skill that AI tools can help develop and a competency that students must apply when using AI. In [[ai-education|AI in education]] [[research-methods-aied|research]], critical thinking appears in two interrelated forms: as a learning objective ([[teacher-role|teaching]] students to think critically) and as a safeguard against uncritical AI reliance. ## Questions to Consider - How confident are you in your ability to spot a false or misleading AI-generated answer? Research suggests [[self-report-measures|self-reported]] AI competence far exceeds actual evaluation ability — how would you test yourself? - Critical thinking here appears in two forms: a skill to teach, and a safeguard against uncritical reliance on AI. Can you think of a situation where a tool that 'teaches' critical thinking is actually training its opposite? - One study found that having students interrogate AI-generated mistakes produced large gains in higher-order thinking. How might deliberately exposing errors — rather than hiding them — be a more powerful teaching move than you assumed? - Easy access to AI answers can displace critical [[student-engagement|engagement]] before students realize it. What design feature, rather than a policy or a ban, could keep the cognitive effort alive? - AI advice has been shown to suppress the willingness to say 'I don't know' — even when the advice is wrong. How does that change what it means to create a classroom culture where questioning is safe? ## Introduction Critical thinking is central to [[ai-literacy]] — students who cannot critically evaluate AI outputs are vulnerable to [[cognitive-offloading|Over-Reliance]], [[hallucination-risk|hallucinated information]], and biased recommendations. Research on [[cognitive-offloading]] shows that easy access to AI answers can displace critical engagement, while [[socratic-method|Socratic approaches]] that withhold direct answers preserve the cognitive effort necessary for deeper thinking. ### Critical thinking in AI education research The knowledge base's articles explore critical thinking through [[design-based-research|design-based]] and empirical lenses. [[ai-agents-constructive-conflict-design-education-2026|Adversarial AI agents]] enact constructive conflict to prompt reconsideration in novice designers — a Socratic variant that forces critical re-evaluation. [[genai-can-harm-teaching-rct-2026|RCT research on GenAI in teaching]] raises the question of whether AI tools that optimize for surface-level outcomes may inadvertently suppress the critical thinking that leads to deeper learning. [[chatgpt-critical-creative-thinking-review|Reviews of ChatGPT's impact on thinking]] document mixed findings: AI can [[scaffolding|scaffold]] critical analysis when used deliberately (e.g., asking students to critique AI-generated arguments), but it can also short-circuit thinking when used as an answer engine. This tension connects to [[ai-literacy-assessment-misalignment]] research showing that self-reported AI competence far exceeds actual critical evaluation ability. - **Higher-order cognitive engagement in student-AI chat.** Chang and Li (2026) find that ~62% of student prompts to AI encode higher-order cognitive demand, with Bloom-level profiles varying by discipline ([[stem-education|STEM]] Apply-prevalent 20.8%, language Understand-prevalent 31.7%, social science Create-prevalent 33.8%). Their within-person design shows the same students produce significantly more higher-order prompts in social science than STEM courses (p < .001), indicating that disciplinary context shapes critical and higher-order engagement with AI. - **AI as a catalyst for critical media literacy in children.** Demir and Akar (2026) evaluate an 18-hour, 5E-model critical media literacy program for fourth-grade Turkish students in which [[generative-ai|generative AI]] (ChatGPT, Grammarly) acted as a [[pedagogical-agent|pedagogical agent]] embedded phase-by-phase rather than an add-on. Paired-samples comparisons showed large gains in media reading (+3.50), writing (+1.67), and total media literacy (+5.17, all p < .01), with between-group post-test effect sizes of Cohen's *d* = 1.12 (reading), 1.18 (writing), and 1.31 (total literacy) favoring the AI-supported group. [[qualitative-research|Qualitative]] analysis (interviews, student posters/drawings/slogans, classroom observation) surfaced six domains of critical media literacy growth — digital self-protection and [[privacy|data privacy]], purposeful and responsible media use, safe communication and boundary awareness, critical evaluation and misinformation awareness, online risk awareness, and media [[ethics]]/digital citizenship — indicating that deliberately interrogating AI-mediated content can cultivate critical analysis and reflection in young learners. - **Dimension-specific critical-thinking gains in primary multimodal writing.** [[lu-ai-multimodal-writing-critical-thinking-2026|Lu et al. (2027)]] followed 60 [[k-12|Grade 5]] students through an eight-week [[conversational-ai]]-supported multimodal writing practice in which they turned narratives into AI-generated images and short videos. Repeated-measures analysis across six critical-thinking dimensions found sustained gains (T1→T2 and T1→T3) in interpretation, analysis, evaluation, and explanation, a short-lived self-[[regulation]] gain, and **no change in inference** — an uneven, dimension-level pattern that an aggregate critical-thinking score would have hidden. The authors argue the AI-generated visuals *externalized* meaning and thereby lowered the inferential demand writing normally imposes, while [[collaborative-learning|peer collaboration]] (peer questions that forced inferring others' interpretations) supplied the occasions for inference the solo [[student-ai-interaction|AI interaction]] did not. The design lesson: [[multimodal|multimodal AI]] composing supports several critical-thinking facets but should be paired with continued [[scaffolding]] and structured peer exchange to preserve inference and [[self-regulated-learning|self-regulation]]. - **AI scaffolding and offloading pull critical thinking in opposite directions.** Davor, Larbi and Boateng (2026) surveyed 533 university students in Ghana and found that AI task scaffolding predicted higher critical thinking (β = .185) while [[cognitive-offloading|cognitive offloading]] tendency predicted lower critical thinking (-.240); AI verification literacy had no direct effect on critical thinking and worked only through [[metacognition|metacognitive self-regulation]], a full mediation pattern the authors read as evidence that teaching students to fact-check AI is not enough on its own. ([[davor-ai-supported-learning-higher-order-outcomes-2026|Davor et al. 2026]]) - **Dependence, not use, is where the association with critical thinking turns.** Shojaei and colleagues (2026) surveyed 412 business students in Oman and found a near-zero bivariate correlation between [[generative-ai|GenAI]] use and self-reported critical-thinking disposition (r = 0.050), with dependence predicting lower disposition (β = -0.389) and weakening the link from use to disposition (β = -0.239), so that the simple slope fell from 0.424 at low dependence to -0.054 at high dependence. ([[shojaei-genai-dependence-critical-thinking-employability-2026|Shojaei et al. 2026]]) - **A short reflection prompt makes reliance on AI advice more discriminative.** In a three-condition experiment with 342 undergraduates, Ren (2026) found that open ChatGPT support produced acceptance of incorrect AI recommendations on 62.4% of trials, falling to 39.7% with a brief metacognitive reflection prompt (OR = 0.40, 95% CI [0.28, 0.56]); reflection also improved awareness calibration (0.59 vs. 0.41) and cut the AI-specific attribution bias index from 0.42 to 0.21 without reducing recommendation accuracy or triggering blanket rejection of useful advice. ([[ren-metacognitive-awareness-genai-reliance-2026|Ren 2026]]) ### Connections to other concepts Critical thinking intersects with [[scaffolding]] (designing AI support that maintains cognitive demand), [[prompt-engineering]] (formulating questions that elicit critical analysis), and [[cognitive-offloading|Over-Reliance]] (knowing when to trust and when to question AI). It is foundational to [[academic-integrity]] and serves as a key dimension of [[ai-literacy]] frameworks across both [[k-12]] and [[higher-ed]] contexts. - **AI errors as provocations for higher-order thinking:** [[pedagogy-ai-mistakes|Hosseini (2026)]] operationalizes Bloom's higher-order levels (Analyze, Evaluate, Create) by having students interrogate AI-generated mistakes in a database course, with significant pre/post gains (Cohen's *d*=1.49) in subject-matter competency. - **Two-sided auditing of AI explanations.** Bernstein and Sibia (2026) used Paul–Elder standards (accuracy, clarity, assumptions, point of view) as interview probes with ten students who had completed CS2, and found mechanism-level scrutiny of [[generative-ai|GenAI]] explanations: students located where an analogy's mapping broke (an island-route analogy for a linked list that implied a circle, a badminton rally offered for recursion that had no guaranteed shrinking input), demanded precise wording over hedging, and treated explanations as arguments carrying a point of view. Crucially, that scrutiny tracked source- or target-domain expertise rather than personal interest — reframing critical [[ai-ed-evaluation|evaluation of AI]] output as a knowledge problem ("two-sided analogy auditing") rather than a dispositional one — and suggesting that assigning flawed AI analogies as objects to inspect and repair is a harder check on conceptual understanding than reading a finished explanation.([[student-reception-genai-analogies-computing-2026]]) - **Critical thinking as non-outsourceable engagement.** Xie (2026) adds a [[philosophy-of-ai-in-education|philosophical]] counterpart from Daoist self-cultivation: because AI is an opaque "black box," a framework oriented to harmonizing uncertainty rather than adjudicating truth in absolute terms better suits the current epistemic landscape, and critical thinking becomes sustained, first-person, non-outsourceable engagement with reality rather than a demonstrable rational procedure. In Neidan (內丹) practice "there are no cognitive shortcuts," so AI is positioned as "not a cognitive surrogate but an instrumental adjunct" to human flourishing.([[daoism-ai-education-philosophy-2026]]) - **Role rotation as a structure for critical human-AI interaction.** Kenzhebayeva and colleagues (2026) report a design-based study in which 62 pre-service educational psychologists rotated through four professional roles (Case Constructor, Research Analyst, Practitioner-Interventionist, Reflective Researcher) that made generated recommendations the object of discussion: participants compared AI output with psychological theory and modified or rejected recommendations that did not fit the case, and later cycles showed more requests for theoretical justification, while overreliance on apparently authoritative responses persisted. ([[kenzhebayeva-ai-role-rotation-pedagogical-model-2026|Kenzhebayeva et al. 2026]]) - **Verification-centered integration in a discipline.** A critical review of [[generative-ai|generative AI]] in university [[chemistry-education|chemistry education]] (Vega-Baudrit and Rivera Álvarez, 2026) argues that because chemical reasoning must be coordinated across macroscopic, submicroscopic, and symbolic representations, students cannot verify what they do not understand, so [[prior-knowledge]] and [[scaffolding]] come first and verification should be designed into [[assessment]] as an assessed activity: identify a false assumption, correct a unit or mechanism error, or justify rejecting a generated answer, keeping prompt logs and revision histories as reasoning traces. ([[vega-baudrit-genai-university-chemistry-education-review-2026|Vega-Baudrit and Rivera Álvarez 2026]]) ## Connected Concepts - [[metacognition]] - [[cognitive-offloading]] - [[ai-literacy]] - [[generative-ai]] - [[higher-ed]] - [[problem-based-learning]] - [[intelligent-tutoring]] - [[educational-development]] - [[teacher-role]] - [[student-experience]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools - [[cognitive-surrender]] ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Critical thinking as understanding and evaluating AI - [[jacome-vasconez-chatgpt-adoption-xai-2026]] — XAI-augmented UTAUT2: habit as strongest predictor, four adoption profiles (Jácome-Vásconez et al. 2026) - [[pearls-epistemic-verification-2026]] — PEARLS framework for epistemic agency and verifying AI output (Wang 2026) - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[critical-thinking-paradox-genai-learning-2026]] — The critical-thinking paradox in GenAI-integrated learning - [[gerlich-ai-tools-cognitive-offloading-critical-thinking]] — AI use negatively correlates with critical thinking via offloading (Gerlich 2025) - [[pedagogy-ai-mistakes]] — The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking (Hosseini 2026) - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and cognitive surrender: how AI reshapes human reasoning (Shaw & Nave 2026) - [[cognitive-commons-ai-expertise-regeneration]] — The tragedy of the cognitive commons: AI and expertise regeneration - [[zhao-genai-higher-order-thinking-meta-2026]] — GenAI and higher-order thinking meta-analysis - [[ai-advice-suppresses-ikt-suspension-2026]] — AI advice suppresses "I don't know" judgment even when the advice is wrong - [[voicu-ai-interpretive-cognition-ssh-2026]] — AI-mediated learning and the restructuring of interpretive cognition in SSH - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a Mediational Agent - [[avraamidou-ai-colonization-science-education]] — Disrupting the AI colonization of science education - [[videla-embodied-ai-education-choreography]] — Embodied cognition and AI in education - [[critical-media-literacy-education-2026]] — Technology, education and critical media literacy - [[li-mroziak-reorienting-critical-ai-literacy]] — Reorienting critical AI literacy - [[panciroli-ai-literacy-episodes-situated-learning]] — AI literacy via Episodes of Situated Learning - [[fowlin-operationalizing-learning-principles-ai]] — Operationalizing age-old learning principles with AI - [[reconceptualizing-community-inquiry-generative-ai]] — Reconceptualizing Community of Inquiry for GenAI - [[genai-thoughtless-use-self-directed-learning-2026]] — Thoughtless GenAI use and college students' self-directed learning - [[genai-counter-learner-groupthink-2025]] — Countering learner groupthink with GenAI-introduced controversy in PBL - [[ai-enhanced-pbl-chatgpt-scaffolding-2026]] — AI-enhanced PBL with ChatGPT adaptive scaffolding for critical thinking - [[luo-ibl-patterns-llm-bloom-2026]] — IBL patterns in LLM-driven environments (Bloom's perspective) - [[jiang-chatgpt-inquiry-steam-review-2026]] — ChatGPT for inquiry-based learning in STEAM - [[ai-tools-academic-work-cheating-2026]] — Student perceptions of AI tools, ethics, and impact on critical thinking - [[probing-ai-generated-physics-solutions-2026]] — Preparing students to critique AI-generated physics solutions - [[student-reception-genai-analogies-computing-2026]] — Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education - [[daoism-ai-education-philosophy-2026]] — Alternative AI Philosophy: Daoism as Method for AI in Education - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant assessment - [[genai-chinese-higher-education-integrity-2026]] — Gen-AI in Chinese higher education: integrity and engagement - [[critical-thinking-biological-sciences-ai-2025]] — Critical thinking in biological sciences and AI - [[zhu-e3-hot-embodied-intelligence-sustainable-learning]] — Fostering Sustainable Learning via Embodied Intelligence (E3-HOT) - [[ai-supported-experimental-design-chemistry-2026]] — AI-supported experimental design in practical chemistry - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions - [[ai-overreliance-complex-adaptive-system-2026]] — AI overreliance modeled as a complex adaptive system - [[llm-adaptive-programming-error-explanations-2026]] — LLM adaptive explanations of programming errors - [[student-ai-conversations-cognitive-engagement-2026]] — Discipline-associated Bloom-level cognitive engagement in student-AI conversations (Chang & Li 2026) - [[ai-assisted-inquiry-ssi-climate]] — AI-Assisted Inquiry in Socio-Scientific Issues on Climate Change - [[demir-akar-ai-media-literacy-children-2026]] — AI-based critical media literacy program for children - [[lu-ai-multimodal-writing-critical-thinking-2026]] — Dimension-specific critical-thinking gains in AI-supported multimodal writing (Lu et al. 2027) - [[critics-lm-critical-thinking-science-education-2026]] — CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education - [[lftutor-logical-fallacy-education-2026]] — teaching fallacy recognition through structured multi-turn dialogue - [[caeai-ai-companions-learning-over-performance-2026]] — learning over performance: what companions should be optimized and measured for - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback - [[davor-ai-supported-learning-higher-order-outcomes-2026]] — AI scaffolding, offloading, and verification literacy via metacognitive self-regulation (Davor et al. 2026) - [[shojaei-genai-dependence-critical-thinking-employability-2026]] — GenAI dependence bounds the use to critical-thinking link in business students (Shojaei et al. 2026) - [[ren-metacognitive-awareness-genai-reliance-2026]] — Reflection prompt cuts acceptance of incorrect AI advice and attribution bias (Ren 2026) - [[kenzhebayeva-ai-role-rotation-pedagogical-model-2026]] — Role rotation as structure for critical human-AI interaction (Kenzhebayeva et al. 2026) - [[vega-baudrit-genai-university-chemistry-education-review-2026]] — Verification-centered GenAI integration in university chemistry education (Vega-Baudrit and Rivera Álvarez 2026) --- ## [Computational Thinking](https://edtechdev.github.io/aied/concepts/computational-thinking/) > **Computational thinking** — a problem-solving approach involving decomposition, pattern recognition, abstraction, and algorithmic design. In AI education, computational thinking is both a prerequisite for understanding AI systems and a skill that AI tools can help develop. ## Questions to Consider - When you solve a problem by breaking it into parts, spotting patterns, abstracting the essentials, and designing steps — you're already doing computational thinking, even without a computer. Where have you done this recently? - A common assumption is that computational thinking is the same as coding or 'computer literacy.' How might they differ, and why might that difference matter for how you teach it? - Research suggests students' deficits in fundamental concepts — not the AI tool itself — are what limit their ability to judge AI suggestions. What must a learner already understand before they can critically evaluate an AI's output? - Some argue computational thinking should move learners from passively consuming AI outputs toward building, critiquing, and designing with AI. What would a classroom that treats students as producers rather than consumers actually look like? - Generative AI can now score students' computational thinking growth — yet both humans and AI struggle with the hardest construct, systems thinking. Where do you think automation of assessment should stop, and why? - Robotics research finds computational thinking only develops when concepts are made explicit and mapped to the curriculum, not treated as isolated tech exercises. What's the risk of teaching 'tech skills' without naming the thinking underneath? ## Introduction ### CT in an AI-era classroom The knowledge base's connected articles converge on a central claim: computational thinking (CT) is the conceptual bedrock students need in order to engage critically with AI, and it is also the skill most directly deepened by well-designed AI-supported learning. Below the evidence is grouped into four themes grounded in the linked articles. - **CT as the foundation of AI literacy and critical engagement.** Several studies show that CT is what lets learners evaluate, rather than just consume, AI outputs. [[chat-debugging-human-ai-collaboration-circuits|Chat debugging research]] found that when undergraduates debugged analog circuits with LLM help, their *deficits in fundamental concepts and critical thinking* — not the tool — were the limiting factor, since students lacked the core ideas needed to judge AI suggestions. [[llm-intervention-design-cs-review|A review of LLM intervention designs]] likewise concludes that the [[cs-education]] push toward computational thinking over syntax mastery is what separates effective interventions from "tool frustration." In early childhood, [[ai-play-framework-early-childhood-2026|the AI-Play framework]] builds unplugged, play-based [[ai-literacy]] by teaching children that "AI is a system built from parts" and "AI learns from examples" — a developmentally grounded first layer of CT. And [[academic-league-of-ai-2026|an AI academic league]] connects CT to real civic AI projects through [[project-based-learning]], embedding [[ai-literacy]] in practice. Together these suggest CT is the transferable cognitive core of AI literacy. - **Educational robotics as a vehicle for CT.** Robotics is the most-studied context for developing CT across [[k-12]] and [[stem-education]]. [[computational-thinking-educational-robotics-secondary-2026|Secondary-school research]] argues that educational robotics enhances problem solving and critical thinking only when CT concepts are made explicit and mapped onto the [[stem-education|STEAM]] curriculum rather than treated as isolated technical exercises. A [[game-based-gamified-robotics-education-review-2026|systematic review of 95 studies]] confirms that robotics fosters CT, creativity, and problem solving, and that [[game-based-learning]] suits informal settings while gamification dominates formal classrooms and supports project-based learning. [[microbit-robotics-machine-learning-teacher-training-2026|Teacher-training evidence]] shows an integrated Micro:bit + robot + machine-learning intervention produced significant CT knowledge gains (d = 0.638) in initial teacher education, arguing robotics should be embedded so future teachers can teach CT. LLMs can lower the barrier further: [[edusim-llm-robotic-simulation-education-2026|EduSim-LLM]] couples an LLM with robot simulation so beginners control robots via natural language, making CT-embedded robotics accessible without low-level coding. - **LLMs as tools for CT assessment and development.** [[generative-ai|Generative AI]] offers scalable ways to measure and scaffold CT. [[llm-computational-thinking-physics-2026|Physics CT assessment research]] showed LLMs can mirror human raters in scoring growth in Data Practices and Computational Problem-Solving Practices across large-enrollment [[physics-education]] courses — while both humans and the LLM struggled with the more complex Systems Thinking construct, marking a clear boundary for automation. [[visual-query-tracer-declarative-logic-learning|Visual query tracing]] shows how visualization can scaffold abstract computation, building intuition that supports CT development. [[student-misconceptions-conditionals-loops-taxonomy|A taxonomy of conditionals-and-loops misconceptions]] provides fine-grained targets for [[scaffolding]] and for automated misconception detection, connecting to [[misconceptions]]. These tools work best, however, when pedagogical design leads: [[llm-intervention-design-cs-review|the CS review]] found semester-long "Virtual Tutor" designs with scaffolded feedback consistently improved CT, whereas unstructured tool access increased frustration. - **CT across K-12, teacher education, and assessment redesign.** CT spans the whole [[k-12]] to [[higher-ed]] spectrum and is reshaping assessment. At the early-childhood end, AI-Play extends CT and AI literacy to Pre-K–K2 learners and non-technical families; at the university end, the [[genai-oop-programming-assessments-2026|OOP assessment study]] found 2026 GenAI systems outperform the average student on authentic programming exams yet still fail on interfaces, abstract classes, and inheritance — recurring conceptual gaps that mark exactly where CT remains hard to automate. [[solving-vs-evaluating-genai-solutions|A randomized A/B crossover study]] showed that evaluation-and-critique tasks produce comparable outcomes to generation, suggesting CT can be exercised through judging flawed AI solutions, though gains require deliberate scaffolding. Underpinning all of this is the teacher: the microbit study links CT instruction directly to [[teacher-education]], and [[hashmi-socratic-physics-chatbot-2025|Socratic chatbot research]] ties the precise problem formulation that CT demands to measurable course performance. ### CT and the shift from AI consumers to producers, creators, and designers A central goal for CT in the AI era is moving students and instructors beyond *passive consumption* of AI outputs toward *creating, building, and designing* with and for AI — an agenda that aligns CT with constructionist learning (learning-by-making). The knowledge base's connected articles increasingly make this producer/creator/designer turn explicit. [[ai-writes-code-student-writes-model-2026|Model-authorship research]] reframes learning-by-construction with GenAI as a measurable "model authorship" process — students author, debug, and iterate on AI models rather than just consuming AI-generated code or answers. [[code-to-learn-genai-artifact-construction-2026|The CtL-GenAI framework]] operationalizes this as constructionism for the GenAI age, treating artifacts students build with AI as the engine of CT development. [[computational-thinking-ai-agent-creation|CT through AI-agent creation]] shows that designing, not merely using, AI agents exercises decomposition, abstraction, and algorithmic reasoning directly. The new meta-analytic evidence sharpens this picture. [[astor-computational-thinking-meta-review-2026|A meta-review of 128 CT systematic reviews]] finds the field converging on a unified definition of CT as reasoning with abstract models that use computational steps and algorithms to solve problems — precisely the kind of model-building (rather than answer-consuming) thinking that production-oriented learning demands. [[tsingidou-ct-robotics-kindergarten-2026|CT-kindergarten robotics research]] shows even early-childhood learners become producers through play-based building with robots, using problem-based learning, storytelling, and scaffolding — a developmental first step toward seeing technology as something one constructs, not just operates. And [[solving-vs-evaluating-genai-solutions|evaluation-and-critique research]] demonstrates that CT can be exercised through judging and debugging flawed AI solutions — a producer stance toward AI output that resists the passive-consumption trap. The practical upshot is that CT instruction should be designed so learners *make things with AI* — authoring models, building agents, constructing artifacts, and critiquing AI output — rather than receiving finished solutions. This both deepens CT and builds [[ai-literacy]] as participatory and creative rather than merely conceptual. Teachers, in turn, need support to move from using AI tools to designing AI-enhanced learning activities (see [[teacher-role]] and [[professional-training]]). ### Practical guidance For educators, the consistent message is that CT is developed through *explicit, scaffolded, observable* engagement rather than passive AI use. Pair robotics with explicit CT-concept mapping to the curriculum; use LLMs for [[simulation]], natural-language control, and scalable assessment of CT growth while reserving human judgment for constructs like Systems Thinking; and redesign assessments to emphasize evaluation and diagnosis of AI output over raw generation. Whatever the setting — unplugged play in early childhood, robots in secondary [[stem-education]], or Virtual Tutors in [[higher-ed]] — structure the activity so students must reason about decomposition, pattern, abstraction, and algorithm rather than receive finished solutions. ### Connections to related concepts Computational thinking is the shared cognitive foundation beneath [[ai-literacy]] and [[critical-thinking]], the curricular core of [[cs-education]] and [[k-12]] computing, and the conceptual target that [[educational-robotics]], [[game-based-learning]], and [[project-based-learning]] are best designed to serve. It is deepened by [[llm|large language models]] and [[generative-ai]] when those are used as scaffolding tools, and it is the skill that student-misconceptions taxonomies and CT-aware assessments aim to measure. Teachers develop it through [[teacher-education]] and [[professional-training]], and it transfers across domains including [[physics-education]] and [[stem-education|STEM]] broadly. - **Computational thinking predicts AI-assistant learning.** [[computational-thinking-aica-2026|Eighth-grade students]] with high computational thinking significantly outperformed low-CT peers in an AI coding-assistant course, using the assistant for understanding rather than answer retrieval. ## Connected Concepts - [[cs-education]] - [[stem-education]] - [[ai-literacy]] - [[k-12]] - [[prompt-engineering]] - [[adaptive-learning]] - [[llm]] - [[generative-ai]] - [[higher-ed]] - [[educational-robotics]] - [[game-based-learning]] - [[project-based-learning]] - [[physics-education]] - [[scaffolding]] - [[critical-thinking]] - [[teacher-education]] - [[simulation]] - [[socratic-method]] - [[misconceptions]] - [[agentic-ai]] ## Connected Articles - [[icet-ml-education-trust-2026]] — Addressing Trust in AI Systems through Education: A Didactic Perspective - [[ai-pbl-computational-thinking-2026]] - [[computational-thinking-ai-agent-creation]] - [[reshaping-cs-education-genai]] - [[panciroli-ai-literacy-episodes-situated-learning]] - [[prompt-problems-nl-programming-mistakes]] - [[llm-computational-thinking-physics-2026]] - [[hashmi-socratic-physics-chatbot-2025]] - [[visual-query-tracer-declarative-logic-learning]] - [[llm-intervention-design-cs-review]] - [[academic-league-of-ai-2026]] - [[ai-play-framework-early-childhood-2026]] - [[edusim-llm-robotic-simulation-education-2026]] - [[computational-thinking-educational-robotics-secondary-2026]] - [[microbit-robotics-machine-learning-teacher-training-2026]] - [[chat-debugging-human-ai-collaboration-circuits]] - [[student-misconceptions-conditionals-loops-taxonomy]] - [[genai-oop-programming-assessments-2026]] - [[game-based-gamified-robotics-education-review-2026]] - [[solving-vs-evaluating-genai-solutions]] - [[conversational-agents-novice-programmers-scoping-2025]] — Scoping review of conversational agents for novice programmers - [[zhang-ct-ai-training-test-2026]] — Computational Thinking in AI Training Test (CTAT) - [[niri-steam-ai-literacy-review-2026]] — STEAM education for AI literacy: systematic review - [[computational-thinking-aica-2026]] — Computational Thinking Levels and AI Coding Assistants (2026) - [[ai-writes-code-student-writes-model-2026]] — Model authorship: theory & measurement for learning-by-construction with GenAI - [[code-to-learn-genai-artifact-construction-2026]] — CtL-GenAI: constructionism framework for artifact construction - [[astor-computational-thinking-meta-review-2026]] — CT meta-review of 128 systematic reviews - [[tsingidou-ct-robotics-kindergarten-2026]] — Systematic review of CT via robotics in kindergarten --- ## [SAMR Model](https://edtechdev.github.io/aied/concepts/samr-model/) > **The SAMR model** — a technology-integration framework developed by Ruben Puentedura that classifies the extent to which a technology transforms learning along four levels: **Substitution, Augmentation, Modification, and Redefinition**. The lower two levels (Substitution, Augmentation) *enhance* an existing task — the technology does what was done before, better or more conveniently; the upper two (Modification, Redefinition) *transform* it — the task itself changes to something not previously possible. In [[ai-education|AI in education]], SAMR is the standard lens for asking whether [[generative-ai|generative AI]] is being used to incrementally improve existing practice or to reconceptualize learning, and it sits alongside [[tpack|TPACK]] as a way of describing how teachers integrate technology rather than why they accept it. ## Questions to Consider - When AI is introduced into a course, is it a Substitution (a [[conversational-ai|chatbot]] replaces a search box) or a Redefinition (a task becomes possible that was not before)? What determines which level is appropriate — and is transformation always the goal? - SAMR and [[technology-acceptance-model|technology adoption models]] answer different questions: adoption theory asks *why* a teacher or institution accepts a tool; SAMR asks *how deeply* the tool changes learning. Which question does a given AI-integration study actually answer? - A [[meta-analysis-systematic-review|systematic review]] of AI integration in [[higher-ed|higher education]] found most uses sat at the Substitution or Augmentation level, with only one study approaching Redefinition. If that is typical, what does it suggest about the gap between AI's potential and its classroom reality? - SAMR is frequently invoked in [[teacher-education|teacher]] [[educational-development|professional development]] to help educators plan technology use. Does classifying a lesson's SAMR level change what a teacher actually does, or is it chiefly a descriptive label? - Critics argue SAMR, like [[tpack|TPACK]], rests on a humanist ontology that treats cognition as unchanged by technology, and that it says nothing about power, data ownership, or equity. Is a level-of-integration lens sufficient, or does AI-era integration need a more critical frame? ## Introduction SAMR describes the depth of technological transformation of learning tasks. Its four levels form an ascending scale from enhancement to transformation: **Substitution** (the tool replaces another with no functional change — a chatbot in place of a search engine), **Augmentation** (the tool replaces and improves — a word processor's spell-check over a typewriter), **Modification** (the task is significantly redesigned — students collaborate in real time on a shared AI-generated draft), and **Redefinition** (new tasks previously inconceivable become possible — learners co-create with a generative model in ways that have no non-AI analogue). The model is widely used in [[ai-technologies|educational technology]] [[research-methods-aied|research]] and [[educational-development|teacher professional development]] as a vocabulary for planning and evaluating technology integration, and it is a standard companion to [[tpack|TPACK]] and to [[technology-acceptance-model|technology-adoption]] frameworks — though it answers a different question than either. ## What the four levels mean - **Enhancement (lower half):** Substitution and Augmentation improve an existing task without changing its nature. Most routine AI use — [[automated-question-generation|question generation]], quick text generation, summarization — sits here: it is faster and more convenient but does not alter what the learner is asked to do. - **Transformation (upper half):** Modification and Redefinition change the task itself. [[vibe-coding|Vibe coding]], real-time co-construction with a model, and adaptive interactive dialogue are tasks that were not possible before generative AI, and represent the transformative end of the scale. SAMR is frequently paired with the [[icap-framework|ICAP framework]] because the two align: the four cooking-to-learning scenarios in [[thermomix-genai-education-analogy-2026|Rummel, Nachtigall & Panadero (2026)]] map ICAP [[student-engagement|engagement]] modes onto SAMR levels — ICAP *Passive* with SAMR *Substitution*, *Active* with *Augmentation*, *Constructive* with *Modification*, and *Interactive* with *Redefinition* — illustrating a progression from passive delegation to interactive co-construction. ## Evidence from the knowledge base - **AI integration in higher education is largely incremental, not transformative.** [[alsheikh-mapping-ai-integration-higher-education-2026|AlSheikh et al. (2026)]], in a PRISMA systematic review of 22 intervention studies screened from 959 records, graded AI integration with the SAMR model and found most studies clustered at the **Substitution or Augmentation** level, with fewer at Modification and only one approaching Redefinition. AI was typically introduced to improve existing practice — dominated by [[automated-assessment|assessment automation]] and [[personalized-learning|personalized-learning support]] — rather than to reconceptualize curricula or [[learning-gains|learning outcomes]]. - **SAMR is a common analytical lens for GenAI-driven [[curriculum-design|curriculum]] change.** [[rewriting-curriculum-genai-pedagogy-2026|Sabani et al. (2026)]] map five curricular shifts (static to dynamic, transmission to capability, local to [[governance|institutional]]) onto established lenses including SAMR and constructive alignment, using the model to clarify the [[pedagogy|pedagogical]] mechanisms by which GenAI reshapes curriculum. - **SAMR and TPACK anchor teacher-technology integration standards.** [[crompton-faculty-technology-integration-standards-2026|Crompton et al. (2026)]] situate their six faculty technology standards against existing frameworks — [[tpack]], RAT, SAMR, SETI — and standards (ISTE, UNESCO, DigCompEdu), most of which target [[k-12]] educators or only the [[teacher-role|teaching]] portion of faculty roles, leaving a gap in higher-education faculty development. - **SAMR is a target of the posthumanist critique.** [[elsayed-pedagogical-symbiosis-posthuman-learner|Elsayed (2026)]] critiques TPACK, SAMR, and [[ai-literacy]] models for sharing a humanist ontology that presumes a bounded learner whose cognition is fundamentally unchanged by technological mediation, arguing these instrumentalist frameworks cannot address AI's constitutive role in cognition. - **A framework for integration, not a framework for power.** [[reclaiming-epistemic-agency-co-agency-2026|Poudyal (2026)]] evaluates SAMR alongside TPACK and other integration frameworks and finds none address equitable power, data ownership, or accountability, motivating an alternative ecological co-agency framework. ## SAMR, TPACK, and technology adoption: how they differ Three frameworks are frequently conflated but answer distinct questions: - **[[technology-acceptance-model|Technology adoption models]]** (TAM, UTAUT) explain *why* an individual or institution accepts and continues using a technology — perceived usefulness, ease of use, and social influence. - **[[tpack|TPACK]]** describes the *knowledge* a teacher needs to integrate technology effectively — the interplay of technological, pedagogical, and content knowledge, extended in the AI era to AI-TPACK/GenAI-TPACK. - **SAMR** classifies *how deeply* a technology transforms a learning task, from enhancement to redefinition. SAMR is best understood as an integration-depth lens used in planning and evaluation, complementing adoption theory (which explains uptake) and TPACK (which explains teacher capability). In the AI era, it is most productively used to ask whether generative AI is being applied to enhance existing tasks or to enable genuinely new ones — a distinction that runs throughout the knowledge base's assessment and curriculum research. ## Implications for practice - **Use SAMR to ask the depth question, not to prescribe transformation.** Enhancement is not inherently inferior; many routine AI uses are legitimate Substitutions. The model is diagnostic, clarifying what a given use actually changes. - **Pair it with ICAP.** SAMR describes what the *task* becomes; ICAP describes how the learner *engages*. Used together they distinguish surface substitution from deep interactive co-construction. - **Complement it with adoption and equity lenses.** SAMR says nothing about why a tool is adopted or who benefits. Pair it with [[technology-acceptance-model|adoption models]] and [[equity-in-ai-education|equity]] analysis to avoid a depth label substituting for a [[critical-thinking|critical evaluation]]. - **Treat evidence of incremental integration as a finding, not a failure.** If most AI integration clusters at Substitution/Augmentation, the design task is not to force Redefinition but to recognize that transformative use requires different task designs, not just better tools. ## Connected Concepts - [[ai-education]] — AI in education (umbrella) - [[tpack]] — Technological Pedagogical Content Knowledge - [[technology-acceptance-model]] — Technology adoption models - [[icap-framework]] — The ICAP framework of cognitive engagement - [[ai-technologies]] — AI technologies and techniques - [[teacher-ai-competency]] — Teacher AI competency - [[educational-development]] — Educational development - [[learning-design]] — Learning design - [[k-12]] — K-12 education - [[higher-ed]] — Higher education - [[generative-ai]] — Generative AI ## Connected Articles - [[alsheikh-mapping-ai-integration-higher-education-2026]] — AI integration in higher ed graded with SAMR: mostly Substitution/Augmentation - [[thermomix-genai-education-analogy-2026]] — ICAP and SAMR mapping the four cooking-to-learning scenarios - [[rewriting-curriculum-genai-pedagogy-2026]] — SAMR among the lenses for GenAI-driven curriculum change - [[crompton-faculty-technology-integration-standards-2026]] — SAMR among the frameworks informing faculty technology standards - [[elsayed-pedagogical-symbiosis-posthuman-learner]] — The posthumanist critique of SAMR's humanist ontology - [[reclaiming-epistemic-agency-co-agency-2026]] — SAMR's silence on power, data ownership, and accountability --- ## [Technological Pedagogical Content Knowledge (TPACK)](https://edtechdev.github.io/aied/concepts/tpack/) > **Technological [[pedagogy|Pedagogical]] Content Knowledge (TPACK)** — the framework (Mishra & Koehler, 2006) describing the integrated knowledge teachers need to use technology effectively in teaching: the interplay of Technological Knowledge (TK), Pedagogical Knowledge (PK), and Content Knowledge (CK), and their intersections. In the AI era, TPACK has been extended to **AI-TPACK** / **GenAI-TPACK**, modeling how teachers integrate generative AI into content-area instruction. It is the dominant theoretical lens for understanding how [[teacher-ai-competency|teacher AI competency]] is structured and built through [[educational-development|professional development]]. ## Questions to Consider - TPACK claims effective teaching with technology isn't the sum of separate skills (knowing your subject + knowing teaching + knowing the tool) but the product of their interplay. Think of a lesson that genuinely worked with a technology. Which combinations of content, pedagogy, and tech knowledge — not any single one — seemed to be doing the work? - A common assumption is that a teacher who is 'tech-savvy' is therefore ready to teach with AI. Where might that assumption fail — for instance, when a technically fluent teacher still uses AI in a pedagogically shallow or content-inaccurate way? What would 'competence' look like that a simple tech-skills test misses? - The AI era reframes the technology in TPACK from a passive tool to an active agent that can plan, generate, and tutor. If a teacher's job shifts from operating a tool to orchestrating an AI that acts on its own, what new knowledge does that demand — and can any static checklist capture it? - The page suggests that professional development should train the intersections, not just the tools. Think about the last technology training you attended or designed. Was it mostly 'how to use the software,' or did it build content and pedagogy together with the technology? Which approach would you expect to change classroom practice more, and why? - Research found different 'teacher archetypes' — optimizers, creators, passive observers — benefit from different kinds of support. Which archetype do you most resemble when using AI in teaching, and what kind of [[scaffolding]] do you think would help you most? Would your learners describe you the same way you do? - Some [[research-methods-aied|researchers]] argue effective AI integration emerges from a teacher's beliefs and sense of efficacy, not just their knowledge. What do you believe about AI's role in learning, and how might that belief — more than your technical skill — shape whether and how you actually integrate it? ## Introduction - **[[crompton-faculty-technology-integration-standards-2026|Crompton et al.]]** DBR operationalizes faculty technology-integration standards that extend the TPACK framework into [[governance|institutional]] practice. ## The Framework TPACK holds that effective technology integration is not the sum of separate knowledge domains but the product of their **dynamic interplay**. The framework comprises three base domains and four intersections: - **Content Knowledge (CK)** — knowledge of the subject matter to be taught. - **Pedagogical Knowledge (PK)** — knowledge of teaching methods, strategies, and how students learn. - **Technological Knowledge (TK)** — knowledge of how to use tools and [[ai-technologies|technologies]], including [[generative-ai|generative AI]]. - **Pedagogical Content Knowledge (PCK)** — how to teach specific content effectively. - **Technological Content Knowledge (TCK)** — how technology shapes and represents content. - **Technological Pedagogical Knowledge (TPK)** — how technology supports or constrains teaching strategies. - **TPACK** — the emergent, integrated knowledge at the center, where all three domains interact to enable technology-enhanced, content-specific teaching. ## AI-TPACK and GenAI-TPACK The AI era has pushed the framework toward a technology-with-intelligence reading. Rather than a passive tool, generative AI is an active agent that can plan, generate content, tutor, and adapt — so integration knowledge increasingly includes **orchestration**: deciding when and how AI acts, scaffolds, or yields to [[human-in-the-loop-ai|human judgment]]. - **Beyond discrete knowledge.** [[ai-tpack-teacher-multi-agent-workflow|AI-TPACK research]] argues effective AI integration emerges not from possessing separate domains but from the dynamic interplay of **systems thinking**, **pedagogical beliefs**, and **[[self-efficacy]]** — challenging static, checklist-based models of teacher AI competency. Teacher archetypes (Systematic Optimizers, Prolific Creators, Passive Observers) emerge from how teachers design multi-agent instructional workflows. - **Cross-level evidence for the pedagogical core.** [[pedagogy-first-technology-second-teacher-knowledge-2026|A multilevel study of 46 teachers and 2,832 secondary students]] found technical AI knowledge alone was insufficient — even slightly dampening students' perceptions of AI for social good — while pedagogical AI knowledge (TPAIK) is what fostered students' perceptions and behavioral intention to learn AI. The result distills to a **"pedagogy first, technology second"** guideline that echoes the mediating-role findings above. - **A review lens for the whole field.** [[edurev-100741-tpack-genai-review|A systematic review from a TPACK perspective]] (Liu & Zhong, 2025) analyzed 71 empirical studies of GenAI in student learning, finding an overall positive effect (Hedges' g = 0.752) and identifying GenAI literacy for students and **GenAI-TPACK professional development for teachers** as the two critical priorities for the field. - **Proficiency alone does not predict pedagogical integration.** [[choi-teacher-ai-interaction-lesson-design-2026|Choi et al. (2026)]] measured teachers' Intelligent-TPACK to segment participants and observed that even AI-proficient novices relied passively on AI output during lesson design, whereas experienced teachers — with lower measured AI-TPACK — critically re-engaged and adapted AI suggestions to pedagogical context. The result reinforces the pattern above: AI-TPACK translates into sound classroom use through experienced pedagogical judgment, and [[teacher-education]] support must therefore target the *application* of AI knowledge, not its mere possession. - **[[teacher-education|Teacher education]] context.** TPACK is instrumental in cultivating teachers' competency to integrate technology into [[curriculum-design|curriculum]]-specific instruction, which is why teacher-education and PD research (e.g., [[genai-pd-ai-pck-learning-gain-2026|intensive GenAI PD programs]], [[ai-tpack-preservice-math-teachers|AI-TPACK readiness among pre-service teachers]]) increasingly measures it as the outcome of interest. - **Pedagogical knowledge mediates GenAI integration.** [[tpack-genai-inservice-teachers-mediation-2026|Mohebi and ElSayary (2026)]] surveyed **325 in-service teachers across 26 countries** and interviewed seven, using an explanatory sequential [[mixed-methods-research|mixed-methods]] design to model TPACK-GenAI. Technological Knowledge (TK), Pedagogical Knowledge (PK), and Pedagogical Content Knowledge (PCK) each associated with overall TPACK-GenAI, but **Technological Pedagogical Knowledge (TPK) mediated** these links — evidence that the "from proficiency to pedagogy" move matters: translating GenAI skill into sound classroom use runs through pedagogical-technological integration knowledge, not tool familiarity alone. - **In [[higher-ed|higher education]], content expertise does not transfer to GenAI integration.** [[sutedjo-faculty-genai-tpack-21-2026|Sutedjo, Chowdhury & Liu (2026)]] adapted the [[ai-literacy|TPACK-21]] instrument to survey **127 faculty** at a U.S. research university on using GenAI to teach 21st-century skills. Faculty reported strong CK (M = 5.15) and PCK (M = 4.70) but low technology-integrated knowledge — TPK (M = 2.62), TCK (M = 2.75), and overall TPACK (M = 2.55, the lowest domain) — and CK showed **no significant correlation with TK or any technology-integrated domain** (r = .11–.15, ns). The three technology-integrated domains inter-correlated so strongly (r = .81–.91) that they may function as a single GenAI-integration factor. The result reinforces the *training the intersections* design principle: [[educational-development|faculty development]] must deliberately build GenAI-integration knowledge through [[discipline-specific-aied|discipline-specific]] activities rather than assume subject-matter expertise will carry over, and treat the technology-integrated domains as a shared GenAI-literacy foundation. - **Institutional mode and training level interact.** [[ai-training-science-teacher-tpack-distance-2026|Mnguni et al. (2026)]] compared 186 final-year science student teachers at a South African campus-based university (n = 85) and a [[online-teaching-and-learning|distance education]] university (n = 97). Self-reported TPACK was higher on campus (64.0% versus 47.4%), Pedagogical Knowledge was the least reported domain in both groups (53% on campus), and the level of AI training was associated with self-reported TPACK only at the distance institution, where completing a five-credit short course predicted *weaker* reported TPACK than no training at all (p = .039, versus p = .607 on campus). Readiness therefore tracks the institutional provision context rather than hours of training alone, and the authors caution that [[self-report-measures|self-reported]] readiness cannot stand in for performance evidence. - **Integrated TPACK moves most when AI is scaffolded inside authentic problem-based tasks.** [[chen-osman-preservice-physics-tpack-ctd-pbl-2026|Chen and Osman (2026)]] compared an eight-week AI-supported CTD-PBL module with conventional instruction for **130 third-year pre-service [[physics-education|physics]] teachers** in an intact-class quasi-experimental design (65 per condition), with groups using DeepSeek through task-specific prompt templates and every AI output required to pass human verification before entering an instructional artifact. The module group reported higher post-test TPACK (M = 4.04 vs. 3.40, p < 0.001, d = 1.02) and higher perceived [[problem-solving|collaborative problem solving]] (M = 3.62 vs. 3.05, d = 0.88), with significant group × time interactions on both outcomes. Where the gains landed is the TPACK-relevant detail: the largest dimensional effects were in the integrated domains — TPCK (d = 1.21), PCK (d = 1.06) and TPK (d = 0.95) — and the [[collaborative-learning|collaboration]] gains were clearest in shared knowledge building and social regulation, i.e. the intersections rather than the base domains. The authors present these as differential change associated with an integrated instructional condition, not as an AI effect: AI was never isolated from the PBL task chain, structured collaboration, instructor scaffolding, [[peer-assessment|peer feedback]] or reflective revision, and both outcomes were [[self-report-measures|questionnaire-based]] perceptions of competence rather than demonstrated classroom performance. **AI-TPACK as the mediator between literacy and classroom integration.** A structural equation model of Chinese pre-service science teachers ([[ai-literacy-ai-integrated-inquiry-science-teaching-2026|Zou, Li, Wang & Du, 2026]]) positions AI-TPACK not as a parallel competency but as the transmission mechanism through which general [[ai-literacy]] becomes teaching practice. AI literacy predicted AI-TPACK strongly, AI-TPACK in turn predicted [[science-education|science teaching]] [[self-efficacy]], and teaching self-efficacy was the strongest single predictor of the intention to teach through AI-integrated [[inquiry-based-learning|inquiry]]. The full serial chain (AI literacy → AI-TPACK → self-efficacy → intention) was significant, as were the separate indirect routes through AI-TPACK and through self-efficacy; taken together, most of AI literacy's association with intention ran through these mediators rather than directly, and the result held after controlling for gender, year of study, major, and AI use frequency. The model's practical claim is a sequence with an entry point: AI literacy is necessary but insufficient, and the work of integration happens where technological, pedagogical, and content knowledge are combined — which is also where teachers' confidence in teaching the subject is built. ## Why It Matters in AI Education TPACK is the organizing framework for the teacher-side of the knowledge base's evidence base. It explains why teacher AI competency is more than tool fluency: teachers must integrate technological, pedagogical, and content knowledge together to turn AI into [[learning-gains|learning gains]]. The knowledge base's [[teacher-ai-competency]] page covers the competency dimensions; TPACK is the *knowledge structure* that underlies them. Research on [[teacher-ai-adoption-confidence|teacher confidence]], [[educational-development|professional development]], and [[teacher-role|the transforming teacher role]] all operate within (or against) this framework. The framework is now being extended to evaluate whole programs: [[wu-li-evaluation-indicator-ai-certificate-programs-2026|Wu & Li (2026)]] build an AHP-Fuzzy-AHP evaluation system for AI certificate programs on TPACK dimensions and find that faculty professional competence — not technical infrastructure — is the largest gap between expert priority and current provision, signaling that credentialing should invest in teachers' integrated pedagogical capacity. ## Design Implications 1. **Train the intersections, not just tools.** PD should build technological, pedagogical, and content knowledge together rather than offering isolated tool training — the core TPACK design principle. 2. **Treat AI as an agent, not an appliance.** AI-TPACK extends the framework toward orchestration of [[agentic-ai|AI agents]], requiring systems thinking and pedagogical judgment about when AI should act. 3. **Differentiate PD by teacher profile.** Different teacher archetypes (optimizers, creators, observers) benefit from different scaffolding — advanced frameworks, rapid feedback, or explicit modeling respectively. 4. **Assess integrated competence.** TPACK-oriented outcomes (e.g., AI-PCK gains) should be measured as integrated capability, not self-reported tool familiarity. ## Connected Concepts - [[teacher-ai-competency]] - [[educational-development]] - [[teacher-role]] - [[generative-ai]] - [[learning-design]] - [[ai-literacy]] - [[scaffolding]] - [[metacognition]] - [[higher-ed]] - [[k-12]] - [[meta-analysis-systematic-review]] - [[teacher-education]] ## Connected Articles - [[wu-li-evaluation-indicator-ai-certificate-programs-2026]] — Evaluation Indicator System for AI Certificate Programs - [[pedagogy-first-technology-second-teacher-knowledge-2026]] — Cross-level study: 'pedagogy first, technology second' — TPAIK outweighs technical TAIK for student outcomes (Shen et al. 2026) - [[tpack-genai-inservice-teachers-mediation-2026]] — In-service teachers' TPACK-GenAI and the mediating role of pedagogical knowledge (Mohebi & ElSayary 2026) - [[reclaiming-epistemic-agency-co-agency-2026]] - [[crompton-faculty-technology-integration-standards-2026]] — Faculty standards for technology integration (TPACK-related DBR) - [[ai-vs-human-assessment-efl-tpck-2026]] — AI-generated vs human-developed assessment tasks in EFL - [[rewriting-curriculum-genai-pedagogy-2026]] — Rewriting the curriculum: GenAI-driven pedagogical change - [[edurev-100741-tpack-genai-review]] — Integrating generative AI into student learning: A systematic review from a TPACK perspective - [[ai-tpack-teacher-multi-agent-workflow]] — Modeling AI-TPACK through teacher multi-agent workflows - [[ai-tpack-preservice-math-teachers]] — AI-TPACK readiness among pre-service mathematics teachers - [[choi-teacher-ai-interaction-lesson-design-2026]] — Teacher-AI interaction in lesson design: Intelligent-TPACK, experience, and proficiency interplay (Choi et al. 2026) - [[genai-pd-ai-pck-learning-gain-2026]] — Intensive GenAI professional development and AI-PCK gains - [[teacher-ai-teaming-five-levels]] — Five-level teacher-AI teaming framework - [[teacher-ai-adoption-confidence]] — Teacher confidence and AI adoption - [[teacher-education-ai-literacy-sdt-2026]] — Teacher education for AI literacy via self-determination theory - [[genai-literacy-training-teacher-education-dbr-2026]] — Design-based research GenAI literacy training - [[rail-ed-genai-literacy-teacher-education]] — Rethinking GenAI literacy in teacher education - [[sec-ai-literacy-narrative-review-2026]] — Social-emotional competencies and AI literacy - [[genai-runaway-object-math-higher-ed]] — GenAI and mathematics in higher education - [[ai-changing-teaching-workflows]] — How AI is changing teaching workflows - [[riandi-teacher-ai-green-energy-education-2026]] — Teacher involvement in AI integration for green energy education (Riandi et al. 2026) - [[utility-value-intervention-teach-responsibly-genai-2026]] — Utility-value intervention effects in learning to teach responsibly with GenAI (Boos, Eder & Lachner 2026) - [[sutedjo-faculty-genai-tpack-21-2026]] — Faculty self-perceived TPACK-21 knowledge for GenAI in higher education (Sutedjo, Chowdhury & Liu 2026) - [[preservice-teacher-agency-genai-design-learning-2026]] — Pre-service teacher agency during GenAI interactions in design for learning (Krushinskaia, Elen & Raes 2026) - [[ai-training-science-teacher-tpack-distance-2026]] — Comparative survey of 186 South African science student teachers: campus advantage in self-reported TPACK, and AI training associated with weaker reported TPACK at the distance institution - [[ai-literacy-ai-integrated-inquiry-science-teaching-2026]] — AI-TPACK and science teaching self-efficacy serially mediate AI literacy's effect on inquiry-integration intention (Zou et al. 2026) - [[chen-osman-preservice-physics-tpack-ctd-pbl-2026]] — AI-supported CTD-PBL module and pre-service physics teachers' TPACK and collaborative problem solving (Chen & Osman 2026) --- ## [Pedagogies and Teaching Strategies](https://edtechdev.github.io/aied/concepts/pedagogy/) > **Pedagogies and teaching strategies** — the methods and approaches educators use to teach and facilitate learning, and the umbrella concept for the knowledge base's coverage of how teaching happens (in contrast to [[learning-theories]], which explains how learning happens). In [[ai-education|AI in education]], pedagogy is central because the choice of teaching strategy shapes how AI tools are deployed: the same generative-AI tool can be a [[scaffolding|scaffold]] under one pedagogy, a [[socratic-method|Socratic]] interlocutor under another, or an answer-generator under a third. The knowledge base documents individual pedagogies and treats them as the instructional lens through which AI's design and classroom use are evaluated. ## Questions to Consider - What's the difference between a teaching strategy (pedagogy) and a theory of how learning happens? Can you name a strategy you use and the learning theory it might rest on? - The page argues every AI tool 'embeds pedagogical assumptions' whether the designer states them or not. Take a [[conversational-ai|chatbot]] that just answers questions—what pedagogy is it quietly enacting, and is that intentional? - The same [[generative-ai|generative AI]] can be a scaffold under one pedagogy, a Socratic partner under another, or an answer generator under a third. Can you describe a single tool being used in these three different ways? - Evidence suggests *how* AI is used matters as much as *whether* it's used. Have you observed the same AI helping one class and harming another? What differed? - If you were advising a school on buying an AI tool, which questions would you ask to uncover the pedagogy embedded in it—rather than just its feature list? - Think of a learner-centered strategy you've tried (active, collaborative, project-based). What made it work or fall flat, and how might an AI tool have changed that outcome? ## Introduction Pedagogy and teaching strategy concern *how* educators teach — the activities, structures, and methods that organize learning — while [[learning-theories]] explains the underlying mechanisms of *how learning happens*. The two are complementary: a pedagogy operationalizes one or more theories, and the knowledge base treats pedagogy as the bridge from theory to classroom practice. Every AI tool embeds pedagogical assumptions about the desired instructional interaction, whether the designer states them or not. ## The pedagogy landscape The knowledge base documents a rich set of individual teaching strategies and pedagogies, organized into families: - **Student-centered and active approaches.** [[active-learning]] (students engaged in doing and thinking rather than passively receiving), [[project-based-learning]] (learning through extended projects), [[experiential-learning]] (learning through direct experience), and [[learning-by-teaching]] (learning by explaining to others). - **Collaborative and social approaches.** [[collaborative-learning]] (learning through [[group-work|group work]]), [[sociocultural-learning]] (learning through social participation and mediation), and [[socratic-method|Socratic questioning]] (learning through guided dialogue and questioning). - **Experience-based approaches.** [[experiential-learning]] (learning through direct experience and reflection), [[situated-learning]] (learning in authentic contexts), and [[embodied-learning]] (learning through physical/embodied interaction). - **Structured and guided approaches.** [[scaffolding]] (temporary, fading support), [[learning-design]] (systematic design of instruction), [[self-regulated-learning]] (learners directing their own learning), and [[sociocultural-learning]] (including structured, teacher-guided sociocultural support). - **Online and distance pedagogies.** [[online-teaching-and-learning|Online teaching and learning]] is itself a pedagogical context, not just a delivery channel: the medium shapes which strategies are viable ([[active-learning]] rethought for asynchronous forums, [[collaborative-learning]] via digital discussion, [[intelligent-tutoring|tutoring agents]] replacing face-to-face interaction). In this medium, AI raises both new opportunities (scalable [[personalized-learning|personalization]], always-on support) and new risks ([[academic-integrity|academic integrity]], [[cognitive-offloading|cognitive offloading]]), making pedagogical intent decisive. - **Motivation and [[student-engagement|engagement]] approaches.** [[game-based-learning]] (learning through games), [[self-determination-theory]] (supporting autonomy, competence, relatedness), and [[motivation]]-oriented strategies. - **[[equity-in-ai-education|equity]]-conscious pedagogies.** [[culturally-relevant-pedagogy|Culturally relevant pedagogy]], [[universal-design-for-learning|Universal Design for Learning]], [[critical-pedagogy]], and [[inclusive-learning]] ensure strategies serve diverse learners. ## How pedagogy appears in AI in education The knowledge base's [[research-methods-aied|research]] examines pedagogy at the intersection of AI and teaching in several ways: - **AI as a pedagogical agent.** AI tools embody pedagogies — a [[intelligent-tutoring|tutor]] built on [[socratic-method|Socratic questioning]] prompts learners to reason, while an answer-generating chatbot may default to direct provision (see [[reducing-ai-misuse]] on why the pedagogical stance matters). The [[agentic-ai|agentic AI]] literature shows that grounding agents in instructional-design theory outperforms raw [[prompt-engineering|prompting]]. [[genai-didactic-pedagogical-mediator-2026|Moganadas et al. (2026)]] reframe this role formally: rather than a dyadic instructor–[[student-modeling|student model]] with GenAI as an external supplement or threat, they propose a nested **instructor–student–GenAI triadic model** in which GenAI operates as a bounded *didactic-pedagogical mediator* within a shared didactic mediation space governed by institutions and stakeholders — yielding five researchable propositions on learning mediation, instructor-role transformation, AI literacy and learner agency, AI-transparent process-oriented assessment, and institutional [[governance]]. - **Pedagogy determines AI's effect.** A recurring finding is that *how* AI is used matters as much as *whether* it is used. [[instructional-guidance-genai-learning|Instructional-guidance research]] and [[generative-ai-guardrails-harm-learning|guardrailed-tutor RCTs]] show the same AI can harm or help depending on the pedagogical wrapper (hints vs. answers, structured vs. open use). - **Teaching strategies for AI literacy.** Teaching students *to use AI well* is itself a pedagogical task — [[ai-literacy]] and [[reducing-ai-misuse]] research develops strategies (think-first/AI-second/reflect, AI-declaration, calibration training) that belong to this umbrella. - **Pedagogy in teacher practice.** [[teacher-role]] and [[teacher-ai-competency]] examine how teachers adopt AI within their existing pedagogical repertoire, and [[pedagogical-llm-training]] / [[pedagogical-agent]] study AI tools trained to follow pedagogical principles. ## Relationship to learning theories Pedagogies and learning theories are closely linked: each pedagogy operationalizes one or more theories. For example, [[project-based-learning]] operationalizes [[constructivist]] and [[experiential-learning|experiential]] theories; [[socratic-method]] draws on [[sociocultural-learning]] and [[metacognition]]; [[scaffolding]] stems from the [[sociocultural-learning|Zone of Proximal Development]]. The knowledge base treats [[learning-theories]] as the conceptual foundation and this page as the instructional-practice umbrella — see also [[learning-design]], which concerns the systematic process of selecting and sequencing strategies. The [[learning-sciences|learning sciences]] sit one step further out again: this page covers the practice of teaching and the strategies educators choose, while the learning sciences study that practice empirically — describing, modeling and evaluating designs to establish which ones change learning. ## Learning gains across pedagogical strategies Different pedagogical strategies produce different kinds and sizes of [[learning-gains|learning gains]], and the knowledge base's evidence lets us compare them: - **Active and experiential strategies** generally produce stronger durable learning than passive reception, though they feel more effortful — [[active-learning]], [[experiential-learning]], [[project-based-learning]], and [[learning-by-teaching]] build understanding through doing. [[generative-ai-reduced-study-time-math|Research]] shows that strategies preserving effortful practice (rather than AI shortcutting it) protect [[learning-gains]]. - **Structured, guided strategies** ([[scaffolding]], [[self-regulated-learning]], [[learning-design]]) produce reliable but more modest gains — the guardrail evidence ([[generative-ai-guardrails-harm-learning|PNAS 2025]]) shows [[guardrails|hint-not-answer]] scaffolding preserves learning that unguarded answer-giving destroys. - **[[game-based-learning|Game-based learning]]** produces engagement and skill gains that are real but often modest and context-dependent — [[genai-educational-outcomes-meta-analysis|meta-analytic evidence]] finds game-assisted GenAI shows no significant added benefit over other formats, so games are best used for motivation and practice, not as a shortcut to gains. - **Collaborative and sociocultural strategies** ([[collaborative-learning]], [[sociocultural-learning]]) show gains mediated by interaction quality, increasingly studied with AI as a partner or peer. - **Socratic and dialogue-based strategies** ([[socratic-method]]) target higher-order thinking and reasoning — gains that are harder to measure than skill gains but central to [[critical-thinking]]. The key cross-cutting finding, consistent with the knowledge base's [[learning-gains]] research, is that **the strategy's effect on learning depends more on how it preserves learner effort and productive struggle than on which label it carries** — any pedagogy, even a "good" one, fails if AI is configured to bypass the cognitive work it was meant to elicit (see [[cognitive-offloading]], [[desirable-difficulties]]). ## Implications for AI in education - **Select pedagogy deliberately with AI:** the teaching strategy determines whether an AI tool supports or undermines learning, so pedagogical intent should drive AI tool selection and configuration. - **Keep learner agency central:** active, Socratic, and scaffolding pedagogies preserve the productive struggle and [[agency]] that AI can otherwise erode (see [[cognitive-offloading]], [[desirable-difficulties]]). - **Design AI to enact good pedagogy:** AI agents and tutors should be grounded in established instructional frameworks, not default answer-generation. - **Teach with and about AI:** pedagogies should both use AI to teach and teach learners how to use AI responsibly. ## Connected Concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[online-teaching-and-learning]] — Online Teaching and Learning - [[learning-theories]] - [[learning-sciences]] - [[learning-gains]] - [[learning-design]] - [[active-learning]] - [[collaborative-learning]] - [[project-based-learning]] - [[experiential-learning]] - [[game-based-learning]] - [[socratic-method]] - [[scaffolding]] - [[learning-by-teaching]] - [[self-regulated-learning]] - [[culturally-relevant-pedagogy]] - [[universal-design-for-learning]] - [[critical-pedagogy]] - [[teacher-role]] - [[teacher-ai-competency]] - [[ai-literacy]] - [[curriculum-design]] - [[higher-ed]] - [[k-12]] - [[pedagogical-llm-training]] — AI tools trained to follow pedagogical principles ## Connected Articles - [[genai-didactic-pedagogical-mediator-2026]] — GenAI as didactic-pedagogical mediator: instructor–student–GenAI triadic model (Moganadas et al. 2026) - [[icet-ml-education-trust-2026]] — Addressing Trust in AI Systems through Education: A Didactic Perspective - [[pedagogy-first-technology-second-teacher-knowledge-2026]] — 'Pedagogy first, technology second' — TPAIK outweighs technical TAIK for student outcomes (Shen et al. 2026) - [[wang-zhang-pedagogical-partnerships-genai-2026]] — Pedagogical partnerships with generative AI - [[ai-communities-of-inquiry-2026]] - [[ai-distance-education-systematic-review-2026]] - [[instructional-guidance-genai-learning]] — How instructional guidance shapes GenAI learning effects - [[generative-ai-guardrails-harm-learning]] — Guardrailed (hint-not-answer) tutoring eliminates the exam penalty - [[agentic-ai-pedagogical-best-practice-2026]] — The automation-vs-learning tension in agentic AI - [[jeon-isd-agent-bench-2026]] — Grounding agents in instructional-design theory - [[ai-tpack-teacher-multi-agent-workflow]] — Teacher TPACK and multi-agent workflows - [[edurev-100741-tpack-genai-review]] — Systematic review of GenAI in student learning from a TPACK perspective - [[ai-learning-tools-engineering-education-needs]] — AI learning tools in engineering education - [[fowlin-operationalizing-learning-principles-ai]] — Operationalizing learning principles with AI - [[learnlm-improving-gemini-learning]] — LearnLM: pedagogical instruction following - [[ai-video-dual-gatekeeping-2026]] — When Saying No Makes Better Videos: Dual Gatekeeping for Pedagogically Grounded AI Content Creation - [[zuo-instructor-power-genai-writing-2026]] — Power relations perceived by college instructors grappling with GenAI in writing (Zuo, Xu & Dunning 2026) - [[kibar-ilgaz-ai-instructional-design-review-2026]] — AI and Instructional Design Practice: A Systematic Review (Kibar & Ilgaz 2026) - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a mediational agent - [[liu-ai-literacy-interventions-meta-analysis-2026]] — Instructional approaches in AI literacy interventions --- ## [Active Learning](https://edtechdev.github.io/aied/concepts/active-learning/) > **Active Learning** — instructional approaches that engage students in doing things and thinking about what they are doing, rather than passively receiving information. In AI in education, active learning research examines both how AI tools can support active learning pedagogies and how active engagement with AI tools — rather than passive consumption — affects learning outcomes. ## Questions to Consider - You've likely heard 'active learning' praised. But is a student who clicks through a dashboard or accepts a generated answer really learning actively? What would make that activity 'active' in a meaningful sense? - The ICAP framework distinguishes active, constructive, and interactive engagement — only the deeper levels build lasting knowledge. When you last used an AI tool to learn something, which level of engagement did it actually push you toward? - AI can enable active learning at scale, but poorly designed AI can also do the cognitive work for the student. Where have you seen AI make a learner more passive rather than more engaged? - An EEG study found interactive student–AI collaboration produced the highest cognitive engagement, while full automation reduced it. Why might 'doing' with AI beat 'watching' AI do the work? - Teach-back — having a learner explain what they understand — surfaces gaps more effectively than passive re-reading. When might prompting a learner to explain to an AI be a better learning move than letting the AI answer for them? - Active learning depends on calibrated scaffolding that fades as competence grows. How hard is it for an AI tutor to know when to step back — and what's the risk if it never does? ## Introduction Active learning is a foundational principle in education research, grounded in [[constructivist]] theories that position learners as active constructors of knowledge. In the context of AI in education, the concept takes on dual significance: AI tools can enable active learning at scale (through [[intelligent-tutoring|interactive tutoring]], [[simulation|simulations]], and [[adaptive-learning|adaptive feedback]]), but poorly designed AI tools can also undermine it by [[cognitive-offloading|doing the cognitive work]] for students. The tension between AI assistance and active cognitive engagement — explored in articles like [[lak2026-hint-button-unproductive-use]] on premature hint use and [[efficiency-gain-illusion-ai-overreliance]] on [[cognitive-offloading|Over-Reliance]] — is a central concern. AI-enabled active learning manifests across multiple forms in this knowledge base: [[intelligent-tutoring]] systems that engage students in problem-solving rather than answer-giving, [[genai-mindtool-generative-learning]] approaches where students use AI as a thinking tool rather than a substitute, [[test-driven-ai-assisted-learning]] where students drive AI interaction rather than follow it, and [[curiobot-llm-tutoring-exploratory-learning]] exploratory learning environments. The [[scaffolding]] concept is tightly coupled — effective active learning requires calibrated support that fades as competence grows, which AI tutors must learn to provide. ## How active learning appears in the knowledge base's research - **Interaction mode determines cognitive engagement.** [[ai-assisted-learning-modes-eeg|An EEG study of high school students]] compared Auto (AI solves independently), Interactive (student–AI collaboration with scaffolding), and Manual (no AI) modes: **Interactive produced the highest cognitive engagement and task accuracy**, while Auto reduced engagement and risked over-reliance. This gives a neurophysiological dimension to the argument that AI must keep students *doing* rather than watching. - **Exploratory and simulation-based active learning.** [[supplynet-visual-exploratory-learning|SupplyNet]] uses a contextual multi-agent LLM simulation to support visual exploratory learning in supply-chain education, pairing an interactive network view with a branching "what-if" timeline so learners trace causal dynamics rather than consume abstract content. [[curiobot-llm-tutoring-exploratory-learning|Curiobot]] and [[genai-assisted-problem-posing-physics-2026|problem-posing in physics]] similarly foreground learner-driven exploration. - **Structured conversational workflows for active review.** [[knowloop-confusion-to-consolidation-2026|KnowLoop]] structures post-lecture review around three stages — Recognize (mark in-situ confusion), Resolve (clarification), and Consolidate (teach-back) — showing that teach-back prompts learners to articulate and reveal conceptual gaps, and that context-grounded AI outperforms general-purpose AI for targeted support. Teach-back instantiates [[learning-by-teaching]]. - **Active learning as a project-based, community structure.** [[academic-league-of-ai-2026|The Academic League of AI]] organizes extracurricular AI education around competition teams, study groups, and AI-for-social-impact projects, embodying active and [[project-based-learning|project-based learning]] through democratic student governance rather than top-down curriculum. - **Mindtools and generative engagement.** [[genai-mindtool-generative-learning|GenAI as a mindtool]] positions AI as a device students think *with* rather than a source of answers, aligning active learning with generative-learning theories where learners integrate new ideas into existing knowledge. ### The ICAP framework as the organizing lens Active learning is precisely operationalized by the [[icap-framework|ICAP framework]] (Interactive–Constructive–Active–Passive), which classifies learner behavior by mode of cognitive engagement and knowledge change. Under ICAP, what is colloquially called "active learning" actually spans three distinct, ordered levels of engagement: *active* (acting on material, e.g. taking notes or answering a prompt), *constructive* (generating new output beyond the given, e.g. self-explaining or drawing), and *interactive* (co-constructing meaning through dialogue). This matters for AI in education because an AI tool can masquerade as "active" while keeping learners in the shallowest modes: clicking through a dashboard or accepting a generated answer is active at best, not constructive or interactive. ICAP thereby sharpens the central design goal of active learning — **push learners from active toward constructive and interactive engagement** — and warns against AI systems that *answer for* the learner, which keep them passive.([[icap-cognitive-engagement-llm-agents]])([[hingle-collaborative-ai-literacy-2025]]) This connects active learning directly to [[icap-framework]], [[student-engagement]], and [[collaborative-learning]], whose highest ICAP mode is interactive dialogue. ## Practical guidance - **Keep the learner in the loop.** Design AI interactions so students act on and with output (interactive, scaffolded modes) rather than receiving finished answers; full automation measurably reduces cognitive engagement. - **Anchor AI support in learners' own activity.** Confusion points, learner-driven questions, and problem-posing give personalized entry points for review and exploration. - **Use teach-back and explanation.** Have learners articulate what they understand; surfacing gaps through explanation is more active than passive re-reading. - **Pair active engagement with calibrated scaffolding.** Support should fade as competence grows — [[scaffolding]] that never withdraws can itself become passive reliance. - **Prefer tools that make thinking visible.** Exploratory simulations, mindtools, and interactive problem-spaces support the causal tracing and comparative reasoning at the heart of active learning. ## Connections to related concepts Active learning is deeply connected to [[collaborative-learning]] (much active learning is social), [[learning-by-teaching]] (explaining to others is maximally active), [[project-based-learning]] and [[experiential-learning]] (learning by doing in authentic contexts), [[embodied-learning]] (physical engagement), [[game-based-learning]], and [[simulation]]. It relies on [[scaffolding]] and timely [[feedback]], and is threatened by [[cognitive-offloading|over-reliance]] when AI substitutes for effort. Grounded in [[constructivist]] and [[learning-theories]], it spans [[higher-ed]], [[k-12]], and [[stem-education]]. Active learning is one of the strongest levers on [[learning-gains|learning gains]] in the AI era. Because active strategies build understanding through effortful doing, they are the most robust to AI short-circuiting — and the knowledge base's evidence shows that preserving that effort protects durable learning while letting AI absorb it erodes it ([[generative-ai-reduced-study-time-math|reduced study time]], [[stromberg-generative-ai-learning-penalty-secondary-2026|the learning penalty]], [[lak2026-hint-button-unproductive-use|hint abuse]]). Instructors who pair active-learning designs with [[learning-gains|measured gains]] on unassisted outcomes get the clearest picture of whether AI-assisted activity actually improved learning. ## Connected Concepts - [[learning-gains]] - [[problem-based-learning]] - [[learning-by-teaching]] - [[scaffolding]] - [[constructivist]] - [[learning-design]] - [[intelligent-tutoring]] - [[student-experience]] - [[higher-ed]] - [[k-12]] - [[stem-education]] - [[generative-ai]] - [[feedback]] - [[cognitive-offloading]] - [[collaborative-learning]] - [[learning-theories]] - [[icap-framework]] - [[student-engagement]] - [[project-based-learning]] - [[experiential-learning]] - [[embodied-learning]] - [[simulation]] - [[game-based-learning]] - [[help-seeking]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education ## Connected Articles - [[espino-ai-business-education-review-2026]] - [[ai-pbl-computational-thinking-2026]] - [[genai-counter-learner-groupthink-2025]] - [[beck-genai-literacy-economics-hands-on]] — Active-learning GenAI framework for economics (Beck & Brodersen 2025) - [[lak2026-hint-button-unproductive-use]] - [[efficiency-gain-illusion-ai-overreliance]] - [[neurodivergent-computing-students]] - [[genai-mindtool-generative-learning]] - [[test-driven-ai-assisted-learning]] - [[curiobot-llm-tutoring-exploratory-learning]] - [[genai-assisted-problem-posing-physics-2026]] - [[ai-assisted-learning-modes-eeg]] — EEG study of AI interaction modes (interactive > auto) - [[supplynet-visual-exploratory-learning]] — SupplyNet: visual exploratory learning via multi-agent simulation - [[knowloop-confusion-to-consolidation-2026]] — KnowLoop: staged conversational post-lecture review - [[academic-league-of-ai-2026]] — Academic League of AI: project-based active learning - [[chatgpt-math-biology-challenge-based-learning-2025]] — ChatGPT in challenge-based biology/math courses - [[critical-thinking-biological-sciences-ai-2025]] — Critical thinking in biological sciences and AI - [[mujib-ai-ibl-creative-math-2026]] — AI-supported IBL and creative mathematical performance - [[pedagogy-ai-mistakes]] — The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking (Hosseini 2026) - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions --- ## [Collaborative Learning](https://edtechdev.github.io/aied/concepts/collaborative-learning/) > **Collaborative Learning** — instructional approaches where students work together to solve problems, complete tasks, or construct knowledge, supported or mediated by AI tools. In [[ai-education|AI in education]], collaborative learning [[research-methods-aied|research]] spans AI as a collaboration partner, AI as a mediator of human collaboration, and the design of collaborative AI tutoring systems. ## Questions to Consider - Think of a time you learned something deeply in a group. What made it work? Now imagine an AI [[conversational-ai|chatbot]] joining that group — how could it strengthen or quietly undermine what you experienced? - Research finds a trade-off: delegating reasoning to AI produces the best task performance but the least self-regulatory engagement, while the mode that builds self-[[regulation]] underperforms on the task. If you had to choose, which would you protect — the outcome or the struggle? - The ICAP framework ranks 'interactive' collaboration as the deepest form of engagement. Could an AI that answers for the group actually downgrade collaboration from interactive to merely passive — even if students feel more satisfied? - One study found AI mediators are trusted only while they stay neutral; when the AI shifts to advising or challenging, that trust erodes. How neutral should a group's AI mediator really be? - When learners use AI to produce a polished artifact, they may skip the epistemic effort that builds understanding. How would you design an AI partner that surfaces disagreement and conflict instead of smoothing it over? - Neurodivergent students report needing structured assignments, small consistent teams, and explicit roles. If AI collaboration tools are built for the 'average' learner, who might they leave out — and how would you design differently? ## Introduction Collaborative learning is grounded in [[sociocultural-learning|sociocultural theories]] of learning that position knowledge construction as fundamentally social. AI introduces new dynamics: AI can serve as a peer, a facilitator, or a participant in collaborative processes. The articles in this knowledge base explore how AI-mediated collaboration affects [[learning-gains|learning outcomes]], epistemic engagement, and [[equity-in-ai-education|equity]] — and how collaborative structures must be designed to accommodate diverse learners. **Collaboration as a construct vs. group work as a structure.** Collaborative learning is the broader theory: knowledge is co-constructed through joint activity and dialogue. [[group-work|Group work]] is its most concrete formal implementation — a team producing a shared outcome, and often a shared grade. The two are closely related but not identical: a group can operate without genuine collaboration (task partitioned into independent parts, work merely merged), and collaboration can happen without formal groups (pairs, whole-class dialogue, or human–[[student-ai-interaction|AI interaction]]). AI presses hardest on exactly this gap — [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] found groups whose individual [[generative-ai|GenAI]] practice never became collective capability because the task never required joint work, alongside groups where the shared grade made GenAI use a coordination problem. The [[group-work|group work]] page examines these dynamics in depth; this page keeps the wider lens on collaborative learning as a whole. **AI as collaborative partner** explores AI's role in group learning. **[[polished-artifacts-fragile-engagement-2026|Kimmerle]]** conceptualizes the risk of reduced epistemic effort when learners use AI to produce polished knowledge artifacts, advocating for AI structured as an argumentative partner that preserves cognitive conflict. Testing this at classroom scale, [[oppenheimer-llms-collaborative-learning-partners-2026|Oppenheimer, Cash & Connell Pensky (2025)]] had introductory social-science students (n = 154) write argumentative essays, receive critiques from [[llm|LLMs]] such as ChatGPT, Gemini, or Claude, and then incorporate or rebut them; blind coders found reflection in 92.7% and active rebuttal of LLM claims in 87.8% of responses (inter-rater κs = 0.81–0.89), evidence that learners behaved as [[critical-thinking|critical]] consumers who preserved rather than surrendered the cognitive conflict of critique. **[[epistemic-emotions-collaborative-problem-solving]]** examines how emotions shape collaborative [[problem-solving]] with AI. **[[hingle-collaborative-ai-literacy-2025]]** explores collaborative approaches to [[ai-literacy|AI literacy]] development. **AI-mediated peer collaboration** examines how AI [[scaffolding|scaffolds]] human-to-human collaboration. **[[golrang-propact-pair-programming-2026]]** and **[[agent-voice-accents-k12-group-learning]]** explore how AI agent characteristics affect group dynamics. **[[ai-agents-peer-learning-discourse]]** documents how [[agentic-ai|AI agents]] teaching each other produce discourse patterns resembling human peer learning. Classroom-wide systems extend this to the *relational* dimension of collaboration: **[[breideband-community-builder-cobi-2026|CoBi]]** uses speech recognition and language understanding to detect "uplifting" small-group discourse (being respectful, equitable, committed to community, moving thinking forward) and returns non-evaluative, classroom-level [[visualization]]s to support community building and collaboration skills, deliberately withholding student- or group-level feedback to protect [[privacy]] and [[trust]]. **Neurodivergent perspectives on collaboration** reveal critical design requirements. **[[neurodivergent-computing-students|Zastudil et al.]]** found that neurodivergent students need structured assignments, small consistent teams with explicitly defined roles, and predictable interaction patterns — requirements that AI collaboration tools must accommodate. This connects collaborative learning to [[inclusive-learning]] and [[neurodiversity]]. **Teacher-AI collaboration** examines how teachers and AI work together. **[[teacher-student-agency-orchestration]]** and **[[teacher-ai-teaming-five-levels]]** explore frameworks for human-AI collaborative teaching, connecting to [[teacher-role]] and [[human-in-the-loop-ai]]. **AI as a [[pedagogy|pedagogical]] mediator** reconceptualizes AI's role in collaboration beyond tool or peer. Drawing on sociocultural theory and [[distributed-cognition|distributed cognition]], **[[niari-ai-pedagogical-mediator-collaborative-learning|Niari]]** positions AI as an active participant in the orchestration of interaction, epistemic sense-making, and regulatory processes, redistributing agency, authority, and responsibility across human and non-human actors without displacing learner or teacher agency. This grounds collaborative learning in a socially mediated, co-regulated view of AI rather than an individualistic one. **Collaboration modes and the efficiency–regulation trade-off.** Empirical research on college students collaborating with AI for complex problem-solving identifies three distinct modes — *Delegated Reasoning*, *Concerted Interpretation*, and *Delegated Elaboration*. The most efficient mode (delegated reasoning) yields the highest task performance but the lowest learners' self-regulatory engagement, while the mode with greatest self-regulation (concerted interpretation) underperforms on task outcomes.([[hao-human-ai-collaborative-problem-solving-cognition]]) This reveals a central design tension: collaborative-learning environments must balance the efficiency of the distributed human–AI system against the depth of learners' [[self-regulated-learning|regulatory]] engagement. **GenAI as agent and space in small groups — mode matters.** [[xu-genai-collaborative-space-2026|Xu et al. (2026)]] observe that *how* a team accesses GenAI shapes collaboration: with a single shared interface in synchronous work, teams co-construct "collective prompts," run a surface–evaluate–embed cycle, and treat the chat as shared memory; in asynchronous work, private prompting and output "de-labeling" fragment [[explainable-ai|transparency]] and raise the cost of sustaining a shared cognitive model. Their GenAI-Supported Cooperative Work (GSCW) lens frames GenAI as a configurable agent (individual assistant to team member) and an interactive collaborative space — connecting access configuration directly to the [[icap-framework|ICAP]]-relevant quality of interactive engagement. **GenAI as group coordination infrastructure — and the risk of flattened cooperation.** [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] show how fifteen pre-service teacher groups handled GenAI in a graded group presentation, and the split runs against the usual assumption that group pressure increases AI reliance. Five groups intensified use to solve a familiar collaboration problem — not knowing what peers' sections contained — feeding that work into a chatbot to make it intelligible and align their own part, with one group rebuilding its cycle as *discussion → externalization to GenAI → collective review → re-discussion*. The authors read this as more than [[cognitive-offloading|cognitive offloading]], since students kept judgment while the tool absorbed coordination, but flag that the smoother workflow may bypass the disagreement through which cohesion is conventionally built, making relational labor the open question. Seven groups cut their GenAI use instead, protecting the [[situated-learning|situated]] knowledge built in shared classrooms ("AI only knows that moment when you type"), [[bias-mitigation|fairness]] to groupmates, originality across groups, and the diversity of perspectives the group already held. Three groups saw no change at all: with the task partitioned into independent sections, individually sophisticated GenAI practice never became a collective capability, even though coherence was an explicit criterion. The pattern suggests group norms, not the tool, decide what a group does with AI — and that collective adoption can lower the perceived [[ai-misuse-learning-harm|risk of misuse]] rather than raise commitment. **Collaboration as the object of instruction.** [[golrang-propact-pair-programming-2026|ProPACT]] is an AI-driven [[intelligent-tutoring|adaptive tutor]] for pair programming that treats the *dyad* — not the individual — as the unit of analysis, modeling joint visual attention, joint mental effort, and pupil-based signals in real time to predict collaborative breakdowns up to 30 seconds in advance and intervene before they occur. Dyads receiving proactive feedback achieved substantially higher debugging success and completed tasks more efficiently, and showed sustained gains in collaborative regulation afterward — evidence that AI can teach collaboration itself, not just support a task. Measuring collaborative competence poses the complementary challenge of assessing collaborative problem-solving (CPS) skill at scale, which traditionally requires manually coding process data from simulated tasks into CPS behaviors — time-consuming and impractical at scale; [[prompt-engineering|context-aware prompting]] of pre-trained language models automates this coding by modeling contextual dependencies and fusing cognitive and social abilities, achieving superior performance over strong baselines. **AI as a neutral mediator — and the tension when it stops being neutral.** [[spritz-ai-disciplinary-mediation-student-teams-2026|Spritz]] is a Discord-based [[llm]] probe that mediates disciplinary boundaries in interdisciplinary student teams by surfacing implicit assumptions and returning anonymized syntheses to shared discussion. Students valued it as both cognitive support and a relational buffer, but a central tension emerged: AI's perceived neutrality was load-bearing, and eroded once the AI moved from neutral mediator to advisor or challenger — a key design constraint for [[pedagogical-agent|agents]] that mediate collaboration while preserving [[human-ai-collaboration]] and [[trust-calibration]]. **A structured facilitator protocol — and its sycophancy risk.** [[ethics-training-agents-group-ethics-discussion-2026|Seo et al. (2026)]] extend collaborative-[[learning-design|learning design]] with a structured divergence–deliberation–convergence protocol: an LLM facilitator stacks speaking turns, times the phases, and summarizes incrementally, a format 45 students found simple and discussion-like and that reduced the need for external facilitation. The design turns on a tension the study measured directly: agents supported breadth of perspective-taking (groups produced larger, more divergent stakeholder and solution sets, and agents consistently voiced minority viewpoints), yet [[ai-sycophancy|sycophantic]], reasoning-free agreement flattened the cognitive conflict that makes collaboration deepen thinking. Participants asked for agent output that exposes intermediate deliberation steps rather than only conclusions, and for counterarguments that preserve genuine disagreement. **Collaborative structures for AI education.** [[academic-league-of-ai-2026|The Academic League of AI]] organizes AI education through democratic student [[governance]] and project teams, embedding [[active-learning]] and [[project-based-learning]] in a collaborative, community-connected structure. ### The ICAP framework: collaboration as the highest engagement mode Collaborative learning occupies the top of the [[icap-framework|ICAP framework]] (Interactive–Constructive–Active–Passive): the *interactive* mode — co-constructing meaning through dialogue, defending a position, or solving jointly — produces the deepest knowledge change in Chi's taxonomy. This makes ICAP both a justification for collaborative pedagogies and a design constraint on AI. An AI that mediates discussion (as [[spritz-ai-disciplinary-mediation-student-teams-2026|Spritz]] or [[golrang-propact-pair-programming-2026|ProPACT]] do) is valuable precisely when it sustains *interactive* engagement; an AI that answers for the group or smooths over cognitive conflict can downgrade collaboration to a mere *active* or *passive* mode. ICAP-based annotation (see [[icap-cognitive-engagement-llm-agents|extended ICAP measurement of collaborative dialogue]]) and facilitation-timing research both treat the quality of interactive discourse as the outcome of interest, grounding collaborative learning in [[student-engagement]] and the ICAP hierarchy.([[icap-cognitive-engagement-llm-agents]])([[llm-facilitation-timing-online-discussions]]) ## Practical guidance - **Model collaboration, not just the individual.** Tools that track dyadic or group state (as [[golrang-propact-pair-programming-2026|ProPACT]] does) can scaffold the collaboration itself, predicting and preventing breakdowns rather than reacting to them. - **Preserve cognitive conflict.** Structure AI as an argumentative partner that surfaces disagreement and implicit assumptions, avoiding the polished-artifacts problem where AI smooths over fragile epistemic engagement. - **Balance efficiency against self-regulation.** Collaborative AI that maximizes task efficiency (delegated reasoning) can undercut learners' regulatory engagement; design should deliberately protect space for concerted interpretation. - **Respect the neutrality constraint.** AI mediators are trusted while neutral; moving into advisory or challenging roles destabilizes that trust, so role switches should be explicit and configurable. - **Accommodate neurodivergent learners.** Structured assignments, small consistent teams, and explicit role definitions are requirements AI collaboration tools must support. - **Prefer non-evaluative, classroom-level feedback.** When supporting the relational dimension of collaboration, class-level aggregated feedback protects [[privacy]] and [[agency|student agency]] where individual scoring would feel surveilled; [[breideband-community-builder-cobi-2026|CoBi]] students preferred [[qualitative-research|qualitative]] visualizations (an organic tree) over [[quantitative-research|quantitative]] ones (a radar chart), and teachers valued using the system's noticings to spark reflection more than live display. - **Design for the viewing/attention that precedes contribution.** Collaborative learning in [[online-teaching-and-learning|online discussion]] forums depends not only on posting but on the reading that precedes it. [[hao-peer-exposure-bridging-social-capital-ai-summaries-2026|Hao & Cukurova (2026)]] show LLM-generated discussion summaries and example posts can act as navigational [[scaffolding|scaffolds]] that broaden students' exposure to peers' contributions and the network conditions for bridging (weak-tie) social capital — support that should complement, not replace, socio-pedagogical strategies for sustaining engagement under academic workload. ## Connected Concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[group-work]] — Group work - [[problem-based-learning]] - [[online-teaching-and-learning]] — Online Teaching and Learning - [[active-learning]] - [[icap-framework]] - [[scaffolding]] - [[teacher-role]] - [[human-in-the-loop-ai]] - [[equity-in-ai-education]] - [[ai-literacy]] - [[k-12]] - [[higher-ed]] - [[inclusive-learning]] - [[neurodiversity]] - [[distributed-cognition]] - [[self-regulated-learning]] - [[project-based-learning]] - [[human-ai-collaboration]] - [[trust-calibration]] - [[pedagogical-agent]] - [[student-modeling]] - [[student-engagement]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Protecting human connection in AI-mediated collaboration - [[jin-emergent-learner-agency-implicit-hai-2026]] — Emergent learner agency in implicit human-AI collaboration: supportive vs. contrarian personas - [[adaptive-ai-scaffold-collaborative-problem-solving-2026]] - [[genai-counter-learner-groupthink-2025]] - [[ai-communities-of-inquiry-2026]] - [[polished-artifacts-fragile-engagement-2026]] - [[epistemic-emotions-collaborative-problem-solving]] - [[hingle-collaborative-ai-literacy-2025]] - [[neurodivergent-computing-students]] - [[teacher-student-agency-orchestration]] - [[niari-ai-pedagogical-mediator-collaborative-learning]] - [[hao-human-ai-collaborative-problem-solving-cognition]] - [[golrang-propact-pair-programming-2026]] — ProPACT: proactive AI adaptive collaborative tutor for pair programming - [[spritz-ai-disciplinary-mediation-student-teams-2026]] — Spritz: AI disciplinary mediation in student project teams - [[academic-league-of-ai-2026]] — Academic League of AI: collaborative, project-based AI education - [[icap-cognitive-engagement-llm-agents]] — Extended ICAP framework for measuring engagement in collaborative dialogue - [[llm-facilitation-timing-online-discussions]] — LLM facilitation timing in online collaborative discussions - [[ba-ai-agents-cscl-review-2026]] — AI agents in computer-supported collaborative learning review - [[wei-perkins-genai-student-collaboration-scoping-2026]] — GenAI and student group work: a scoping review (Wei & Perkins 2026) - [[context-aware-prompting-cps-skill-identification-2026]] — Context-aware prompting for automated collaborative problem-solving skill coding - [[astra-multi-agent-tutoring-benchmark-2026]] — ASTRA synthetic benchmark for multi-agent tutoring and participation-balanced collaboration - [[xu-genai-collaborative-space-2026]] — GenAI as agent and collaborative space in small-group dynamics (Xu et al. 2026) - [[breideband-community-builder-cobi-2026]] - [[oppenheimer-llms-collaborative-learning-partners-2026]] - [[hao-peer-exposure-bridging-social-capital-ai-summaries-2026]] — AI-Generated Summary-Driven Learning Design in Online Discussion Forums - [[chen-zou-genai-group-assessment-agency-2026]] — GenAI as coordination infrastructure in student groups: intensified, restrained, and non-enacted use - [[ethics-training-agents-group-ethics-discussion-2026]] — Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration - [[durable-skills-measurement-ai-teammates-2026]] — Toward Scalable Measurement of Durable Skills - [[peer-group-vs-ai-feedback-2026]] — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education --- ## [Project-Based Learning](https://edtechdev.github.io/aied/concepts/project-based-learning/) > **Project-based learning (PBL)** — an active, learner-centered [[pedagogy]] in which students learn by engaging in extended, real-world projects that require inquiry, problem solving, and the application of knowledge to produce tangible outcomes. PBL emphasizes student [[agency|autonomy]], collaboration, and [[authentic-assessment|authentic tasks]], and is widely used with technology — including [[educational-robotics|educational robotics]] and AI — to give learners hands-on, meaningful projects. It contrasts with purely theoretical or lecture-based instruction. ## Questions to Consider - Think of the last time you truly 'learned by doing' — building, designing, or creating something real. What made that experience stick compared with a lecture you sat through the same week? How might that difference transfer to how students learn AI or robotics? - PBL and problem-based learning are frequently conflated, yet they differ: one centers on producing a tangible project, the other on resolving an ill-structured problem. Before you read, how would you distinguish the two — and does your answer matter for how a course should be designed? - A robotics study tackled a 'theory-practice gap' with an agile, semester-spanning project. When you think about your own field, where does the gap between what students are taught and what they can actually do tend to open up — and what would a project-based approach need to close it? - PBL emphasizes student autonomy, collaboration, and authentic tasks. But learners differ in their readiness to direct themselves. What could go wrong if you dropped extended projects into a classroom without support for self-direction, and who would it harm most? - PBL is widely coupled with gamification and educational robotics in the research. From your experience, is fun/[[student-engagement|engagement]] always a reliable proxy for deep learning, or can a well-scored 'game' mask shallow understanding? - Before reading on, ask yourself: what concrete evidence would convince you that project-based learning actually beats lecture-based instruction for a specific learning outcome — and how hard is that evidence to gather in a real course? ## Introduction PBL is closely related to [[active-learning]], [[experiential-learning]], [[collaborative-learning]], and [[constructivist]] pedagogy. It is especially valuable for AI and robotics education because these fields are inherently applied: learners best understand robots, algorithms, and systems by building and testing them in project contexts. PBL also fosters [[computational-thinking]], problem solving, and [[self-regulated-learning|self-direction]]. Project-based learning is closely related to — but distinct from — [[problem-based-learning|problem-based learning]]: both are learner-centered and context-driven, but problem-based learning centers on an ill-structured *problem* whose solution requires inquiry and knowledge construction, whereas project-based learning centers on producing a tangible *project* or artifact. The two are frequently conflated, and many AI-in-education frameworks draw on both (see the [[problem-based-learning|problem-based learning]] page for the AI-era treatment). ### How PBL appears in the knowledge base's research - **Robotics projects:** [[bots-blocks-project-based-robotics-education-2026|Bots and Blocks]] presents an agile, semester-spanning project-based approach to teach robotics in an applied [[cs-education|computer science]] program, addressing the theory-practice gap. - **The Project Approach in early childhood with AI agents:** [[creative-project-approach-ai-early-childhood-2025|Yang, Li and Lee (2025)]] extend PBL's foundational form — the Project Approach (Katz & Chard), an extended collaborative investigation of a real-world topic — into [[early-childhood-elementary-ai-education|early childhood]], proposing a five-step **Creative Project Approach** that integrates [[agentic-ai|AI agents]] and [[educational-robotics|robots]] (coding robots and generative social robots) into projects to foster young children's [[creativity|creative learning]]. The five steps — identify learning needs, facilitate [[teacher-role|teacher]]-guided child–robot interaction, situate AI in contexts, calibrate the automation/creativity balance, and evaluate outcomes — keep the teacher as a facilitator guiding inquiry, positioning PBL as the natural vehicle for developmentally appropriate AI use with the youngest learners. - **Gamification coupling:** [[game-based-gamified-robotics-education-review-2026|A systematic review]] found [[game-based-learning|Gamification]] in robotics education strongly favored project-based learning (p = .009). - **AI literacy and co-design:** PBL underlies many [[ai-literacy|AI literacy]] and [[teacher-education]] interventions, where learners co-create AI tools or resources. - **AI-agent-supported software PBL:** [[spec-driven-development-ai-agents-sdpbl-2026|Tanaka et al. (2026)]] embedded Spec-Driven Development with [[agentic-ai|AI agents]] into a team-based undergraduate software PBL course, structuring projects into investigation, planning, implementation, and review phases paired with instructor-run comprehension checks. - **Immersive VR studios with an embedded teaching agent:** [[ai-ive-pbl-vocational-design-creativity-2026|Jin et al. (2026)]] specify the **AI-IVE-PBL** model for vocational design education, pairing PBL with an AI-enabled immersive virtual environment ([[virtual-and-augmented-reality|VR]] headsets plus an [[llm]]-backed teaching assistant). PBL's customary constraints for vocational learners — limited equipment, hard-to-replicate scenarios, delayed teacher [[scaffolding]] — are absorbed by immersion plus an in-session agent, and the model is stated as a five-phase loop (discovery, envisioning, modeling, communication, refinement) with a named actor and artifact per phase, driven by sustained idea-developing discourse. In a 12-week quasi-experiment (n = 63) the condition raised design ability and creative ability and lifted cognitive and behavioral [[student-engagement|engagement]], while leaving ideational novelty (innovative thinking) and affective engagement unchanged — a reminder that the design-specific and the ideational parts of a project's value do not move together. PBL connects to [[active-learning]], [[experiential-learning]], [[collaborative-learning]], [[educational-robotics]], [[game-based-learning]], [[computational-thinking]], and [[higher-ed]]/[[k-12]] pedagogy. - **PBL supports AI-powered robotics learning.** [[educational-robotics-pathways-2026|Pathways research]] shows project-based robotics+AI curricula let high school students learn through engagement in real-world practice, designing, and playful creative expression. ## PBL in AI Literacy Courses - **Measuring PBL in AI literacy courses.** Zhu and Kong (2026) developed and validated an AI project-based learning scale (AI-PBLS) grounded in Hong Kong secondary and university students' experiences, and used it to show that perceived PBL fosters AI literacy course satisfaction through the mediating mechanisms of empowerment in AI [[problem-solving]] and AI [[ethics|ethical]] awareness. The scale offers [[research-methods-aied|researchers]] a validated instrument, and the mediation finding strengthens the case for PBL as a vehicle that builds confidence and ethical reasoning — not just content — in [[ai-education|AI education]]. ### PBL With Digital Storytelling in the AI Era - Project-based learning combined with [[storytelling-in-education|digital storytelling]] offers a pedagogical response to [[generative-ai|generative AI]] in art and design education. A 15-week embedded case study with 426 undergraduates implemented a PBL-DS framework in which digital storytelling served as the primary methodology for students to translate local cultural heritage into emotionally resonant, [[multimodal]] narratives, cultivating the creative capacities that AI lacks. ## Connected Concepts - [[problem-based-learning]] - [[active-learning]] - [[experiential-learning]] - [[collaborative-learning]] - [[educational-robotics]] - [[game-based-learning]] - [[computational-thinking]] - [[higher-ed]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[arts-design-and-media-education]] ## Connected Articles - [[mechanical-engineering-ai-curriculum-2026]] — Project-Based AI Education Curriculum in Thermal Engineering - [[pbl-structural-conditions-ai-2026]] - [[genai-counter-learner-groupthink-2025]] - [[bots-blocks-project-based-robotics-education-2026]] — Bots and Blocks - [[game-based-gamified-robotics-education-review-2026]] — Game-Based and Gamified Robotics Education - [[genai-literacy-training-teacher-education-dbr-2026]] — AI Literacy Training for Teachers - [[roboblockly-conversational-block-robotics-ct-2026]] — RoboBlockly Studio - [[white-wu-robotics-ai-education-2026]] — Robotics and AI in Education - [[academic-league-of-ai-2026]] - [[teachlm-post-training-llms-education]] — TeachLM: project-based tutoring data from Polygence - [[educational-robotics-pathways-2026]] — Pathways to Learning AI-Powered Educational Robotics (2026) - [[tsingidou-ct-robotics-kindergarten-2026]] — PBL is a dominant CT learning strategy - [[ai-literacy-course-satisfaction-pbl-scale-2026]] — AI-PBLS scale; empowerment and ethical awareness mediating PBL-to-satisfaction in AI literacy courses (Zhu & Kong 2026) - [[project-based-digital-storytelling-art-design-2026]] — Project-based digital storytelling framework for art/design education in the AI era - [[creative-project-approach-ai-early-childhood-2025]] — The Creative Project Approach: AI agents and robotics within the Project Approach in early childhood (Yang, Li & Lee 2025) - [[spec-driven-development-ai-agents-sdpbl-2026]] - SDD with AI agents in team software PBL - [[ai-ive-pbl-vocational-design-creativity-2026]] — AI-IVE-PBL: an immersive VR design studio with an embedded teaching agent, evaluated against traditional PBL (Jin et al. 2026) --- ## [Problem-Based Learning](https://edtechdev.github.io/aied/concepts/problem-based-learning/) > **Problem-based learning (PBL)** — a learner-centered pedagogy in which students acquire knowledge and skills by working to understand and resolve a realistic, often ill-structured problem, with a facilitator guiding inquiry rather than delivering instruction. In the AI era, PBL's structural features — problem-driven inquiry, [[collaborative-learning|collaborative knowledge construction]], facilitation over instruction, and [[metacognition|metacognitive reflection]] — have emerged as the same conditions under which [[generative-ai|generative AI]] integration becomes educationally productive rather than substitutive. ## Questions to Consider - In problem-based learning, the *problem* is the driver of the [[curriculum-design|curriculum]]—not an illustration of content taught elsewhere. Can you recall a time you learned something deeply because you were solving a real problem first? - The page argues PBL's features (problem-driven inquiry, collaboration, facilitation, reflection) are exactly the conditions under which AI becomes a productive partner rather than a shortcut. Why might that be? - AI severs the link between a submitted artifact and the effort that produced it—the 'artifact-as-proxy' problem. How does assessing process instead of just product help, and what makes process [[assessment]] hard? - PBL originated in medical education partly because professional competence needs adaptive judgment, not routine execution. What's the difference between those two, and how does AI change which one we train for? - If an AI can now make 'wicked problems' and complex real-world cases accessible to more students, what might be lost when the challenge becomes easier to reach? - Where should an AI in a PBL setting draw the line between giving a problem, a hint, or an answer? What determines which is the right move at any moment? ## Introduction PBL originated in medical education (McMaster University, 1960s) and has since spread across [[medical-education|health professions]], engineering, and [[k-12]]. The learner takes responsibility for diagnosing learning needs, identifying resources, and constructing solutions, while the facilitator scaffolds the process. The distinctive claim of PBL is that the *problem* is the curriculum driver — not an illustration of content taught elsewhere but the context in which content is learned. ## PBL in the AI era The knowledge base's [[research-methods-aied|research]] shows that PBL has become a focal point for thinking about productive AI integration. - **PBL's structural conditions make AI integration productive.** [[pbl-structural-conditions-ai-2026|Rowe (2026)]] argues that PBL's core features — problem-driven inquiry, collaborative construction, facilitation, metacognitive reflection — are exactly the conditions under which AI functions as a partner in [[educational-development|professional development]] rather than a shortcut around it. The alignment is structural, not retrospective: PBL was designed around these conditions before AI existed, rooted in the recognition that professional competence requires adaptive judgment rather than routine execution. - **AI raises the ceiling on problem complexity.** The same argument holds that AI expands what category of problem PBL can engage, making "wicked problems" and complex real-world cases accessible to students who previously could not reach them. - **The artifact-as-proxy problem.** AI severs the link between a submitted artifact and the [[student-engagement|engagement]] that produced it — a problem PBL is well positioned to address because it assesses the *process* and demonstrated understanding, not just the product. This connects PBL to [[authentic-assessment|authentic]] and process-based assessment and to the knowledge base's [[cognitive-offloading|over-reliance]] literature. - **ChatGPT as [[scaffolding|adaptive scaffolding]].** [[ai-enhanced-pbl-chatgpt-scaffolding-2026|La Sunra et al. (2026)]] show ChatGPT integrated as adaptive scaffolding within an AI-enhanced PBL framework to improve [[critical-thinking|critical thinking]] and [[personalized-learning|personalized learning]] in K-12 (120 eighth graders). This treats AI as a within-PBL support rather than an answer machine. - **Quasi-experimental evidence from pre-service teacher preparation.** [[chen-osman-preservice-physics-tpack-ctd-pbl-2026|Chen and Osman (2026)]] compared an eight-week AI-supported CTD-PBL module with conventional instruction for 130 third-year pre-service [[physics-education|physics]] teachers (65 per condition), with DeepSeek used through task-specific prompt templates as a bounded scaffold — lesson-idea generation, resource organization, explanation comparison, peer-feedback prompts and revision planning — and every AI output required to pass verification against content accuracy, pedagogical appropriateness, feasibility and [[ethics]] before entering an instructional artifact. The module group reported higher post-test [[tpack|TPACK]] (M = 4.04 vs. 3.40, d = 1.02) and higher perceived [[problem-solving|collaborative problem solving]] (M = 3.62 vs. 3.05, d = 0.88), and the largest interaction effects fell on shared knowledge building and social regulation — precisely the collaborative process dimensions the structural argument above predicts AI should strengthen rather than shortcut. The authors are explicit that AI was not isolated as an independent variable: the differential change attaches to the whole condition of problem-based tasks, structured collaboration, [[teacher-role|instructor]] scaffolding, [[peer-assessment|peer feedback]] and reflective revision, and both outcomes were [[self-report-measures|self-reported]] perceptions rather than observed practice. - **Responsible-use frameworks.** [[learn-framework-responsible-genai-pbl-2026|Uden & Hwang (2026)]] advance the neuroscience-informed **LEARN** framework ([[lifelong-learning|Lifelong Learning]], Engagement, Active Processing, Reflection, Neuro-based Design) for ethically grounding generative AI use within PBL, countering [[cognitive-offloading|cognitive offloading]] and integrity erosion. - **Domain implementations.** PBL frameworks for [[engineering-education|biomedical engineering]] use GenAI for summarization and coding support with replication packages ([[pbl-biomedical-engineering-genai-2026|Nnamdi et al. 2026]]); AI-supported PBL enhances [[computational-thinking|computational thinking]] in [[educational-robotics|robotics]] ([[ai-pbl-computational-thinking-2026|AI-supported PBL for computational thinking]]); and genAI-enabled virtual patients support medical history-taking tutorials ([[genai-simulate-patient-history-pbl-2026|Mool et al. 2026]]). - **Evidence synthesis.** [[educators-engagement-ai-pbl-review-2026|Amdan et al. (2026)]] [[meta-analysis-systematic-review|systematically review]] 50 studies (2015–2024) on educators' engagement with AI in PBL within a [[human-ai-collaboration|human-computer interaction]] framework. - **GenAI in scenario-based PBL.** [[genai-scenario-based-healthcare-education-2026|Neto and colleagues (2026)]] include problem-based learning among the scenario-based approaches in their systematic review of GenAI in healthcare education, finding that [[prompt-engineering|prompt design]] functions as instructional specification and that hybrid human-AI collaboration outperforms fully automated approaches. The review highlights the need for stronger alignment of generated content with instructional frameworks and more reproducible prompting practices in PBL contexts. ## PBL vs. related pedagogies PBL is closely related to [[project-based-learning|project-based learning]] — both are learner-centered and problem/context-driven — but PBL centers on an ill-structured *problem* whose solution requires inquiry and knowledge construction, whereas project-based learning centers on producing a tangible *artifact* or project. PBL is also kin to case-based and [[inquiry-based-learning|inquiry-based learning]], which share the problem-driven, facilitation-led structure. In the knowledge base's mapping, the structural argument for productive AI integration applies to any pedagogy sharing these features, not only PBL. ## Why PBL matters for AI integration Because PBL foregrounds process, collaboration, and demonstrated understanding over artifact production, it is a natural home for [[ai-literacy|responsible AI use]]: students learn *with* AI as a partner rather than *from* AI as a substitute. The facilitator's role and the design of the problem become the levers that determine whether AI supports learning or displaces it — the same [[human-in-the-loop-ai|human oversight]] and assessment-design concerns that recur across the knowledge base. ### Productive failure and PBL [[productive-failure|Productive failure (PF)]] shares PBL's [[constructivist]] core — [[learners]] engage problems before instruction — but differs in structure: PF deliberately withholds instruction and scaffolds until *after* learners struggle to generate solutions, whereas PBL embeds facilitation throughout. The AI-era PF literature directly informs how AI should behave inside problem-based settings: [[kim-ai-productive-failure-adult-2026|Kim et al. (2026)]] show AI should preserve [[desirable-difficulties|productive struggle]] with non-directive support across PF phases; [[puech-pedagogical-steering-llm-productive-failure-2025|Puech et al. (2025)]] demonstrate [[llm]] [[intelligent-tutoring|tutors]] can be steered to withhold answers and elicit multiple solution attempts; and [[rhaimi-productivemath-2025|ProductiveMath]] uses AI to help design high-quality PF/PBL-style problems. The design question — when AI should give a problem, a hint, or an answer — is shared across both [[pedagogy|pedagogies]]. ## Connected Concepts - [[project-based-learning]] - [[active-learning]] - [[collaborative-learning]] - [[scaffolding]] - [[critical-thinking]] - [[metacognition]] - [[cognitive-offloading]] - [[authentic-assessment]] - [[generative-ai]] - [[ai-literacy]] - [[medical-education]] - [[engineering-education]] - [[higher-ed]] - [[productive-failure]] — Productive Failure ## Connected Articles - [[pbl-structural-conditions-ai-2026]] — PBL and the structural conditions for productive AI integration (Rowe 2026) - [[educators-engagement-ai-pbl-review-2026]] — Systematic review of educators' engagement with AI in PBL (Amdan et al. 2026) - [[genai-simulate-patient-history-pbl-2026]] — GenAI virtual patient for medical PBL tutorials (Mool et al. 2026) - [[learn-framework-responsible-genai-pbl-2026]] — The LEARN framework for responsible GenAI in PBL (Uden & Hwang 2026) - [[ai-enhanced-pbl-chatgpt-scaffolding-2026]] — ChatGPT as adaptive scaffolding in AI-enhanced PBL (La Sunra et al. 2026) - [[pbl-biomedical-engineering-genai-2026]] — PBL in biomedical engineering in the GenAI era (Nnamdi et al. 2026) - [[ai-pbl-computational-thinking-2026]] — AI-supported PBL for computational thinking - [[genai-counter-learner-groupthink-2025]] — GenAI agent counters groupthink in interprofessional PBL - [[dai-chatbots-problem-posing-primary-2026]] — GenAI chatbots and problem posing in primary science - [[jiang-chatgpt-inquiry-steam-review-2026]] — ChatGPT for inquiry-based learning in STEAM - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[rhaimi-productivemath-2025]] — ProductiveMath: AI to Support PF Problem Design - [[genai-scenario-based-healthcare-education-2026]] — Systematic review of GenAI in scenario-based healthcare education (Neto et al. 2026) - [[chen-osman-preservice-physics-tpack-ctd-pbl-2026]] — AI-supported CTD-PBL module with 130 pre-service physics teachers: TPACK and collaborative problem solving (Chen & Osman 2026) --- ## [Productive Failure](https://edtechdev.github.io/aied/concepts/productive-failure/) > **Productive Failure (PF)** — an instructional approach, grounded in [[constructivist|constructivist theory]] and developed by Manu Kapur, that engages learners with problems targeting concepts they have **not yet learned**, having them struggle to generate solutions *before* receiving direct instruction (Kapur, 2008; Kapur & Bielaczyc, 2012). Rather than treating failure as something to avoid, PF treats initial struggle and error as a powerful catalyst: learners activate and differentiate [[prior-knowledge|prior knowledge]], surface [[misconceptions]], and prepare to learn better from subsequent instruction — leading to deeper understanding, better retention, and enhanced [[transfer-of-learning|knowledge transfer]]. ## Questions to Consider - Have you ever learned more from failing at something first than from being shown how to do it? What made that failure 'productive' rather than just discouraging? - The page's core claim is that the *order* matters: struggling to generate solutions *before* instruction produces deeper learning. What do you think happens in the brain during that struggle that later instruction can build on? - Not all failures are equal: 'mistakes' (slips) may carry little diagnostic value, while 'errors' reveal genuine misconceptions. How might telling these apart change how you respond to a struggling learner—or how an AI tutor should respond? - The 'Safety Gap' is the divergence between a student's AI-assisted performance and their unassisted capability. When does help that makes a student look capable actually mask what they can't yet do? - An AI that withholds answers to preserve struggle can be perceived as less helpful. If you were a student, how would you react to a tutor that refused to give you the answer—and would that reaction match what's best for your learning? - How might the fear of making mistakes in front of peers (or an AI) shut down the very productive struggle the page describes? What would a 'safe space' for failing need to include? ## Introduction Productive failure is closely related to, but distinct from, other "learning from difficulty" constructs: [[desirable-difficulties]] (Bjork) focuses on introducing desirable challenges into practice; [[problem-based-learning]] and [[inquiry-based-learning]] emphasize learner-driven [[problem-solving|problem solving]]; and learning-from-mistakes/errors (error-correction learning) emphasizes the value of errorful processing and corrective feedback. Productive failure is distinctive in its two-phase structure — **generation & exploration before instruction**, then **consolidation & knowledge assembly after** — and its claim that the *order* (failure before instruction) is what produces the learning advantage. ## The core mechanism Learners generate solutions without cognitive support, relying on prior knowledge and producing suboptimal or even incorrect solutions. They then compare and contrast these attempts with the canonical solution during consolidation. The struggle: - **Activates and differentiates prior knowledge**, making learners aware of gaps and misconceptions. - **Prepares learners to learn from instruction** — they know what they don't know and can connect new material to their attempts. - **Enhances knowledge transfer and durable skills** ([[critical-thinking]], resilience, communication), reduces fear of making mistakes, and promotes positive attitudes toward learning. The PF framework has been extended through related designs including **vicarious failure**, **solution diversity**, and **adaptive guidance** (Braas et al., 2025; Brand et al., 2025). ## Learning from mistakes and learning from errors Productive failure sits within a broader family of error-centered learning theory, and the concept page covers these related ideas: ### Learning from errors Errorful processing can aid retention and conceptual change, particularly when learners are given opportunities to reflect on and reorganize their understanding (Kapur, 2008; Schwartz & Martin, 2004). In the AI era, this is operationalized in systems where learners diagnose and correct their own errors rather than receiving direct corrections — e.g., [[lukesova-clue-before-correction-2026|clue-before-correction]] tasks where AI gives guided hints and learners infer the correct solution, reducing cognitive load and supporting autonomous revision. Elaborative feedback produces significantly higher [[learning-gains|learning gains]] than verification-only feedback (Hattie & Timperley), and the timing of feedback matters. ### Mistakes vs. errors A useful distinction: **mistakes** are typically slips or lapses (often from carelessness or overload) that may carry limited diagnostic value, whereas **errors** reflect genuine misunderstanding or flawed reasoning and are more productive learning material because they reveal a misconception that can be addressed. Error analysis — identifying *why* an answer is wrong — is central to learning from errors, and is a key [[pedagogy|pedagogical]] skill (and a target for [[ai-feedback-quality|AI feedback]] design). ### The role of corrective feedback Learning from errors depends on feedback that helps learners see what was wrong and why. Clue-based and elaborative feedback (guiding learners to the correction) is more effective than simply supplying the right answer — an insight that links [[feedback]] theory directly to productive-failure design, and to how [[intelligent-tutoring|AI tutors]] should respond to student mistakes. ## Productive failure and AI in education A major theme in the knowledge base's [[research-methods-aied|research]] is the tension between AI's helpfulness and the preservation of productive struggle: - **The risk: AI erases the struggle.** Overly "helpful," Oracle-style AI that supplies answers directly can eliminate the productive struggle necessary for schema construction, creating what [[wang-safety-gap-productive-struggle-2026|Wang & Shan (2026)]] call the **Safety Gap** — the divergence between a student's AI-assisted performance and their internal, unassisted capability. This connects to [[cognitive-offloading]]: AI that substitutes for effort erodes the very capacities education builds. - **The design response: AI that scaffolds struggle.** [[kim-ai-productive-failure-adult-2026|Kim et al. (2026)]] derive five design principles for AI supporting productive-failure-based learning ([[human-ai-collaboration|human-AI collaboration]], [[usability-research|usability]], reflective design, emotional design, open knowledge), emphasizing that AI should preserve struggle while offering non-directive support. [[puech-pedagogical-steering-llm-productive-failure-2025|Puech et al. (2025)]] show [[llm]] tutors can be *steered* to follow productive-failure pedagogy (withhold solutions, elicit multiple attempts), at the cost of perceived helpfulness. CoMeT (Hou et al. 2026) prices that cost and locates the middle: its escalating tutor was 0.42 scale points more frustrating than the tutor that answered on request (p < .001) and 0.29 less frustrating than the one that only asked questions (p = .011), with distress detected in 8.4% of its sessions against 16.0% under the withholding tutor, and it surrendered the full answer in 6.1% of sessions. A floor that opens on an explicit statement of giving up rather than on frustration follows from the same design — without one, the learners who most need the demand route around it. - **AI as a tool for PF design:** [[rhaimi-productivemath-2025|ProductiveMath]] uses [[generative-ai|generative AI]] to help teachers create high-quality PF problems — addressing the challenge that designing productive-failure activities is effortful. - **AI-generated errors as provocations:** the [[pedagogy-ai-mistakes|pedagogy of AI mistakes]] deliberately leverages AI errors and [[hallucination-risk|hallucinations]] as [[teacher-role|teaching]] tools, aligning with productive-failure thinking by treating erroneous output as a cognitive provocation. This makes productive failure a central lens for [[ai-ed-evaluation|evaluating AI in education]]: the question is not whether AI helps, but whether it helps in a way that **preserves the struggle through which durable learning is built**. ## Implications for instructors and instructional design - **Design for struggle before instruction.** Sequence learning so students attempt problems before direct teaching, then consolidate — rather than the traditional instruction-first approach. - **Withhold solutions strategically.** Scaffold with hints, clues, and guiding questions ([[socratic-method]]) that keep learners cognitively engaged rather than handing over answers. - **Use AI to scaffold, not substitute.** Choose and configure AI tools that give non-directive support, preserve productive struggle, and surface errors for reflection — not answer-givers. This applies to both tutor design (steering LLMs) and classroom practice. - **Attend to the emotional side of failure.** [[affective-computing|Emotional design]], a safe space for experimentation, and reducing the fear of mistakes are essential — productive failure requires learners to be willing to struggle and fail. - **Build error analysis into learning.** Ask learners to diagnose *why* their (or AI's) answer is wrong, using elaborative/clue-based feedback, to convert errors into learning gains. - **Distinguish mistakes from errors pedagogically.** Not all failures are equally productive; attend to whether errors reflect misconceptions worth addressing. ## Connections to learning gains and other measures - **Learning gains:** PF is associated with improved [[learning-gains|learning]] compared with traditional instruction-first approaches (Kapur, 2008, 2015; Schwartz & Martin, 2004), especially on measures of transfer and deep understanding. Clue-based/elaborative feedback produces significantly higher learning gains than verification-only feedback (Hattie & Timperley). - **Transfer and retention:** the benefits of PF are strongest on transfer of knowledge to novel problems — a durable-skills outcome rather than short-term exam performance. - **Affective and [[motivation|motivational]] outcomes:** PF reduces fear of mistakes, increases [[student-engagement|engagement]], and cultivates resilience and positive attitudes toward learning. - **AI-specific measures:** PF-oriented AI research is evaluated on strategy fidelity (e.g., [[puech-pedagogical-steering-llm-productive-failure-2025|StratL's PF score]], number of elicited solution attempts) and on teacher/learner perceptions, alongside traditional learning-outcome measures. ## Connections to related concepts Productive failure connects to [[learning-theories]] (constructivism), [[desirable-difficulties]] (valuing difficulty in learning), [[problem-based-learning]] and [[inquiry-based-learning]] (learner-driven problem solving), [[metacognition]] (reflection on one's own attempts), [[cognitive-offloading]] (the risk that AI erases struggle), [[scaffolding]] (support that preserves effort), [[feedback]] (elaborative, corrective), [[prior-knowledge]] (activation and differentiation), [[transfer-of-learning]] (durable outcomes), and [[socratic-method]] (questioning to provoke reasoning). In the AI era it is a central evaluative lens: does the AI help in a way that preserves the struggle through which durable learning is built? ## Connected Concepts - [[constructivist]] - [[learning-theories]] - [[desirable-difficulties]] - [[problem-based-learning]] - [[inquiry-based-learning]] - [[metacognition]] - [[cognitive-offloading]] - [[scaffolding]] - [[feedback]] - [[prior-knowledge]] - [[transfer-of-learning]] - [[socratic-method]] - [[critical-thinking]] - [[human-ai-collaboration]] - [[hallucination-risk]] - [[ai-ed-evaluation]] - [[learning-gains]] - [[student-engagement]] ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Designing for productive struggle in educational technology - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[rhaimi-productivemath-2025]] — ProductiveMath: AI to Support PF Problem Design - [[wang-safety-gap-productive-struggle-2026]] — The Safety Gap: Restoring Productive Struggle - [[lukesova-clue-before-correction-2026]] — Clue Before Correction: ChatGPT for Autonomous Learning - [[pedagogy-ai-mistakes]] — The Pedagogy of AI Mistakes - [[finkelstein-principled-ai-education-2025]] — Principled AI in Education - [[crewscaler-ai-upskilling-framework]] — AI Upskilling Framework (productive failure as a tutoring protocol) - [[adaptive-scaffolding-contingency-comet-tutor-2026]] — Adaptive Scaffolding Needs Contingency: An AI Tutor That Escalates and Fades on What the Learner Does --- ## [Inquiry-Based Learning](https://edtechdev.github.io/aied/concepts/inquiry-based-learning/) > **Inquiry-based learning (IBL)** — a learner-centered [[pedagogy]] in which students develop understanding by posing questions, exploring independently, and constructing knowledge through a cycle of inquiry, reflection, and revision, with the instructor [[scaffolding]] rather than lecturing. In the AI era, IBL's question-driven, exploration-centered structure has become a focal point: [[generative-ai|generative AI]] and [[llm]] tools can serve as interactive "co-inquirers" that support questioning and investigation — but only when designed to preserve rather than bypass the cognitive work of inquiry. ## Questions to Consider - Inquiry-based learning centers on student-driven questions and the inquiry process itself. How is learning driven by your own question different from learning driven by someone else's? - The page asks whether AI serves as a 'co-inquirer' supporting your investigation or as a substitute that bypasses the cognitive work. When you use AI to explore a question, who is doing the inquiring? - Evidence on AI-supported inquiry is mixed: it improved creative performance and attitudes but not critical [[problem-solving]] in one study, and deepened conceptual understanding without boosting [[ai-literacy|AI literacy]] in another. What does that suggest about what AI alone can and cannot develop? - If treating AI output as authoritative risks over-reliance and superficial conclusions, how would you teach students to evaluate AI-generated answers as part of the inquiry cycle? - The page connects inquiry to productive failure — learners exploring before instruction, so struggle itself prepares learning. How might AI that withholds answers and elicits multiple attempts support that struggle rather than remove it? - What is the facilitator's role when AI lowers the friction of asking questions? If students can get answers instantly, what does the instructor now need to scaffold that they didn't before? ## Introduction Inquiry-based learning centers on student-driven questions and the inquiry process itself, typically moving through phases (orientation → conceptualization → investigation → discussion → conclusion). It is the broader family under which [[problem-based-learning|problem-based learning]] and [[project-based-learning|project-based learning]] are often nested: IBL emphasizes the *questioning and discovery process*, PBL the ill-structured *problem*, and project-based learning the tangible *artifact*. ## How inquiry-based learning appears in the knowledge base **AI as co-inquirer.** A [[meta-analysis-systematic-review|systematic review]] of ChatGPT for inquiry-based learning in STEAM ([[jiang-chatgpt-inquiry-steam-review-2026]], 24 studies) found ChatGPT used mainly in the conceptualization, investigation, and discussion phases — as learning tool, tutor, learning peer, domain expert, and [[teacher-role|teaching]] assistant — improving performance, [[critical-thinking|critical thinking]], [[student-engagement|engagement]], and motivation, while posing risks of over-reliance, hallucination, and superficial conclusions when output is treated as authoritative. **Problem posing as the starting point.** A quasi-experiment with 97 third-graders ([[dai-chatbots-problem-posing-primary-2026]]) showed GenAI [[conversational-ai|chatbots]] significantly outperformed search engines for science problem posing, improving question quality and producing more integrated epistemic [[network-analysis|networks]] while lowering cognitive load. **Cognitive-level patterns in LLM-driven IBL.** An exploratory study of 117 interview transcripts and interaction records ([[luo-ibl-patterns-llm-bloom-2026|Luo et al.]]) identified 14 interaction patterns across Bloom's cognitive levels, showing how students' prior knowledge shapes LLM use and highlighting the need for scaffolding that targets higher-order thinking stages and mitigates over-reliance. **Outcomes evidence is mixed.** An AI-supported IBL experiment in [[math-education|mathematics]] ([[mujib-ai-ibl-creative-math-2026|Mujib et al.]]) improved creative mathematical performance and attitudes but not critical problem-solving skills — suggesting AI-IBL mainly supports [[creativity]] and [[affective-computing|affective]] development. A meta-analysis of 29 experiments ([[zhao-genai-higher-order-thinking-meta-2026|Zhao et al.]]) found GenAI has a moderate positive effect on higher-order thinking, strongest for problem-solving and with 8–16 week interventions and higher [[self-regulated-learning]] learners benefiting most. A quasi-experiment with 48 pre-service science teachers in Türkiye ([[ai-supported-inquiry-photosynthesis-respiration-2026|Aydın]]) using an 8-week AI-supported guided inquiry program (integrating problem- and design-based learning) found significant group-by-time gains in conceptual understanding of photosynthesis and cellular respiration, but *no* significant effect on AI literacy or self-perceived [[computational-thinking|computational thinking]] — evidence that AI-IBL can deepen domain understanding while the development of AI/CT competencies requires more explicit, targeted design. **[[equity-in-ai-education|Equity]] and context.** A conceptual framework bridges generative-AI co-design with open educational practices to support inquiry-led STEM teaching in under-resourced contexts, using AI-generated, [[multilingual-learning|multilingual]], contextually relevant [[simulation|simulations]]. **AI-scaffolded evidence comparison.** [[ai-information-extraction-undergraduate-thesis-2026|An and colleagues (2026)]] show how an AI-powered information-extraction system shifts the undergraduate literature review from summarization toward **evidence organization and comparison** — students moving from isolated reading to cross-study comparison and evidence-based justification in support of research-based learning and thesis completion. This concretizes how [[generative-ai|AI]] can scaffold inquiry processes in [[stem-education|STEM]] undergraduate research. ## Why inquiry-based learning matters for AI integration IBL's question-driven, process-focused structure is the natural home for productive AI use: students learn *with* AI as a partner rather than *from* it as a substitute. The knowledge base's evidence converges on a core tension — AI can lower the friction of information retrieval and question formulation (reducing [[cognitive-offloading|cognitive load]] and enabling deeper [[metacognition|reflection]]), but without design scaffolding it risks over-reliance and bypassing higher-order cognition. The facilitator's role, the design of scaffolds, and the explicit teaching of evaluation skills are the levers determining whether AI deepens or displaces inquiry. ### Productive failure and inquiry [[productive-failure|Productive failure (PF)]] is the most structured cousin of inquiry-based learning: learners explore problems and generate solutions *before* direct instruction, then consolidate. Both share the premise that learner-generated attempts (even failed ones) [[prior-knowledge|activate prior knowledge]] and prepare learners to learn from instruction. The AI-era PF [[research-methods-aied|research]] sharpens how AI should scaffold inquiry without short-circuiting it: [[kim-ai-productive-failure-adult-2026|Kim et al. (2026)]] derive AI design principles for preserving struggle through problem exploration and solution generation; [[puech-pedagogical-steering-llm-productive-failure-2025|Puech et al. (2025)]] show LLM tutors can be steered to withhold answers and elicit multiple attempts; [[lukesova-clue-before-correction-2026|clue-before-correction]] tasks exemplify clue-based (vs. direct) scaffolding that keeps learners doing the reasoning. These connect inquiry and PF to the broader imperative that AI must not remove the [[desirable-difficulties|productive struggle]] through which durable learning forms. ## Connected Concepts - [[problem-based-learning]] - [[project-based-learning]] - [[active-learning]] - [[critical-thinking]] - [[metacognition]] - [[self-regulated-learning]] - [[scaffolding]] - [[cognitive-offloading]] - [[generative-ai]] - [[llm]] - [[stem-education]] - [[collaborative-learning]] - [[higher-ed]] - [[k-12]] - [[productive-failure]] — Productive Failure ## Connected Articles - [[jiang-chatgpt-inquiry-steam-review-2026]] — ChatGPT for inquiry-based learning in STEAM (systematic review) - [[dai-chatbots-problem-posing-primary-2026]] — GenAI chatbots and problem posing in primary science - [[luo-ibl-patterns-llm-bloom-2026]] — IBL patterns in LLM-driven environments (Bloom's perspective) - [[mujib-ai-ibl-creative-math-2026]] — AI-supported IBL and creative mathematical performance - [[zhao-genai-higher-order-thinking-meta-2026]] — GenAI and higher-order thinking meta-analysis - [[niri-steam-ai-literacy-review-2026]] — STEAM education for AI literacy - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[lukesova-clue-before-correction-2026]] — Clue Before Correction: ChatGPT for Autonomous Language Learning - [[ai-supported-inquiry-photosynthesis-respiration-2026]] — AI-supported guided inquiry in photosynthesis & respiration (science teacher education) - [[ai-information-extraction-undergraduate-thesis-2026]] — AI-powered information extraction supporting undergraduate thesis and research-based learning (An et al. 2026) - [[ai-assisted-inquiry-ssi-climate]] — AI-Assisted Inquiry in Socio-Scientific Issues on Climate Change --- ## [Experiential Learning](https://edtechdev.github.io/aied/concepts/experiential-learning/) > **Experiential learning** — learning through direct experience, reflection, and the application of knowledge in authentic or hands-on contexts ("learning by doing"). Drawing on Kolb's experiential learning cycle (concrete experience, reflective observation, abstract conceptualization, active experimentation), experiential approaches emphasize that learners learn most deeply when they act, observe the results, and reflect. In [[ai-education|AI education]], experiential learning includes hands-on labs, project-based work, [[educational-robotics|robotics]], [[simulation|simulations]], and real-world [[problem-solving|problem solving]]. ## Questions to Consider - Kolb's cycle describes learning by doing: concrete experience, reflection, observation, conceptualizing, then active experimentation. Think of a skill you genuinely learned. Did it follow that loop — and could a lecture alone have produced the same depth? - Experiential learning is often the default in AI, cybersecurity, and robotics education, where students learn by working with real tools. What's the argument for why hands-on, applied practice closes the theory-practice gap that lectures leave open? - Some experiential approaches now use AI assistants inside virtual labs and simulated robots. When the 'experience' itself is simulated or AI-assisted, is it still genuinely experiential — or does the absence of real consequences change what's learned? - When have you seen 'learning by doing' fail to produce learning? What conditions — reflection, feedback, a real problem — seem necessary for experience to actually teach? ## Introduction Experiential learning is closely related to [[active-learning]], [[project-based-learning]], [[embodied-learning]], and [[simulation]]. It is particularly relevant to AI, cybersecurity, and robotics education, where students develop skills by working with tools and systems in applied contexts rather than through lectures alone. A key rationale is closing the theory-practice gap in professional preparation. ### How experiential learning appears in the knowledge base's research - **Emergency substitution, with a rubric for how far it gets you.** [[ai-personas-fieldwork-experiential-learning-2026|Elhajj et al. (2026)]] document a substitution forced by the 2024 conflict in Lebanon: students in a graduate Experiential Learning course at the American University of Beirut, unable to reach communities for needs assessments, interviewed ChatGPT-generated stakeholder personas instead. Two raters scored all ten group prompts and found the split instructive — alignment with educational goals (mean 5.00) and diversity of perspectives (4.90) were strong, while authenticity and realism (4.38) and especially group dynamics and coherence (3.80, with one 30-persona focus group collapsing into sequential interviews) and limitations and gaps (3.20) were weak, the last because emotional flatness, absent contradiction and thin cultural specificity recurred in every context. The authors' conclusion is a boundary rather than a verdict: personas work as rehearsal and as a stopgap where access is impossible or unsafe, but not where emotional complexity, cultural specificity and interpersonal dynamics *are* the learning objective. Their mitigation is structural — pair simulated role-play with real interviews so students can compare, and train students to interrogate persona output for bias and generalization instead of treating it as field evidence. - **Cybersecurity labs:** [[genai-cybersecurity-ocr-multimodal-instruction-2025|LLM-assisted cybersecurity instruction]] integrates a [[generative-ai]] instructional assistant into a virtual lab platform, supporting hands-on experiential skill building. - **Robotics projects:** [[bots-blocks-project-based-robotics-education-2026|Bots and Blocks]] uses a project-based, hands-on approach to teach robotics, addressing the lack of practical experience in classic programs. - **Simulation and embodied learning:** [[edusim-llm-robotic-simulation-education-2026|EduSim-LLM]] lets beginners experiment with simulated robots, and [[embodied-learning|embodied]] robot interaction grounds learning in direct experience. ### Two forms of hands-on learning in an AI-supported curriculum A thematic review of 32 peer-reviewed studies of hands-on learning in AI-supported design education ([[hands-on-learning-necessity-age-of-ai-review-2026|Yu, Liu & Zhu, 2026]]) argues that AI reorganizes rather than replaces experiential learning, and draws a distinction the knowledge base's other sources tend to collapse: *Embodied Hands-on*, which depends on bodily action, tools, and materials, and *Cognitive Hands-on*, which develops through continued operation, judgment, and adjustment of AI-generated outputs. Both run the same action–feedback–reflection–refinement cycle, but they draw feedback from different sources — real-world material resistance in the first case, language and visual outcomes in the second — so they should not be treated as equivalent or as substitutes. The review's caution is that generation efficiency can compress the exploratory phase: several included studies report reduced exploratory sketching and gradual trial-and-error, so more iterations enabled by AI need not mean greater iterative depth, and students can miss the failure and material-constraint encounters that make hands-on work educative. Consistent with this, [[prompt-engineering|prompting]] alone did not raise [[creativity]] in the reviewed work, whereas multi-step operations (generate, modify, select) did — again locating the learning in the learner's judgment rather than in the generation. Experiential learning connects to [[active-learning]], [[project-based-learning]], [[embodied-learning]], [[simulation]], [[educational-robotics]], and [[higher-ed]] professional preparation. ## Connected Concepts - [[active-learning]] - [[project-based-learning]] - [[embodied-learning]] - [[simulation]] - [[educational-robotics]] - [[higher-ed]] - [[learning-theories]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[virtual-and-augmented-reality]] — immersive practice as deliberate experience ## Connected Articles - [[ying-genai-journalism-assessment-2026]] - [[espino-ai-business-education-review-2026]] - [[genai-counter-learner-groupthink-2025]] - [[workforce-readiness-smart-manufacturing-wrl-2026]] — Workforce Readiness Level framework for smart manufacturing in the AI era - [[zhu-e3-hot-embodied-intelligence-sustainable-learning]] — Fostering Sustainable Learning via Embodied Intelligence (E3-HOT) - [[genai-cybersecurity-ocr-multimodal-instruction-2025]] — GenAI in Cybersecurity Education - [[bots-blocks-project-based-robotics-education-2026]] — Bots and Blocks - [[edusim-llm-robotic-simulation-education-2026]] — EduSim-LLM - [[white-wu-robotics-ai-education-2026]] — Robotics and AI in Education - [[ai-lms-middle-school-longitudinal]] — AI-Integrated LMS Longitudinal Study - [[vargas-situated-learning-ai-review-2024]] - [[li-ai-science-situated-learning-teachers-2025]] - [[vargas-ai-catalyst-situated-learning-2026]] - [[panciroli-ai-literacy-episodes-situated-learning]] - [[fowlin-operationalizing-learning-principles-ai]] - [[educasim-cs1-instructional-practice]] — EducaSim: role play with simulated students for teacher training - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions - [[ai-personas-fieldwork-experiential-learning-2026]] — AI personas substituting for community fieldwork, with a five-indicator rubric for where the substitution fails (Elhajj et al. 2026) - [[hands-on-learning-necessity-age-of-ai-review-2026]] — Thematic review distinguishing Embodied Hands-on from Cognitive Hands-on in AI-supported design education (Yu, Liu & Zhu 2026) - [[shi-genai-experiential-learning-management-education-2026]] — a strategic management course reconfigured around three generative AI mechanisms --- ## [Game-Based Learning](https://edtechdev.github.io/aied/concepts/game-based-learning/) > **Game-based learning (GBL)** — the use of games themselves (digital or physical) as the medium and context for learning, where the game's mechanics, challenges, and progression carry educational content. Learners learn *through* playing. Relatedly, **gamification** applies game-design elements (points, badges, levels, leaderboards) to non-game learning activities without turning them into full games. In AI and [[educational-robotics|robotics]] education, both approaches are used to make technical content engaging and motivating. ## Questions to Consider - Game-based learning uses the game itself as the medium for learning — you learn *through* playing. Gamification just layers points, badges, and levels onto a non-game activity. How different do you think those two are in effect on real learning, versus on short-term engagement? - A comparative review found game-based learning more prevalent in informal settings while gamification dominated formal classrooms and favored project-based learning. Why do you think each approach found a different home — and what does that tell us about where each works best? - Gamification is grounded in self-determination theory — [[agency|autonomy]], competence, relatedness. If motivation is about satisfying those needs, why might a points-and-badges system succeed or fail depending on how it shapes perceived effort and attention? - [[research-methods-aied|Research]] suggests the motivational benefit of game-like and AI-supported designs depends on how they shape perceived workload and attention, not on gamification alone. When have you seen a game or badges boost engagement without actually improving learning — or vice versa? ## Introduction GBL is grounded in [[motivation]], [[student-engagement]], and [[active-learning]] theories: games provide intrinsic motivation, immediate feedback, and authentic problem contexts. It overlaps with [[simulation]], [[project-based-learning]], and [[educational-robotics]]. GBL is particularly relevant to [[educational-robotics]], [[computational-thinking]], and [[cs-education]], where games can make abstract technical concepts concrete and fun. ### How GBL appears in the knowledge base's research - **Robotics education:** [[game-based-gamified-robotics-education-review-2026|A comparative systematic review]] of game-based learning and gamification in robotics education found GBL more prevalent in informal settings, while gamification dominated formal classrooms and favored [[project-based-learning|project-based learning]]. - **Robot-mediated games:** [[remind-robot-mediated-roleplay-antibullying-2026|REMind]] is a robot-mediated role-play game for anti-bullying intervention, and [[motibo-digital-storytelling-robots-motivation-2026|MotiBo]] uses interactive [[storytelling-in-education|digital storytelling]] to boost motivation. - **AI [[conversational-ai|conversational agents]] in simulation games:** Wenzel, Geiger, and Liening (2026) derive the CAIS-GBL framework — four design principles and fifteen design features for AI conversational agents in digital game-based learning — from theory-driven meta-requirements spanning cognitive, motivational, [[affective-computing|affective]], and [[sociocultural-learning|socio-cultural]] engagement, with an [[equity-in-ai-education|equity]]-by-design stance. Their instantiated agent (Lara) in a business simulation game was positively received for cognitive and [[community-of-inquiry|social presence]] and [[self-regulated-learning]] support, addressing the common gap of limited [[formative-assessment|formative]] feedback and structured reflection in simulation games. ### Gamification **Gamification** is the application of game-design elements (points, badges, levels, leaderboards, challenges, progress bars) to non-game contexts to motivate and engage users. Unlike game-based learning — where learning happens *through* a game — gamification layers game mechanics onto an existing learning activity without turning it into a full game. It is used in education to boost motivation, [[student-engagement]], and persistence, and is widely applied in formal classroom settings. Gamification is grounded in motivational theory, particularly [[self-determination-theory]] (supporting autonomy, competence, and relatedness) and behavior-change frameworks. It has shown particular synergy with [[project-based-learning]] in applied domains like robotics. In the knowledge base's research: - **Robotics education:** the comparative review found gamification dominated formal classrooms in robotics education (p < .001) and strongly favored [[project-based-learning|project-based learning]] (p = .009), while game-based learning was more common in informal settings. - **Engagement and motivation:** gamification is used across the knowledge base to increase learner engagement and motivation in AI, [[cs-education|programming]], and [[stem-education|STEM]] learning contexts. [[genai-motivation-engagement-2026|Generative AI, motivation, and engagement]] research examines how game-like elements combine with AI to sustain learner interest. Two 2026 studies extend this by comparing gamified and AI-supported conditions against traditional instruction: [[nasa-tlx-workload-gamified-ai-2026|a NASA-TLX study]] measured perceived workload across traditional, gamified, and AI-supported learning conditions, and [[arcs-motivational-ergonomics-gamified-ai-2026|an ARCS study]] examined motivational "ergonomics" in gamified and AI-supported learning with implications for [[professional-training|workplace training]]. Together they clarify that the motivational benefit of game-like and AI-supported designs depends on how they shape [[motivation|perceived effort]], workload, and attention (e.g., ARCS attention/relevance dimensions), not on gamification alone. GBL and gamification together connect to [[educational-robotics]], [[student-engagement]], [[motivation]], [[self-determination-theory]], [[active-learning]], [[simulation]], [[project-based-learning]], and [[computational-thinking]]. ## Connected Concepts - [[educational-robotics]] - [[student-engagement]] - [[motivation]] - [[self-determination-theory]] - [[active-learning]] - [[simulation]] - [[project-based-learning]] - [[computational-thinking]] - [[cs-education]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[virtual-and-augmented-reality]] — immersive and gamified practice overlap in design and evidence ## Connected Articles - [[game-based-gamified-robotics-education-review-2026]] — Game-Based and Gamified Robotics Education - [[remind-robot-mediated-roleplay-antibullying-2026]] — REMind - [[motibo-digital-storytelling-robots-motivation-2026]] — MotiBo - [[bots-blocks-project-based-robotics-education-2026]] — Bots and Blocks - [[white-wu-robotics-ai-education-2026]] — Robotics and AI in Education - [[genai-motivation-engagement-2026]] — Generative AI, Motivation, and Engagement - [[nasa-tlx-workload-gamified-ai-2026]] — NASA-TLX workload across gamified/AI conditions - [[arcs-motivational-ergonomics-gamified-ai-2026]] — ARCS motivation and AI-supported gamification - [[conversational-agents-business-simulation-gaming-2026]] — CAIS-GBL framework for AI conversational agents in business simulation games (Wenzel et al. 2026) - [[play-ai-pre-k-kindergarten-ai-literacy-2026]] — Play With AI (PL-AI): play-centered AI literacy curriculum for pre-K and kindergarten (Lee 2026) --- ## [Learning by Teaching](https://edtechdev.github.io/aied/concepts/learning-by-teaching/) > **Learning by teaching (LbT)** — the instructional framework, grounded in the protégé effect, in which students deepen their understanding by explaining material to a peer, tutee, or agent. Decades of work in LbT and peer tutoring show that explaining concepts, anticipating misunderstandings, and responding to questions consolidate understanding and support transfer. In the AI era, **teachable agents** — and increasingly **LLMs configured as novice tutees** — operationalize LbT at scale, positioning students as instructors who must explain, correct, and fill gaps. ## Questions to Consider - Recall a time you truly understood something only after explaining it to someone else. What was happening mentally — and why do you think teaching produces deeper understanding than just studying alone? - A common view is that teaching is for experts, and novices have nothing to offer. Yet 'learning by teaching' rests on the opposite premise: preparing to teach forces you to organize knowledge and find your own gaps. How does that reframe who benefits from teaching? - The page describes 'teachable agents' — software that students teach as part of learning. With an LLM, you can configure a chatbot as a fallible novice tutee that asks questions and makes mistakes. What would you need to design into such a tutee for it to actually improve learning rather than just chat? - One challenge is 'engineering fallibility': AI models are trained to give expert, fluent answers, which is the opposite of the struggling novice the learning-by-teaching paradigm wants. Why might an error-prone tutee be more effective for learning than a correct one? - A ChatGPT-based teachable agent improved learning but its tendency to generate correct code limited error-correction practice. How might a tool that always gives the right answer actually shortchange the learner who needs to practice spotting and fixing mistakes? - If you were to design a learning-by-teaching activity for your own class, what would make the teaching task *consequential* enough that students put real effort into it rather than copy-pasting an answer? ## Introduction Learning by teaching is the finding, usually traced to the protégé effect, that preparing to teach — and actually explaining to another person or a teachable agent — produces deeper processing than studying alone. The demands of teaching force learners to organize knowledge, anticipate [[misconceptions|misunderstandings]] and generate explanations, which surfaces gaps in their own understanding and strengthens [[metacognition]]. AI enters the idea from both directions: [[intelligent-tutoring|tutoring systems]] and teachable agents can play the student, while a growing literature asks what happens to learning when the machine, rather than the learner, supplies the explanation ([[generative-ai]], [[cs-education]]). ## The Protégé Effect Learning by teaching rests on the finding that preparing to teach and actually explaining to another person produces deeper processing than studying alone. The demands of teaching — articulating ideas, anticipating misunderstandings, and answering questions — force learners to organize knowledge, identify gaps in their own understanding, and generate explanations that support retention and transfer. Benefits are most evident in [[collaborative-learning]] contexts and in well-structured domains that support teachable agents (e.g., Betty's Brain). The protégé effect names the mechanism: students put forth more effort and reflect more deeply when they feel responsible for teaching something, so they clarify [[misconceptions]] and fill gaps through explanation and [[metacognition]]. ## Teachable Agents: From Rule-Based to Conversational **Teachable agents** are the software systems through which learning by teaching is operationalized — a learner teaches a system as part of learning. Traditional teachable agents were rule-based or retrieval-based and could respond only to limited commands; their key limitation was an inability to engage in natural-language dialogue. [[llm|Large language models]] change this: they can flexibly adopt roles via [[prompt-engineering|prompting]] — including the role of a "tutee" that asks questions or makes mistakes — and engage in open-ended dialogue, enabling LbT in less-structured domains (writing, vocabulary) than was previously possible. The knowledge base's evidence base traces this shift to **conversational, LLM-based teachable agents**: - **ChatGPT as a teachable agent** ([[chatgpt-teachable-agent-programming-lbt-2024|Chen et al.]]) supports LbT in programming, improving knowledge gains, programming ability, and [[self-regulated-learning|self-regulated learning]] — though its tendency to generate correct code limits error-correction practice. - **Explique at scale** ([[explique-teachable-agent-algorithms-546-students-2026|Wang et al.]]) deployed an AI teachable agent (Algorithm Apprentice) to 546 students over an 11-week semester, finding that explanation-oriented dialogue predicts fewer incorrect quiz submissions, while external-content reuse predicts more. - **Vocabulary teaching** ([[teaching-ai-vocabulary-lbt-llms-2026|Uchida et al.]]) used an LLM as a student to generate dynamic questions, improving retention at 3 and 7 days. ## Engineering Fallibility: LLMs as Novice Tutees A central design challenge for LLM-based teachable agents is that LLMs are trained to produce expert-level, fluent responses by default — the opposite of the fallible novice the LbT paradigm wants. Making an LLM a good tutee requires **engineering fallibility**: - **[[prompting-teachability-novice-personas-lbt-2026|Prompting for teachability]]** (Miller & Bosch) found that constraint-based prompts explicitly forcing error production (e.g., "answer incorrectly" or "get 2–3 wrong") elicit novice-like behavior far more reliably than persona-, misconception-, or uncertainty-based prompts. - **[[socrates-students-instructors-llms-lbt-2025|Engineered knowledge gaps]]** (Yang et al.) design problems the LLM cannot solve without knowledge only the student possesses, making teaching a necessity and countering the passive over-reliance of LLM-as-tutor use. - **Explique's apprentice constraints** (Wang et al.) instruct the tutee to (a) stay a novice, (b) keep requesting clarification until the student's explanation is accurate, and (c) never reveal the target explanation — and to *resist* students who try to reverse the roles and have the tutee explain back. ## Questioning, Self-Regulation, and Active Learning Two further affordances recur across the knowledge base: - **Questions identify knowledge gaps.** LbT systems use learner-generated questions to expose gaps and reinforce comprehension, and [[teaching-ai-vocabulary-lbt-llms-2026|LLM-generated questions]] replace rigid template-based generators. - **LbT [[scaffolding|scaffolds]] self-[[regulation]].** Teaching a [[conversational-ai|conversational agent]] fosters [[self-efficacy]] and the implementation of self-regulated learning strategies, and connects LbT to [[desirable-difficulties]] — the effortful act of explaining and correcting is itself a productive struggle that AI's friction-removal would otherwise erase. ## Why It Matters in AI Education Learning by teaching is the constructive, [[active-learning]] counterpoint to the dominant LLM-as-tutor pattern. Where a tutor gives answers (and risks [[cognitive-offloading|Over-Reliance]]), an LbT setup makes the student the teacher, forcing explanation, gap-detection, and knowledge construction. This positions LbT as a key strategy for turning [[generative-ai|generative AI]] from a crutch into a tool for deeper learning, and connects to [[desirable-difficulties]], [[active-learning]], and [[constructivist]] [[pedagogy]]. ## Putting Learning by Teaching into Practice ### Design patterns for an AI tutee The [[research-methods-aied|research]] above converges on a few reusable patterns for turning a default-expert LLM into a productive tutee: - **Constraint-based novice prompts (most reliable).** Rather than asking the model to "pretend to be a confused student," explicitly force fallibility and a teaching loop, e.g.: *"You are a novice student learning about [concept]. Ask me to teach it to you. Ask clarifying questions and deliberately get 2–3 things wrong during our conversation. Never state the correct answer yourself — wait for me to explain, then tell me whether I made sense."* - **The reverse-teaching guard.** Add a rule that the tutee must *decline* to explain the answer back when the student tries to flip the roles: *"If I ask you to solve the problem or explain the concept, remind me that I'm the teacher and ask me to explain it instead."* Explique shows this resistance is what preserves the LbT interaction. - **Engineered knowledge gaps.** Structure the task so the model *cannot* answer without information only the student holds — the student's knowledge becomes genuinely necessary, not optional. This converts the interaction from optional chat into required teaching. - **An external success criterion.** Give the teaching a real consequence — a gatekeeper quiz that unlocks only after the student teaches successfully (Explique), or a code-judging platform the student must make the agent's output pass (Chen). Accountability is what sustains genuine effort and prevents the whole exercise becoming a checkbox. ### Tips for instructors - **Make the teaching task consequential, not busywork.** The strongest evidence for [[student-engagement|engagement]] comes from activities that matter — Explique gated a graded quiz behind the teaching exercise; Chen tied the tutee's output to passing a judging platform. If teaching is purely optional, students will rationally skip the hard part. - **Give students a teaching protocol, not just a chat window.** Scaffold the interaction with a structure — "explain the concept → give a concrete example → answer the tutee's questions → check for understanding" — so open-ended dialogue becomes a deliberate teaching sequence rather than aimless conversation. - **Address content-dumping head-on.** Explique found that direct copy-paste of external content rose from under 15% to 30–35% of interactions by the end of the semester. Tell students why pasting defeats the purpose, and consider an accountability step (e.g., "explain the agent's misunderstanding in your own words"). - **Pair LbT with debugging practice.** Because AI writes correct code, students may lose error-correction practice. Deliberately ask the tutee to *misimplement* something, or follow the teaching session with a bug-finding task, so debugging stays in the loop. - **Watch the effort gradient.** Expect novelty to fade; plan to vary the target concepts, add challenge, or rotate which students take the [[teacher-role|teaching role]] to sustain cognitive effort across a term. ### Tips for developers - **Prefer hard constraints over persona alone.** Prompting for "uncertainty" or "a student persona" is unreliable; explicitly force errors and a clarification loop. (See [[prompting-teachability-novice-personas-lbt-2026]].) - **Build a completion criterion.** Define *when the student has explained enough* (Explique used an LLM tool function keyed to the concept's learning objectives) so the interaction ends on understanding, not on a time limit or a fixed turn count. - **Log and code the dialogue.** Explique used word-per-minute detection plus LLM semantic coding to classify interactions as Detailed / Minimal / External Content Use — that signal is how you detect circumvention and declining engagement before it becomes a problem. - **Give instructors a dashboard.** Completion rates and [[qualitative-research|qualitative]] patterns in teaching interactions let a human intervene when effort drops (Explique's instructors monitored exactly this). ### Implications and open questions - **LbT is a scalable antidote to AI over-reliance** — it inverts the tutor/student role and keeps the learner cognitively active, which matters more as AI gets more fluent and more "helpful." - **Fallibility is a feature, not a bug.** A tutee that is *too* correct removes the error-correction and gap-detection that make LbT work; design for the productive struggle rather than against it. - **Open questions remain:** How do LbT interactions sustain beyond a semester as novelty fully fades? Does LbT transfer to non-CS, less-structured domains at the same scale? Can automated dialogue coding become a practical, real-time engagement monitor for instructors? And how do we keep the teaching role meaningful for *every* student rather than a motivated few? ## Connected Concepts - [[generative-ai]] - [[active-learning]] - [[constructivist]] - [[scaffolding]] - [[self-regulated-learning]] - [[metacognition]] - [[desirable-difficulties]] - [[cognitive-offloading]] - [[cs-education]] - [[collaborative-learning]] - [[intelligent-tutoring]] - [[pedagogical-agent]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education ## Connected Articles - [[chatgpt-teachable-agent-programming-lbt-2024]] — ChatGPT as a teachable agent in programming - [[explique-teachable-agent-algorithms-546-students-2026]] — Explique: teachable agent for 546 students - [[prompting-teachability-novice-personas-lbt-2026]] — Designing novice personas for teachability - [[socrates-students-instructors-llms-lbt-2025]] — Students as instructors of LLMs (Socrates) - [[teaching-ai-vocabulary-lbt-llms-2026]] — Vocabulary learning by teaching AI - [[knowloop-confusion-to-consolidation-2026]] — Teach-back consolidation in a conversational review system - [[simulating-students-java-programming-errors-llms]] — Simulating student errors with LLMs --- ## [Scaffolding](https://edtechdev.github.io/aied/concepts/scaffolding/) > **Scaffolding** — structured support that helps learners accomplish tasks they cannot yet complete independently, with support fading as competence grows. In [[ai-education]], scaffolding is the primary design principle for ensuring AI tools support learning rather than replace it. ## Questions to Consider - Think of a time someone 'helped' you with something you were learning, and the help did the work so well you learned less. Where's the line between support that lets you grow and support that replaces you? The page argues this is the central design question for AI tutors. - Scaffolding is rooted in the Zone of Proximal Development — enough support to enable progress, not so much that learning is bypassed. What does 'too much support' look like in an AI tutor, and can you detect it from the student's behavior alone? - A key finding: students often *prefer* the more directive AI tutor roles even though they *perform better* with collaborative peer and [[teacher-role]]-assistant roles. Does learner preference reliably track what's best for learning — and what does this divergence imply for letting students choose their own scaffolding? - Scaffolding that never fades creates dependency. The page notes automated scaffolds risk staying static instead of being withdrawn as competence grows. Why is 'fading' essential, and why might an AI system fail to do it if it isn't deliberately designed to? - The design principle is 'scaffold, do not substitute.' Students themselves asked for AI that 'does not provide any solutions for you — you still learn as you have to find the correct answer yourself.' Does that match how you've experienced effective help, or have you preferred the shortcut even while knowing it cost you? - Set a goal before reading: pick a task you teach, and sketch what a hint looks like that preserves the learner's effort versus an answer that removes it. How will you know your hints are in the productive-struggle zone? ## Introduction ### How scaffolding appears in AIED - **Prompt-based scaffolding:** [[guided-llm-scaffolding-independent-learning|Guided LLM scaffolding]] teaches structured [[prompt-engineering|prompting]] as a learning intervention. [[scaffolding-critical-engagement-genai-minority-students|Critical engagement scaffolding]] uses [[culturally-relevant-pedagogy|culturally responsive]] approaches. - **Socratic scaffolding:** [[socratic-method|Socratic AI dialogue]] withholds direct answers, using questions to guide discovery — a form of [[desirable-difficulties]] scaffolding. - **Adaptive fading:** [[intelligent-tutoring|Intelligent tutoring systems]] adjust scaffolding based on [[knowledge-tracing]] estimates, providing more support for unmastered concepts and less for known ones. - **Need-triggered adaptive scaffolding in [[medical-education|clinical]] interview training:** the MeduAI-SP [[rct|randomized trial]] ([[ai-standardized-patient-scaffolding-medical-2026|Yang et al., 2026]]; N = 100 third-year medical students) operationalized scaffolding as adaptive, need-triggered support: a turn-level evaluator agent flagged when a learner was stuck, omitted key history, risked premature closure, or impaired rapport, and only then did a [[pedagogical-agent|tutor agent]] issue a [[socratic-method|Socratic]] prompt. Expert annotation of 207 consultations showed that scaffolding need was strongly phase-dependent, rising from about 16.1% of student utterances early in the encounter to 34.8% late (24.1% flagged overall) — empirical evidence that novice consultation support is most needed during integration and diagnostic reasoning, not initial information gathering. The scaffolding condition outperformed a structured progressive-disclosure control on the final OSCE-aligned [[summative-assessment|examination]] (71.8% vs. 55.6%; β = 16.4 percentage points; P = 3.30e-4), with the largest gain in communication (Hedges' g = −0.79). - **Hint systems:** [[correct-answer-trap-ai-tutor|AI tutor hint research]] examines when hints help versus when they encourage [[cognitive-offloading|Over-Reliance]]. - **Conceptual scaffolds:** [[concept-catalyst-engineering-scaffolds|Concept Catalyst]] and [[rethinking-scaffolding-llm-tutors|LLM tutor rethinking]] explore design patterns for cognitive support. - **"Scaffold, do not substitute" as a design principle:** [[substitution-to-scaffolding-ai-harm-cycle-2026|Favero et al. (2026)]] argue that the central risk of AI in education is misalignment — AI that substitutes for human effort erodes the capacities education is meant to build — and derive a single design principle, *scaffold, do not substitute*. Scaffolding must be a first-class capability of [[ai-technologies|AI systems]]: knowing *when to withhold an answer, ask a question, surface uncertainty, or present alternative perspectives*. Their analysis of student essays shows learners themselves converge on this — asking for AI that "does not provide any solutions for you, you still learn as you have to find the correct answer yourself." The principle positions scaffolding as the alternative to a self-reinforcing harm cycle of substitution across cognition, [[agency]], emotion, and [[ethics]]. - **Scaffolding embedded in the medium, not bolted on:** [[wang-chatgpt-comments-video-learning-scaffolding-2026|Wang, Du and Jin (2026)]] operationalize four scaffolding principles inside a video player — [[sociocultural-learning|ZPD]]-based adjustment of comment depth, *fading* scaffolding (knowledge support thins over the timeline), *distributed* scaffolding (every comment is either knowledge or emotional support), and cognitive-load-derived **timing** — by computing frame-level video entropy and inserting comments only in low-information intervals. Their four-condition ablation with 20 learners found the entropy-timing module produced the strongest and most robust effect on perceived quality when removed (Z = −2.85, r = 0.45, p = .004), which makes the *scheduling* of scaffolding a measurable design variable rather than a packaging detail. The study also carries a warning for automated scaffolding: [[generative-ai|ChatGPT]]-generated comments were consistently harder to read, less lexically diverse, and less topically aligned than instructor comments, with the largest relevance gap in emotional support ([[ai-feedback-quality|feedback quality]]). - **Preferred scaffolding is not always the most effective:** [[preferred-scaffolding-ai-mathematical-modeling|Zhu, Yang and Yang (2026)]] found in a within-subjects experiment that students performed best with Peer and Teaching Assistant AI roles (which foster [[collaborative-learning|collaborative]] reasoning) yet preferred the more directive Tutor and Excellent Student roles — a divergence between preference and performance that cautions against equating learner preference with effective scaffolding in AI-supported mathematical modeling. - **Automated scoring as a deliberate scaffold:** [[chen-automated-scoring-interpreting-self-regulated-learning-2026|Chen and Liu (2026)]] treated an automated interpreter-scoring system as [[formative-assessment|formative]] scaffolding rather than as a measurement instrument, in a 14-week comparison with 46 English Translation and Interpreting sophomores. The scaffold converted into a gain only where the deficit it targeted was decomposable, reliably scored and sensitively scaled: linguistic accuracy and logical coherence rose, while information fidelity and delivery fluency stayed flat, fidelity being the dimension on which automated and human ratings disagreed most (r = 0.12). The cycle it reinforced was also partial. Only practice-phase execution and monitoring correlated with score gains (r = 0.42), while pre-learning planning sat near the scale midpoint (M = 3.01) and students set goals from the previous score rather than from the task ahead. A scaffold can be well placed inside the performance phase and still leave the planning that would make it unnecessary untouched. - **Closed-loop scaffolding on the learner's current boundary:** [[zhu-adaptive-teaching-assistance-genai-big-data-2026|Zhu, Luo and Li (2026)]] built a music-education loop in which multimodal error detection over performance audio and score becomes the reward that steers [[reinforcement-learning|reinforcement]]-generated practice tracks, so difficulty follows the learner's current boundary instead of a fixed syllabus: generated practice trajectories matched learner skill profiles at a peak cosine similarity of 0.962, and the Group x Time interaction favored the scaffolded group in a 12-week quasi-experiment with 120 undergraduates (beta = 0.52, 95% CI [0.31, 0.73]). Two limits follow for automated scaffolding. Detection is triage rather than judgment, since rhythm error precision of 89.7% means roughly one flagged error in ten is a false alarm, and blind expert review rated support for musical expression the system's weakest area, leaving interpretation to the teacher. ## The ZPD connection [[sociocultural-learning|Vygotsky's Zone of Proximal Development]] provides the theoretical foundation: scaffolding targets the space between what learners can do independently and what they can achieve with support. AI tools should operate in this zone — enough support to enable progress, not so much that learning is bypassed. A configuration that inverts the usual direction of adaptation appears in Sidorkin's (2026) graduate course, where the learner rather than the system set the support level: weekly readings were generated on demand and students dialed comprehension level through iterative prompting (pacing, definitions, vocabulary density, depth). Analysis of three reading logs found definitional markers 3.4x to 8.7x more frequent in AI responses following comprehension-oriented prompts than in baseline explanatory text, with the clearest cases building a definitional layer and then a numbered procedural one, which makes scaffold density a measurable property of the learner's request rather than only of a system's mastery estimate. Requiring at least three follow-up questions per reading made that dialing routine, turning the text into an interaction that surfaced comprehension gaps the instructor otherwise would not see. Scaffold form has to match what the learner can already hold, not only what the task requires. In a systematic scoping review of 24 evidence sources on [[creativity|creative thinking]] in children aged 6 to 15, [[niu-genai-children-creative-thinking-cognitive-development-review-2026|Niu et al. (2026)]] report that younger children lacked the precise linguistic and metacognitive control that text-based prompting demands and needed multimodal, adult-facilitated interfaces, while text-based [[llm|LLMs]] showed stronger reported outcomes with older children and young adolescents. Prompt dependence was strongest in the lower grades, and no included study examined developmental readiness thresholds, so the authors treat modality and [[scaffolding|scaffolding]] choices as an open design question to be settled by the learner's capacity rather than by convenience. ### Connections Scaffolding connects to [[cognitive-offloading|Over-Reliance]] (scaffolding that doesn't fade creates dependency), Cognitive Load Theory (scaffolding manages cognitive load), [[feedback|Feedback Loop]] (scaffolding provides [[formative-assessment|formative]] feedback), and [[ai-literacy]] (learners must recognize when scaffolding is beneficial vs. when it displaces learning). Agents must scaffold dynamically, not statically: [[agentic-ai-pedagogical-best-practice-2026|Woollaston et al. (2026)]] identify that automated scaffolds risk staying static instead of being withdrawn as competence grows, and recommend dynamic scaffolds that adapt and fade — a key guardrail for [[agentic-ai]]. CoMeT (Hou et al. 2026) supplies both the separation and the warrant for *when* to fade: its support climbs one rung each time a learner does not use it and drops to the lightest rung on take-up, holding [[metacognition|metacognitive demand]] statistically equivalent to a withholding tutor (p_TOST = .004) while delivering an artifact in 48.1% of sessions against 23.7% — more system labor, not less. The trigger it validated is aim rather than depth: after a full demonstration the tutor later conceded 30.3% of what was still open, after a pasted artifact 25.9%, after a bare assertion 20.4% and after a request to build 33.7%, whereas turns not aimed at the decision under support drew a later concession 40.3% of the time against 28.8% for aimed turns (11.5-point difference, 95% bootstrap interval [2.1, 22.4]). Take-up was sparse — 37.0% after the first ask and 21.8% after the third — so a fade rule keyed to learner effort would read a non-answer as readiness. Scaffolding must be situation-appropriate, not maximal: [[zhang-tutormoments-2026|Zhang et al. (2026)]] introduce TutorMoments, which evaluates whether LM tutors scaffold only when support is needed, push for rigor when the student is ready, and avoid over-scaffolding (reducing cognitive demand more than the situation requires). Minimally prompted frontier models default to over-scaffolding at the expense of productive struggle. - **AI that scaffolds productive struggle.** [[kim-ai-productive-failure-adult-2026|Kim et al. (2026)]] derive AI design principles (non-directive support, reflective design, [[human-in-the-loop-ai]]) that keep scaffolding in the productive-struggle zone rather than collapsing to answer-giving; [[puech-pedagogical-steering-llm-productive-failure-2025|Puech et al. (2025)]] show [[llm]] tutors can be steered to give help only when strictly necessary — scaffolding that preserves the learner's own effort. **Scaffold withdrawal as the enforcement mechanism for verification.** [[kumar-genai-computing-education-systematic-review-2026|Kumar, Wongsirichot and Nanthaamornphong (2026)]] synthesize 72 computing-education studies and locate scaffold withdrawal, alongside guardrail tools and [[self-regulated-learning]] designs, as one of three ways courses enforce critical engagement with AI output — structurally (constraining what the tool returns), procedurally (reflection logs, self-testing) and temporally (progressively restoring conditions under which independent reasoning is required). Their evidence is that efficiency gains under AI assistance do not [[transfer-of-learning|transfer]] to unaided performance, and that the failure mode — the *pseudo-apprenticeship* pattern, where students watch AI generate code without performing the task — is exactly modeling without whole-task practice. Graduated access therefore functions as a fading schedule for a powerful new form of support, and the review grounds it in 4C/ID: assistance helps only when the learner already has enough schema to engage critically with it (the zone of proximal development, [[cognitive-offloading]]). ## Rule-Guided vs. Ad-Hoc Scaffolding - **Rule-guided vs. ad-hoc scaffolding.** Looi, Liu, and Sun (2026) formalize a distinction central to scaffolding design: **rule-guided scaffolding**, in which tutoring is governed by an auditable three-layer architecture (diagnosis → intent selection → constrained response generation), versus **ad-hoc scaffolding**, where helpful moves are difficult to audit and replicate. Their primary-school math study showed rule-guided scaffolding improves interactional consistency, reduces premature answer-giving and early closure, and sustains cognitive [[student-engagement|engagement]] — evidence that explicitness and auditability of scaffold moves matter for both consistency and learning in procedural domains. - **Scaffolding as the constrained pathway between bypass and offloading.** The Neuroplasticity-[[student-ai-interaction|AI Interaction]] Model names scaffolding as the third of three pathways for LLM help, alongside direct bypass and cognitive offloading, and defines it by whether the model preserves the effortful processing the task is meant to train ([[naim-bypass-offload-scaffold-llm-learning-2026]]). The model's calibration evidence is a natural experiment in scaffold constraint: unrestricted GPT-4 access in a study of nearly 1,000 high-school [[math-education|mathematics]] students produced a 48% practice gain but a 17% deficit on the unassisted exam, while hint-constrained GPT Tutor produced a 127% practice gain with the exam deficit largely eliminated. The design lesson matches the rule-guided versus ad-hoc distinction above at a coarser grain: it is the constraint on what the tutor is allowed to supply, not the presence of a tutor, that determines whether the scaffold is removed successfully. The rule-guided versus ad-hoc distinction reaches past the tutor and into the assignment itself. A [[ai-integration-instructional-design-collaboratory-2026|cross-institutional faculty collaboratory]] found that critique of AI output did not happen on its own, so requirements to audit, compare, revise and justify generated content had to be written into the task, and those who protected independent disciplinary analysis before AI entered found candidates could evaluate AI output more critically. [[bondurant-shaughnessy-ai-pedagogies-practice-2026|Rehearsal evidence]] agrees: pre-service teachers practicing with AI partners used more probing and exploring questions when structured post-rehearsal feedback was provided. In both cases support is specified in advance rather than improvised, separating designed guidance from the ad-hoc help that is hard to audit and replicate. - **A correctly constrained scaffold still fails if the constraint is not administered.** [[ai-literacy-tool-design-programming-education-2026|Azimi (2026)]] built what the distinction above prescribes — a budget of 25 hints per session, a 15-minute cap on AI use, a required end-of-session reflection, and no generated code — and randomized 33 master's students between it and unrestricted [[generative-ai]] use across seven weeks. Assignment performance and concept-inventory gains did not differ between conditions. The hint budget did not act as a rationing mechanism: some students spent most of it on the first problems and had none left for the demanding ones, others finished with most unused, and within the Coach condition it was the students who already had a deliberate strategy for spending a hint who scored higher. The constraint raised reported [[self-efficacy|confidence]] and followed students out of the classroom as a self-questioning habit (whether a question was worth asking the tool), but it rewarded existing self-governance rather than developing it. Auditability of a scaffold and a learner's capacity to use it are separate conditions. - **A scaffold the learner cannot verify is a fluent substitute.** Because chemistry reasons across observable phenomena, particulate models and symbolic notation, a generated answer can be locally persuasive and globally wrong. [[vega-baudrit-genai-university-chemistry-education-review-2026|Vega-Baudrit and Rivera Álvarez (2026)]] place [[scaffolding|scaffolding]] on the productive side of their Presage-Process-Product analysis and uncritical copying on the failure side, and require scaffolds to demand representational translation in both directions, since a response that describes neutralization correctly can still claim that every equivalence point has a pH of 7. Verification is designed into the task rather than announced as a rule: students identify a false assumption, correct a unit or mechanism error, compare a symbolic structure with a submicroscopic model, or justify rejecting a generated answer. Because students cannot verify what they do not yet understand, [[prior-knowledge|prior knowledge]] and scaffolding come first, and prompting is treated as an epistemic act in which the learner specifies the constraints the answer must satisfy. ## Connected Concepts - [[problem-based-learning]] — PBL embeds fading scaffolds around ill-structured problems - [[learning-by-teaching]] — Scaffolding knowledge building through explanation - [[sociocultural-learning]] — Vygotskian foundation: ZPD and socially mediated learning - [[cognitive-offloading]] — Scaffolding that never fades creates over-reliance - [[feedback]] — Scaffolding delivers formative feedback as support fades - [[ai-literacy]] — Recognizing when scaffolding supports vs. displaces learning - [[intelligent-tutoring]] — ITSs adapt scaffold intensity to mastery estimates - [[socratic-method]] — Questioning that withholds direct answers - [[metacognition]] — Scaffolds that build self-monitoring and self-regulation - [[adaptive-learning]] — Adaptive systems modulate support within the learner's ZPD - [[learning-design]] — Scaffolding is a core instructional-design strategy - [[help-seeking]] — Scaffolding shapes when and how learners request help - [[teacher-role]] — Teachers scaffold, then fade as competence grows - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[agentic-ai]] - [[productive-failure]] — Productive Failure ## Connected Articles - [[kumar-genai-computing-education-systematic-review-2026]] — Scaffold withdrawal as the mechanism enforcing verification (VIE framework) - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[wang-chatgpt-comments-video-learning-scaffolding-2026]] — Entropy-timed AI comments as in-video knowledge and emotional scaffolding (Wang, Du & Jin 2026) - [[adaptive-ai-scaffold-collaborative-problem-solving-2026]] - [[guided-llm-scaffolding-independent-learning]] — Guided LLM prompting as a structured learning intervention - [[scaffolding-critical-engagement-genai-minority-students]] — Culturally responsive critical-engagement scaffolding with GenAI - [[rethinking-scaffolding-llm-tutors]] — Design patterns for scaffolding in LLM tutors - [[concept-catalyst-engineering-scaffolds]] — Concept Catalyst scaffolds for conceptual change - [[correct-answer-trap-ai-tutor]] — When hints help vs. when they encourage over-reliance - [[critical-thinking-genai-scaffolding]] — Scaffolding critical thinking with GenAI - [[veriforge-narrative-drafting-scaffolding-2026]] — Scaffolded narrative drafting with Veriforge - [[ai-cognitive-partner-co-regulation-learning]] — AI cognitive partner supporting co-regulation of learning - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[preferred-scaffolding-ai-mathematical-modeling]] — Preferred scaffolding in AI-supported mathematical modeling - [[agentic-ai-pedagogical-best-practice-2026]] — Dynamic (fading) scaffolds as a guardrail for agentic AI - [[zhang-tutormoments-2026]] — When Help is Unhelpful: evaluating AI tutors for productive struggle - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[brcic-effortless-trap-productive-struggle-2026]] — Guarded vs. unguarded AI: the placement rule (Brcic & Frljic 2026) - [[stanford-evidence-base-ai-k12-2026]] — Tutoring-specific AI preserves productive struggle vs. general-purpose task completion - [[young-people-learning-generative-ai-rapid-review-2026]] — Guardrailed GenAI tools as scaffolds vs answer sources - [[ai-supported-experimental-design-chemistry-2026]] — AI-supported experimental design in practical chemistry - [[ai-video-dual-gatekeeping-2026]] — When Saying No Makes Better Videos: Dual Gatekeeping for Pedagogically Grounded AI Content Creation - [[scaffolding-systematic-reviews-2026]] — Scaffolding Systematic Reviews with Mentoring and AI (Wang 2026) - [[lukesova-clue-before-correction-2026]] — Clue Before Correction: ChatGPT for Autonomous Language Learning - [[wang-safety-gap-productive-struggle-2026]] — The Safety Gap: Restoring Productive Struggle - [[rhaimi-productivemath-2025]] — ProductiveMath: AI to Support PF Problem Design - [[computational-thinking-aica-2026]] — Computational Thinking Levels and AI Coding Assistants (2026) - [[ai-writes-code-student-writes-model-2026]] — Model authorship: theory & measurement for learning-by-construction with GenAI - [[code-to-learn-genai-artifact-construction-2026]] — CtL-GenAI: constructionism framework for artifact construction - [[pedagogy-ai-mistakes]] — The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking (Hosseini 2026) - [[llm-adaptive-programming-error-explanations-2026]] — LLM adaptive explanations of programming errors - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a mediational agent - [[tsingidou-ct-robotics-kindergarten-2026]] — Scaffolding is a dominant CT learning strategy - [[cogevol-learning-environment-generation-2026]] — CogEvol: Learning Environment Generation - [[preservice-teacher-agency-genai-design-learning-2026]] — Pre-service teacher agency during GenAI interactions in design for learning (Krushinskaia, Elen & Raes 2026) - [[bondurant-shaughnessy-ai-pedagogies-practice-2026]] — AI across the pedagogies of practice in mathematics teacher education: structured rehearsal feedback raised probing questions (Bondurant & Shaughnessy 2026) - [[ai-integration-instructional-design-collaboratory-2026]] — AI integration as instructional design: critique of AI output must be assigned, and disciplinary analysis comes first - [[making-ai-annoying-constrained-writing-2026]] — Making AI annoying on purpose: constraint in AI-supported writing (Konradt, Boote & Taub 2026) - [[rule-integrated-llm-tutoring-primary-math-2026]] — Rule-guided vs ad-hoc scaffolding in an LLM tutoring system for primary mathematics (Looi et al. 2026) - [[sidorkin-ai-generated-course-readings-2026]] — Comprehension prompts as a scaffold dial in AI-generated course readings (Sidorkin 2026) - [[ai-literacy-tool-design-programming-education-2026]] — A hint-budgeted AI Study Coach: scaffolded vs unrestricted GenAI use, and why the constraint alone did not produce learning (Azimi 2026) - [[scan-framework-task-assignment-generative-ai-2025]] — SCAN: scaffolding reframed as deciding which sub-zone a task belongs to - [[adaptive-scaffolding-contingency-comet-tutor-2026]] — Adaptive Scaffolding Needs Contingency: An AI Tutor That Escalates and Fades on What the Learner Does - [[pivot-generative-video-tutors-stem-2026]] — From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning - [[chen-automated-scoring-interpreting-self-regulated-learning-2026]] — Automated scoring as formative scaffolding: it converts into a gain only where the targeted deficit is decomposable, reliable and sensitively scaled (Chen & Liu 2026) - [[zhu-adaptive-teaching-assistance-genai-big-data-2026]] — Closed-loop adaptive scaffolding that tracks the learner's current boundary in music practice, with error detection as triage (Zhu et al. 2026) - [[niu-genai-children-creative-thinking-cognitive-development-review-2026]] — Scaffold form must match children's developmental capacity, not only the task (Niu et al. 2026) - [[vega-baudrit-genai-university-chemistry-education-review-2026]] — Scaffolds must require representational translation and verification, since students cannot verify what they do not understand (Vega-Baudrit & Rivera Álvarez 2026) --- ## [Socratic Method](https://edtechdev.github.io/aied/concepts/socratic-method/) > **Socratic Method** — a [[pedagogy|pedagogical]] approach rooted in guided questioning and dialogue rather than direct instruction, now being adapted for generative AI tutoring systems. In [[ai-education|AI in education]], the Socratic method is operationalized through LLMs that ask probing questions, scaffold reasoning, and withhold direct answers — aiming to promote deeper understanding and [[desirable-difficulties|productive struggle]] rather than answer-fetching.([[hashmi-socratic-physics-chatbot-2025]])([[favero-critical-ai-tutors-empower-enslave-2025]]) ## Questions to Consider - Think of a time a teacher (or friend) answered your question with another question and it actually helped you think. What made it work, and when did it instead just feel frustrating or evasive? - The Socratic approach withholds direct answers to provoke 'productive struggle.' Do you believe struggle is necessary for deep learning, or is it sometimes just unnecessary friction — and how would you tell the difference? - An AI Socratic tutor must decide when to guide, when to hint, and when to give a direct answer, based on a student's real-time signals. How do you think a system (or a human) knows which move to make at a given moment? - The page notes a frustrated student may need a brief direct answer before returning to Socratic questioning. What do you think this implies about the limits of a one-size-fits-all question-only approach? - If a [[conversational-ai|chatbot]] that only asks questions can produce measurable reasoning gains, what might be lost compared to the original Socratic dialogue with a human mentor — and what might be gained? ## Introduction The Socratic method is one of the oldest pedagogical techniques — originating with Socrates in ancient Athens — and it has found new relevance in the age of [[generative-ai|generative AI]]. In AI education [[research-methods-aied|research]], the Socratic method refers to AI systems that engage learners through guided dialogue, posing questions that lead students to discover answers rather than providing them outright. Asking structured questions rather than providing answers is one of the strongest pedagogical scaffolds for deep learning; when automated via AI, it produces measurable reasoning gains but also requires careful calibration to avoid frustrating learners or displacing human mentorship.([[hashmi-socratic-physics-chatbot-2025]])([[favero-critical-ai-tutors-empower-enslave-2025]]) ## How it works in AI tutoring Unlike direct-instruction AI tutors that give answers, Socratic AI tutors use question sequences that: - **Elicit [[prior-knowledge|prior knowledge]]** — asking what the student already knows about a topic - **Probe reasoning** — "Why do you think that?" or "What if the situation were different?" - **Surface [[misconceptions]]** — through carefully chosen counterexamples - **Guide toward insight** — without giving the answer away The Socratic approach directly embodies the principle from [[pedagogical-llm-training|EduQwen]]: **reward "guiding" over "answering."** However, real-time Socratic calibration is harder than paper-bench pedagogy: EduQwen optimizes for correct guiding on a multiple-choice [[benchmark]], whereas a live Socratic tutor must decide *when* to guide, *when* to hint, and *when* to answer — based on real-time student signals. [[affective-tutoring|Affective state]] is a critical moderator: a frustrated student may need a brief direct answer before returning to Socratic mode. ## Evidence of effectiveness A custom Socratic AI chatbot deployed in a large-enrollment introductory mechanics course (150 first-year [[stem-education|STEM]] majors) produced measurable reasoning gains: | Metric | Result | |---|---| | **Sample** | 150 first-year STEM majors | | **Knowledge-based skills rating** | Median **4.0/5** | | **Overall effectiveness rating** | Median **3.4/5** (notable gap) | | **Question specificity (first turn)** | ~10–15% | | **Question specificity (final turn)** | **100%** | | **Specificity × grade correlation** | Pearson **r = 0.43** | **Interpretation:** Students began with vague, generic questions but progressively sharpened them through Socratic interaction — a clear indicator of developing expert-like reasoning. The positive correlation between question specificity and self-reported expected grade suggests that learning to ask better questions is itself a domain skill. ### The effectiveness gap The gap between "knowledge-based skills" (4.0/5) and "overall effectiveness" (3.4/5) suggests a tension: students recognize that the Socratic bot improved their reasoning, yet do not fully endorse it as a complete tutoring solution. Possible reasons: - Socratic dialogue is effortful; students may prefer direct answers for efficiency - The chatbot cannot provide the relational support of a human tutor - Some students may get stuck in Socratic loops without resolution ### A counter-finding: unrestricted access can outperform constrained modes Not all evidence favors constraining the AI. [[socratic-nuclear-ai-learning|Socrates went Nuclear (Clin Deffarges, Kosmyna & Maes, 2026)]], a randomized EEG study of 50 participants comparing an unrestricted ChatGPT-style bot, a Socratic hint-only mode, and an adaptive question-limited mode on a nuclear-safety learning task, found that the **unrestricted chatbot produced higher learning gains** than both constrained modes (*p* < .03, *d* > 0.80) — even though the **adaptive condition generated significantly higher EEG-measured [[student-engagement|cognitive engagement]]** (*p* = .018). The result complicates the assumption that pedagogically constrained (Socratic) interaction always yields deeper learning: on short-horizon factual acquisition, free access won, while restricting access raised measured cognitive engagement without converting it into higher immediate post-test gains. This is a useful calibration point alongside the stronger [[learning-gains|learning-outcome]] results above: constraint can boost engagement, but the engagement-to-retention translation is not automatic, and over-constraining may simply frustrate learners seeking answers. In [[medical-education|clinical]]-interview training, [[ai-standardized-patient-scaffolding-medical-2026|the MeduAI-SP trial (Yang et al., 2026)]] had the tutor agent deliver Socratic prompts only on a flagged need — missing key history, premature closure, conversational impasse, or communication breakdown — phrasing them as reflective questions such as whether the gathered information sufficed to support the leading diagnosis. Students trained under this Socratic scaffolding scored 31 percentage points higher on the observable "expressing empathy" checklist item (Holm-corrected P = 8.30e-4) and 0.90 points higher on the 1–5 OSCE communication domain (P = 4.50e-4), linking non-answer-giving questioning to measurable patient-centered communication gains rather than to diagnostic accuracy (84% vs. 86%; P = 1.000). ## Research in the knowledge base The **[[hashmi-socratic-physics-chatbot-2025|Socratic Physics Chatbot]]** provides empirical evidence that the Socratic method can be operationalized through generative AI at scale, serving simultaneously as a [[teacher-role|teaching]] tool and data-collection instrument for [[learning-analytics]]. Unlike rule-based Socratic systems of the past, [[llm]]-based approaches can adapt question sequences dynamically based on student responses. **[[ai-agents-constructive-conflict-design-education-2026|Adversarial AI agents]]** enact constructive conflict — a Socratic variant — [[prompt-engineering|prompting]] novice designers to reconsider their assumptions, leading to more design iterations and higher-rated final work. This connects Socratic questioning to [[design-thinking]] and [[critical-thinking]]. **[[syal-multimodal-dialogue-stem-2026|Multimodal dialogue systems]]** extend Socratic tutoring to visual domains, using a zero-retraining intervention protocol that asks models to describe, reason, and self-correct — a [[multimodal]] Socratic scaffold. **[[retrieval-augmented-tutoring-algorithm-kite|Retrieval-augmented tutoring]]** operationalizes Socratic principles through retrieval, anchoring each response in authoritative course content rather than relying only on the model's parametric knowledge — addressing the gap that pedagogical quality alone is insufficient without content fidelity. [[lftutor-logical-fallacy-education-2026|LFTutor (Shi et al., 2026)]] applies Socratic questioning to a subject where withholding the answer is the whole task: teaching laypeople to see the logical fallacy in a persuasive text they believe is valid. Its dialogue agent decomposes the learner's own argument with the Toulmin model (claim, grounds, warrant), detects the learner's intent, and then selects exactly one of four strategies - Responding, Evidence, Assumption, Refutation - in a fixed priority order that mirrors the Toulmin structure, with a separate verifier agent checking after generation that the reply actually executed the chosen strategy and rephrasing it when it did not. The evaluation metrics are the Socratic failure modes rather than learning gains: divergence from the topic, stance change (caving to the learner's position), repetition, failure to refute, failure to ask for evidence, strategy fixation, unexplained fallacy terminology, and passive guidance. Across 1,000 simulated dialogues per framework with a GPT-4o backbone, LFTutor passed 84.5% of dialogues on average against 61.5% for a prompt that listed those same pitfalls and 31.2% for plain role-play prompting, and the ablation shows the gain is not from the Toulmin vocabulary but from verified strategy execution and intent-based selection. With 20 human participants debating the tutor, LFTutor scored significantly better on eight of nine Likert metrics, including helpfulness (4.15 against 1.65), with repetition the one dimension where the difference was not significant. ## Agency and critical use Favero et al. (2025) caution that even Socratic AI can undermine [[agency]] if students become dependent on the questioning structure rather than internalizing it. The goal is not permanent Socratic scaffolding but **scaffolded transfer** — students eventually Socratize themselves. ## Connections to other concepts The Socratic method is closely tied to [[scaffolding]] (providing just enough support), productive-struggle (letting students wrestle with difficulty), and [[intelligent-tutoring]] (adaptive question sequencing). It contrasts with [[cognitive-offloading|Over-Reliance]] — students who receive direct answers may bypass learning, while Socratic guidance maintains cognitive engagement. It supports [[self-regulated-learning]] and [[metacognition]] by making reasoning visible, and connects to [[formative-assessment]] when used to probe understanding in real time. ## Open Questions 1. Does Socratic dialogue transfer across domains, or is [[discipline-specific-aied|domain-specific]] reasoning non-transferable? 2. How does Socratic specificity correlate with *actual* (not self-reported) course performance? 3. Can Socratic AI be combined with [[becerra-aicofe-feedback-2026|peer feedback]] for social amplification? - **Withholding answers to provoke reasoning.** [[puech-pedagogical-steering-llm-productive-failure-2025|Puech et al. (2025)]] engineer LLM tutors to follow [[productive-failure|productive failure]] pedagogy by withholding solutions and eliciting multiple attempts — a Socratic-style refusal to give help except when strictly necessary; [[wang-safety-gap-productive-struggle-2026|Wang & Shan (2026)]] recommend Socratic and Adversarial AI architectures that preserve constructive cognitive friction. ## Connected Concepts - [[scaffolding]] - [[intelligent-tutoring]] - [[learning-analytics]] - [[stem-education]] - [[student-modeling]] - [[student-experience]] - [[agentic-ai]] - [[metacognition]] - [[knowledge-tracing]] - [[adaptive-learning]] - [[generative-ai]] - [[cognitive-offloading]] - [[self-regulated-learning]] - [[formative-assessment]] - [[ai-literacy]] - [[agency]] - [[critical-thinking]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[productive-failure]] — Productive Failure ## Connected Articles - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[hashmi-socratic-physics-chatbot-2025]] - [[physics-chatbot-epistemological-beliefs-2026]] - [[ai-agents-constructive-conflict-design-education-2026]] - [[syal-multimodal-dialogue-stem-2026]] - [[retrieval-augmented-tutoring-algorithm-kite]] - [[genai-performance-vs-learning]] - [[structured-llm-feedback-programming]] - [[zerkouk-comprehensive-review-its-2025]] - [[embodied-inquiry-ai-facilitator-physics-2026]] - [[prober-ai-inquiry-writing]] - [[critical-thinking-genai-scaffolding]] - [[generative-ai-guardrails-harm-learning]] - [[pedagogy-ai-mistakes]] - [[stanford-evidence-base-ai-k12-2026]] — Structured Socratic hints vs. open-ended general-purpose Q&A - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[wang-safety-gap-productive-struggle-2026]] — The Safety Gap: Restoring Productive Struggle - [[rhaimi-productivemath-2025]] — ProductiveMath: AI to Support PF Problem Design - [[lukesova-clue-before-correction-2026]] — Clue Before Correction: ChatGPT for Autonomous Language Learning - [[socratic-nuclear-ai-learning]] — Socrates went Nuclear: Comparing Interaction Strategies for AI in Learning - [[lftutor-logical-fallacy-education-2026]] — Socratic questioning plus critical argumentation in a four-step fallacy-tutoring framework --- ## [Critical Pedagogy](https://edtechdev.github.io/aied/concepts/critical-pedagogy/) > **Critical Pedagogy** — an educational approach, rooted in the work of Paulo Freire and later critical theorists such as Henry Giroux, that treats teaching and learning as inherently political acts. Rather than merely transmitting skills or knowledge, critical pedagogy asks who benefits from education, whose knowledge is privileged, and how schooling reproduces or resists systems of power and oppression. In AI-in-education, critical pedagogy interrogates the corporate and capitalist logics shaping AI adoption, centers the voices and epistemologies of marginalized communities, and treats [[ai-literacy|AI literacy]] as a practice of resistance and social transformation rather than mere technical competence. ## Questions to Consider - Who does your education — or the AI tools in it — actually serve? Critical pedagogy insists this question is unavoidable, not optional. What's your honest answer? - Critical thinking asks 'is this reasoning sound?' Critical pedagogy asks 'who does this education benefit, and whose knowledge counts?' How are those two questions different, and when does one need the other? - Some scholars describe generative AI's spread into education as a form of 'colonization' that extracts data, labor, and resources from marginalized communities. Does that framing feel extreme, or does it name something real? - Critical AI literacy can include 'resisting AI' — refusing the inevitability of tech as a solution. In what situations might strategic refusal or non-use be a more responsible choice than adoption? - Who gets to decide what counts as authoritative knowledge? If [[ai-technologies|AI systems]] are positioned as authoritative, what happens to learners' own lived and community epistemologies? - Under critical pedagogy, a teacher is not a neutral transmitter of AI skills but a facilitator who helps [[learners]] interrogate the politics of AI. How comfortable are you with that role, and what would it ask of you? ## Introduction Critical pedagogy is distinct from [[critical-thinking]]. Critical thinking is a cognitive skill — evaluating arguments, questioning assumptions, and reasoning carefully — that can be practiced within almost any framework. Critical pedagogy, by contrast, is a *political and [[ethics|ethical]] stance* that connects those skills to questions of power, justice, [[equity-in-ai-education|equity]], and social transformation. The knowledge base treats them as related but distinct concepts: critical thinking asks "is this reasoning sound?", while critical pedagogy asks "who does this education serve, and whose knowledge counts?". ### Critical pedagogy and AI in education [[ai-education|AI in education]] raises sharp questions for critical pedagogy, and the knowledge base documents several strands: - **Critique of the "AI colonization" of education.** Critical and feminist scholars argue that [[generative-ai|generative AI]] is reshaping education in ways that repeat patterns of colonial exploitation and extraction — of data, labor, and natural resources — enriching the powerful at the expense of marginalized communities.([[avraamidou-ai-colonization-science-education]]) This framing rejects the "techno-utopia" narrative that presents AI as a neutral silver bullet. - **Feminist and justice-centered AI.** A feminist approach to AI prioritizes justice over profit, asks critical questions about the nature and ownership of knowledge, who benefits from AI, and who is accountable when systems fail. It calls for critical AI literacy framed within feminist [[pedagogy|pedagogies]].([[avraamidou-ai-colonization-science-education]]) - **Resisting AI as a critical literacy practice.** Critical AI Literacy (CAIL) can encompass "resisting AI" — a stance that refuses the inevitability and techsolutionism of dominant discourse, and instead cultivates collective [[agency]] through dialogic, [[collaborative-learning|collaborative]] pedagogies that imagine alternative futures.([[li-mroziak-reorienting-critical-ai-literacy]]) - **[[multimodal]] [[writing-education|composition]] as critical AI literacy pedagogy.** [[burriss-multimodal-composition-critical-ai-literacy-2026|Burriss et al. (2026)]] document a classroom unit in which 22 eleventh-grade students composed video public service announcements about AI [[ethics]] issues of their choosing, translating surveillance, [[privacy|data consent]] and algorithmic accusation into 90-second to 3-minute films built from sound, image, text and their own bodies. Framed by Critical Posthumanist Literacy, the work shows the political stance in practice: across all seven films harm emerged from tangled human–machine responsibility rather than from a villainous tool, and students still closed on [[agency]] — "the solution is in our reach" — even as one participant dissented that youth "completely lack any kind of credibility." The authors argue that creative, [[collaborative-learning|collaborative]] composition is itself a form of critical AI literacy, and that existing AI literacy scales and competency frameworks, which assume individually measurable performance, systematically exclude it. - **Redistributing epistemic authority.** Community-based AI learning grounds AI [[student-engagement|engagement]] in learners' lived and community epistemologies, challenging the positioning of AI systems as authoritative knowledge sources.([[ojeda-ramirez-community-based-ai-learning]]) This involves epistemic fine-tuning, redistribution of authority, and [[situated-learning|situated]] discernment. - **Digital mediation as critical practice under resource constraint.** [[beyond-the-algorithm-academic-developers-digital-mediators-2026|Sithole (2026)]] applies Critical Digital Pedagogy with decolonial epistemologies of the South to [[educational-development|academic development]] in South African Historically Disadvantaged Institutions, where developers filter institutional AI rhetoric through the question of what it means "for our students from poor schools." Their mediation is intervention rather than facilitation — interrogating rather than adopting tools, asking what is gained and whose labor disappears, and treating context as an epistemic principle rather than a design preference. The author's analytical contribution is to separate [[digital-divide|digital inequality]] from **algorithmic coloniality** and to insist, following Adams, that calls to decolonize AI stay material rather than metaphorical. - **Situated and cultural-historical ethics.** Critical approaches also argue that AI ethics in education must be *situated* — grounded in cultural-historical and ecological context rather than abstract principles.([[raffaghelli-situated-ai-ethics-2026]]) - **Efficiency-first alignment discourse as an erasure of context.** [[mcinnes-salvaging-constructive-alignment-genai-2026|McInnes et al. (2026)]] apply Fairclough's three-dimensional model to 14 pieces of gray literature (November 2022 – April 2025) advising higher-education practitioners to use [[generative-ai|generative AI]] for constructive alignment, and find a techno-solutionist discourse in which the tool is anthropomorphized as "an educational expert and assistant" and academic staff are positioned as supplying "subject matter expertise" while the system performs the pedagogical work. The analysis names the erasure of situated, disciplinary and critical context as one of three failure modes — alongside performativity (alignment that only looks aligned) and shallow alignment that conflates the constructive dimension with the aligned one — evidence that even advice about [[learning-design|course design]], ostensibly a neutral technical matter, carries the depoliticising logic critical pedagogues critique elsewhere. - **Human rights education as the test case for AI governance as formation.** [[kasa-malksoo-ai-human-rights-education-2026|Kasa-Mälksoo (2026)]] works through the UN's *about, through, for* framework in a law program and locates the difficulty not in the technology but in the pedagogy: students submitted polished written work and polished session designs without the engagement those artifacts are supposed to evidence, and the most critical thinking appeared in end-of-course writing rather than in class dialogue. Her conclusion shifts the educator's role rather than shrinking it — [[governance|AI governance]] becomes a professional responsibility students are trained to contest, and the [[llm|tool]] is directed rather than banned. ### The role of the educator Under critical pedagogy, educators are not neutral transmitters of AI skills but critical interlocutors and facilitators who help learners interrogate the politics of AI. This connects to the knowledge base's [[teacher-role]] and [[ai-literacy]] concepts, and to the broader concern with [[equity-in-ai-education]] and [[reducing-ai-misuse]]. The educator's task is to cultivate spaces where communities can collectively question, appropriate, or refuse AI — keeping education a site of imagination and social transformation. [[burriss-multimodal-composition-critical-ai-literacy-2026|Burriss et al. (2026)]] add a practical condition to that role: open-ended critical AI literacy work does not require educators to arrive as AI experts, only to be willing to explore alongside students — a joint student–teacher investigation of one specific system, electronic "hall passes" (not all of which are AI), surfaced and worked through [[misconceptions]] while deepening technical knowledge. [[miles-prompt-literacy-human-centered-genai-framework-2026|Miles, Haber-Curran and Arar (2026)]] give that role a constructionist, ethics-of-care shape: they treat [[prompt-engineering|prompt literacy]] as a rhetorical, ethical and reflective process rather than a technical optimization skill, and the framework's closing phase has students co-design a Personal AI Use Policy with their instructor, so the norms governing classroom AI use are authored by the people they govern rather than issued to them. ## Connected Concepts - [[critical-thinking]] - [[ai-literacy]] - [[equity-in-ai-education]] - [[culturally-relevant-pedagogy]] - [[reducing-ai-misuse]] - [[agency]] - [[ethics]] - [[teacher-role]] - [[social-emotional-learning]] - [[ai-education]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education ## Connected Articles - [[kasa-malksoo-ai-human-rights-education-2026]] — Human rights education, reflective practice, and teaching AI governance as professional formation (Kasa-Mälksoo 2026) - [[burriss-multimodal-composition-critical-ai-literacy-2026]] — Multimodal composition as critical AI literacy pedagogy - [[mcinnes-salvaging-constructive-alignment-genai-2026]] — Critical discourse analysis of techno-solutionist GenAI constructive-alignment advice - [[benali-genai-academic-writing-2026]] - [[avraamidou-ai-colonization-science-education]] — Critical, feminist critique of the "AI colonization" of science education - [[li-mroziak-reorienting-critical-ai-literacy]] — "Resisting AI" as a community-rooted praxis of critical AI literacy - [[ojeda-ramirez-community-based-ai-learning]] — Redistributing AI's epistemic authority through community-based learning - [[raffaghelli-situated-ai-ethics-2026]] — Situated, cultural-historical and ecological framework for AI ethics - [[voicu-ai-interpretive-cognition-ssh-2026]] — Developmental-critical model for interpretive cognition in the humanities - [[mechanical-compliance-human-flourishing-ai-literacy-2026]] — Socialist humanist AI literacy + fair use - [[alsuhaymi-sustainable-education-ai-digitalization-2026]] — Value-critical approach to sustainable education and AI (Alsuhami & Atallah 2026) - [[emancipatory-ai-learner-flourishing-2026]] — Emancipatory vision for designing generative AI toward learner flourishing - [[miles-prompt-literacy-human-centered-genai-framework-2026]] — Human-centered GenAI engagement framework built on Freire, constructionism and an ethics of care - [[beyond-the-algorithm-academic-developers-digital-mediators-2026]] — Critical Digital Pedagogy and decolonial epistemologies applied to academic development in the Global South --- ## [Pedagogical Partnerships](https://edtechdev.github.io/aied/concepts/pedagogical-partnerships/) > **[[pedagogy|Pedagogical]] partnerships** (also known as **students as partners** or **SaP**) — a relationship-centered approach to teaching and learning in which students and educators work together as collaborators to co-create curriculum, teaching activities, [[assessment]], open educational resources, and educational policies, rather than treating students as passive recipients of faculty-designed instruction. Rooted in values of **respect, reciprocity, and responsibility** (Cook-Sather et al., 2014), pedagogical partnership repositions students as co-creators of knowledge and educational innovation and challenges traditional top-down power dynamics in education. In the AI era, partnership has become a key counter-narrative to deficit views of students as AI cheaters or victims — positioning students as partners in co-designing AI policy, tools, and practice. ## Questions to Consider - When you design a course, assignment, or policy, who typically makes the decisions — and whose perspective is left out? What would change if the people most affected by those decisions were genuine collaborators rather than recipients? - A common assumption is that students lack the expertise to contribute to teaching design. Yet partnership [[research-methods-aied|research]] argues students are experts in their own experience as learners. Where is the line between a student's legitimate experiential expertise and the faculty member's disciplinary expertise? - The word "partnership" can paper over real power differences — a student asked to "co-design" a course still depends on the instructor for grades and references. What conditions make a partnership genuine rather than performative? - In the age of generative AI, students are often framed as either cheaters or victims. How might reframing them as partners in designing AI use and policy change both the tools built and the conversations educators have with students? - Partnership can extend to whole classes, small groups, or individuals. Which form is most equitable, and what practical and resource constraints limit broader participation? - What does trust have to do with pedagogical partnership — and how might surveillance- or deficit-oriented responses to AI fracture the relationships partnership depends on? ## Introduction Pedagogical partnerships, also called **students as partners (SaP)**, describe a spectrum of practices in which students collaborate with academic staff to shape teaching, learning, assessment, curriculum, and the broader [[student-experience|student experience]]. The approach draws on the foundational guide by Cook-Sather, Bovill, and Felten (2014) and the higher-education framework of Healey, Flint, and Harrington (2014), which describe partnership as grounded in three core values: **respect, reciprocity, and responsibility**. Partnership is not about making students pedagogical experts, nor about faculty relinquishing disciplinary authority; rather, it recognizes the complementary roles each partner brings — faculty as disciplinary experts and intellectual guides, students as experts in their own experience as learners with diverse talents and interests. Pedagogical partnership challenges the entrenched assumption that students cannot be trusted to collaborate on teaching and learning, and it actively disrupts traditional faculty-centered hierarchies. It overlaps conceptually with [[collaborative-learning]], [[agency|learner agency]], and [[student-engagement]], but is distinct in centering students in the *design* of education itself — curriculum, assessment, teaching activities, open educational resources, and policy — rather than merely as active participants within a pre-designed course. In the [[generative-ai|generative AI]] era, pedagogical partnership has taken on new urgency. As AI amplifies deficit narratives that frame students as either cheaters or victims (see [[ai-misuse-learning-harm]] and [[framing-ai-use-for-students]]), partnership offers a counter-narrative: rather than policing students or assuming misuse, educators and institutions can engage students as co-designers of AI policy, tools, and practice. This is reflected across the knowledge base's research — from students co-designing course AI policies to co-creating custom AI tools and co-constructing assessment criteria. ## Forms and scope of partnership Partnership can take many forms and operate at multiple scales, each with distinct affordances and constraints: - **Whole-class partnership (co-creation in the classroom):** All students in a course collaborate with faculty on course structure, content, goals, and policies. Whole-class approaches may be the most equitable form of partnership because they allow broad participation and benefit, though they require substantial "high-touch" personnel resources to support increased student decision-making (Bovill, 2020). - **Individual or small-group partnership:** A small number of student partners work with staff on a specific project — for example, co-designing [[assessment]] criteria, evaluating AI tools, or developing curriculum. These are well-suited to dialogic, in-depth collaboration. - **Project- or program-based partnership:** Students partner on a bounded initiative, such as co-designing a course AI policy, building a custom AI tool, or developing an algorithmic-literacy program. - **Partnership in [[curriculum-design]] and [[learning-design]]:** Students contribute to shaping what and how learning happens, from individual activities to entire programs. Across these forms, partnership is a *process and a relationship*, not just an activity. The outcomes are often co-created and not fully defined in advance, which requires staff to let go of control and accept the risks of the unknown — values that run counter to the hidden rules of much of [[higher-ed|higher education]]. ## Pedagogical partnership and AI in education The intersection of pedagogical partnership and AI is one of the fastest-growing strands of the knowledge base. Key themes: - **Students as co-designers of AI policy.** Rather than institutions imposing AI rules on students, partnership positions students as co-creators of course- and institution-level AI policy. For example, a guided-inquiry activity in which students co-designed a [[generative-ai]] course policy surfaced student priorities around training, standardized disclosure, institutional support, and involvement in decision-making ([[guided-inquiry-genai-course-policy-2026]]). - **Students as co-creators of AI tools.** In students-as-partners frameworks, students have co-designed and refined custom AI [[conversational-ai|chatbots]] aligned with pedagogical goals — an approach that extends the SaP paradigm to include AI tools themselves and positions student voice as central to responsible AI innovation ([[lo-co-creating-custom-gpts-sap-2026]]). - **Co-creation in assessment and AI.** Student-staff partnerships have co-evaluated AI-generated output in coursework assessments and co-designed evaluation criteria, improving understanding of AI's benefits and limits and supporting [[self-regulated-learning]] ([[williams-ingle-assessment-co-creation-ai-2025]]). - **Partnership to preserve pedagogical trust.** In the face of AI-driven uncertainty and "AI shame," partnership practices that nurture *pedagogical trust* — a confident, reciprocal learning relationship open to uncertainty and co-navigated through dialogue — offer a way to sustain the relational core of education ([[matthews-five-guiding-principles-ai-sap-trust-2025]]). - **Youth as co-designers of AI systems.** Participatory design approaches engage historically minoritized students as partners in designing the AI systems that will affect their classrooms, surfacing students' values and [[ethics|ethical]] commitments ([[chang-co-designing-ai-youth-relational-privacy-2025]]). - **Students' lived realities as a design input.** Frameworks like entangled pedagogy illuminate the complex, "messy" realities students navigate with AI, arguing that policy and guidance should be co-designed with students in light of their lived experiences rather than imposed from above ([[fawns-entangled-pedagogy-genai-students-2026]]). These strands share a core claim: that the people who will live with [[ai-education|AI in education]] — students — must be meaningfully engaged in shaping it, or the resulting tools, policies, and practices risk being untrustworthy, inequitable, or ineffective. ## Relationships to related concepts Pedagogical partnership is related to but distinct from several neighboring concepts: - **[[agency|Learner agency]] and [[self-directed-learning]]:** Partnership supports learner agency by giving students a voice in what and how they learn, and it overlaps with self-directed learning. However, partnership is fundamentally *relational* — agency and direction emerge through the instructor–student relationship rather than in isolation. Some frameworks conceptualize this as "shared agency" that exists in the relationship between student and instructor, not within either party alone. - **[[collaborative-learning]]:** Both involve working together, but collaborative learning typically refers to students learning *with each other* within a designed activity, whereas pedagogical partnership centers students and staff collaborating on the *design* of education itself. - **[[student-engagement]]:** Partnership is a deep form of engagement, but engagement describes a student's involvement in learning while partnership describes a changed power relationship in which students share responsibility for shaping that learning. - **[[teacher-role]]:** Partnership redefines the teacher's role from sole authority and gatekeeper of knowledge to co-designer and intellectual guide, without diminishing faculty expertise. - **[[assessment]] and [[authentic-assessment]]:** Co-designing assessment — from evaluation criteria to rubrics to whole assessment formats — is a central arena of partnership, and AI has amplified both the need for and the difficulty of genuine co-design in assessment. - **[[educational-policy-ai]] and [[governance]]:** Partnership extends into policy and governance, positioning students as co-creators of AI policy rather than passive subjects of it. - **[[equity-in-ai-education]] and [[inclusive-learning]]:** Partnership is argued to be a more equitable form of educational participation, centering voices — especially those of historically minoritized or marginalized students — that are often excluded from educational decision-making. ## Key research themes - Whether and how partnership improves [[learning-gains|learning outcomes]], [[student-engagement]], [[motivation]], and a sense of belonging. - The relational conditions — trust, psychological safety, dialogue, shared responsibility — required to bring partnerships about and sustain them. - How AI both enables and complicates partnership: as students co-design AI tools and policies, and as deficit narratives around AI threaten the trust partnership depends on. - Equity in partnership: which students get to participate, whose voices are centered, and how whole-class vs. selective forms distribute benefit. - The resource and institutional constraints on partnership, including the time-intensive nature of genuine co-design. - How partnership relates to decolonizing, critical, and [[critical-pedagogy|critical pedagogies]] that challenge traditional hierarchies and center multiple ways of knowing. ## Practical implications For educators and institutions seeking to adopt pedagogical partnership, the knowledge base's research suggests: 1. **Start small and iterate.** Begin with a bounded project — a pilot with a limited number of participants — and honor participant effort with recognition (employment, credit, certificates, publication). 2. **Build on existing infrastructure.** [[educational-development|Faculty development]] programs, student employment relationships, and cross-campus partnerships provide foundations for adapting partnership practice. 3. **Lead with relationship and trust.** Partnership works when trust, psychological safety, dialogue, and shared responsibility are deliberately cultivated — not assumed. 4. **In the AI era, invite students in rather than policing them.** Co-designing AI policy, tools, and assessment with students is both more equitable and more effective than deficit-based responses. 5. **Consider equity of participation.** Whole-class partnership offers the broadest participation; selective programs must attend to which students are invited and included. ## Connected Concepts - [[pedagogy]] — Pedagogies and teaching strategies - [[collaborative-learning]] - [[agency]] — Learner agency - [[self-directed-learning]] - [[curriculum-design]] - [[learning-design]] - [[assessment]] - [[authentic-assessment]] - [[teacher-role]] - [[student-engagement]] - [[student-experience]] - [[educational-policy-ai]] - [[governance]] - [[equity-in-ai-education]] - [[inclusive-learning]] - [[human-ai-collaboration]] - [[student-ai-interaction]] - [[generative-ai]] ## Connected Articles - [[guided-inquiry-genai-course-policy-2026]] — A Guided Inquiry Approach to Students Co-Designing Generative AI Course Policies - [[lo-co-creating-custom-gpts-sap-2026]] — Co-creating custom GPTs: an autoethnographic study of undergraduate students as partners in generative AI innovation - [[williams-ingle-assessment-co-creation-ai-2025]] — Assessment design through co-creation: Student-staff partnership in evaluating AI - [[matthews-five-guiding-principles-ai-sap-trust-2025]] — Five guiding principles for navigating AI in students as partners practice to preserve pedagogical trust - [[chang-co-designing-ai-youth-relational-privacy-2025]] — Co-designing AI with youth partners: a relational privacy ethical framework - [[fawns-entangled-pedagogy-genai-students-2026]] — Illuminating complex student realities of AI through an entangled pedagogy framework - [[anastasia-shared-agency-partnership-framework-2026]] — Shared Agency: The Agency Partnership Framework for Instructor–Student Collaboration - [[maybee-disruptive-partnerships-sap-2025]] — Disruptive Partnerships: Collaborating with Students in Information Studies - [[student-centered-genai-responsible-framework-2026]] — Student-centered framework for responsible generative AI use - [[physics-faculty-learning-community-ai-2026]] — A Workshop Series for Effective Use of AI in Uncertain Times: Building a Physics Faculty Learning Community --- ## [Storytelling in Education](https://edtechdev.github.io/aied/concepts/storytelling-in-education/) > **Storytelling in education** — the use of narrative as a [[pedagogy|pedagogical]] tool to engage learners, convey meaning, and support knowledge construction, creativity, and emotional connection. Storytelling is a natural and motivating way for learners to make sense of the world, and it is increasingly combined with technology — including AI and [[educational-robotics|social robots]] — to create interactive, adaptive narrative experiences. Digital and robot-mediated storytelling can add interactivity, [[personalized-learning|personalization]], and embodiment that conventional (paper-based or slide-based) storytelling lacks. ## Questions to Consider - Think of a lesson or concept you remember vividly because it came as a story. What did the narrative add that a plain explanation didn't, and can that power survive being outsourced to an AI or a robot? - The page describes robot- and LLM-mediated storytelling that responds to the learner and supports co-creation. What do you imagine is gained — and potentially lost — when a story is co-built with a machine instead of told by a person? - If storytelling builds engagement, creativity, and emotional connection, where might its use in education be more than entertainment — and where might it risk substituting feeling for understanding? - A robot storyteller increased behavioral and cognitive engagement over paper and PowerPoint in the [[research-methods-aied|research]] cited. Do you trust that engagement translates into learning, or could the novelty of the robot be doing the work? - How could you tell whether a learner is genuinely making meaning from an AI-co-created story versus just enjoying the interactive experience? ## Introduction Storytelling is grounded in [[motivation]], [[student-engagement]], and [[constructivist]] theories of learning. It supports [[language-learning]], [[creativity]], [[social-emotional-learning]], and comprehension. In the AI era, [[llm|LLM-powered]] and robot-mediated storytelling enables co-creation, where learners and [[agentic-ai|AI agents]] build stories together, and interactive narrative that responds to the learner. ### How storytelling appears in the knowledge base's research - **Robot-mediated storytelling:** [[motibo-digital-storytelling-robots-motivation-2026|MotiBo]] uses a human-like interactive digital storytelling robot, finding significant gains in behavioral and cognitive engagement over paper and PowerPoint methods; [[robobuddy-llm-social-robots-classroom-2025|RoboBuddy]] lets teachers create LLM-powered scenario-based storytelling activities from [[curriculum-design|curriculum]] content. - **Co-creative narrative HRI:** [[icub-humanoid-storytelling-llm-hri-2025|The iCub narrative study]] explores human-robot co-creation of stories, integrating generative models for contextually appropriate interaction. - **Narrative and creativity:** Storytelling supports [[creativity]] and [[language-learning|language]] development, and is used to enhance motivation and engagement in [[k-12]] settings. - **Digital storytelling as a counterweight to AI:** As [[generative-ai|generative AI]] ascends, digital storytelling is reframed as a vehicle for the emotional, cultural, and narrative capacities that AI lacks. In a [[project-based-learning|project-based learning]] model for art and design education, students translated local cultural heritage into [[multimodal]] narratives across 92 digital storytelling works, demonstrating how storytelling sustains human creativity in the AI era. Storytelling connects to [[student-engagement]], [[motivation]], [[creativity]], [[educational-robotics]], [[language-learning]], [[social-emotional-learning]], and [[educational-robotics]]. ## Connected Concepts - [[student-engagement]] - [[motivation]] - [[creativity]] - [[educational-robotics]] - [[language-learning]] - [[social-emotional-learning]] ## Connected Articles - [[motibo-digital-storytelling-robots-motivation-2026]] — MotiBo - [[robobuddy-llm-social-robots-classroom-2025]] — RoboBuddy - [[icub-humanoid-storytelling-llm-hri-2025]] — iCub Narrative HRI - [[remind-robot-mediated-roleplay-antibullying-2026]] — REMind - [[white-wu-robotics-ai-education-2026]] — Robotics and AI in Education - [[project-based-digital-storytelling-art-design-2026]] — Project-based digital storytelling framework for art/design education in the AI era - [[adapted-stories-social-story-intervention-2026]] — AI-Assisted Social Story Intervention for Special Education: The Design of AdaptED Stories --- ## [Online Teaching and Learning](https://edtechdev.github.io/aied/concepts/online-teaching-and-learning/) > **Online teaching and learning** — the pedagogy and practice of teaching and learning that happens through digital, network-mediated environments rather than in a shared physical classroom. It spans fully online courses, Massive Open Online Courses (MOOC), blended and hybrid formats, and distance education. For the knowledge base, the central question is how [[generative-ai]] reshapes the opportunities, challenges, and recommended practices of teaching at a distance — from scalable [[personalized-learning|personalization]] to new [[academic-integrity]] and [[cognitive-offloading]] risks. ## Questions to Consider - You've likely taken an online course or taught one. What did you lose and what did you gain when the physical classroom was removed—and how did that reshape what instructors could rely on? - The page argues online teaching is a distinct pedagogy, not just a delivery mechanism. In what concrete ways does the online medium change which teaching strategies are even possible or effective? - An [[rct]] found unguarded AI assistance raised practice performance but lowered unassisted exam scores, while a 'hint-not-answer' tutor removed the harm. Before reading further, can you explain why giving students the answer might inflate immediate performance yet erode durable learning? - Online assessment can't always tell assisted from independent work. If detection tools are a 'partial, contested response,' what alternative assessment designs might reveal genuine understanding instead? - How does the perceived availability of an effortless AI shortcut reshape student motivation in a self-paced, screen-based course? What might you design to counter it? - An autonomous agent can now log into a learning management system, read the material, answer the quiz and post to the discussion. If producing the artifact no longer demonstrates learning, what would you need to see instead? - AI can now generate a MOOC-equivalent course in minutes at a fraction of the cost. What are the pedagogical trade-offs of 'N agents for one student' versus 'one video for N students'? ## Introduction Online teaching and learning is a distinct [[pedagogy|pedagogical]] context, not merely a delivery mechanism. It removes the physical co-presence that scaffolds attention, [[motivation]], and informal interaction, and it substitutes structured digital interaction — discussion forums, asynchronous materials, video, [[intelligent-tutoring|tutoring agents]] — for face-to-face contact. This changes what instructors can rely on, what students can access, and how learning is designed and assessed. As an umbrella concept in the [[pedagogy]] landscape, it sits alongside [[active-learning]], [[collaborative-learning]], and [[self-regulated-learning]] but is distinguished by the medium: the constraints and affordances of the online environment shape which strategies are viable. The rise of generative AI lands directly in this context. Online learners already work through screens and software, so AI tools are natural neighbors; at the same time, online assessment is harder to invigilate, making misuse easier and the stakes higher. The evidence in this knowledge base shows that AI can be a powerful ally for online teaching and learning — and, configured poorly, a significant source of learning harm. [[lock-integrating-ai-online-learning-higher-ed-2025|Lock, Arteaga & Johnson (2025)]]'s critical literature review (63 citations across 32 countries) organizes this landscape into four interconnected themes that recur throughout the page below: the types and purposes of AI integration, pedagogical approaches (AI literacy, self-regulated learning), benefits, and challenges. Their central caution — that these themes *overlap* and that online AI integration is a sociotechnical undertaking anchored in pedagogy and human relationships rather than technology adoption alone — aligns with the page's framing of online teaching as a distinct pedagogy. Notably, they report that students using ChatGPT *alongside* teacher tutoring perceived greater [[learning-gains]] than those using it alone, reinforcing the hybrid human–AI collaboration emphasis threaded through this page. The development that moved this from a design question to an urgent one is that generative AI no longer only writes text a student could have written. Autonomous [[agentic-ai|agents]] now log into a learning management system, read the course materials, answer the quizzes, read classmates' posts, and submit the work. Three demonstrations on a live undergraduate psychology course show what that means in practice: two quiz completions, one in roughly **12 minutes** and one in **under 5 minutes**, both scoring **10/10**, and a discussion post in which the agent mined its peers' posts and then fabricated a credible first-person life history to answer them — against a public record of at least **15 documented agent runs** across Canvas, Moodle and Brightspace using **seven agent tools**. [[ai-agents-complete-lms-assessment-validity-2026|Hadjisolomou and El-Haddad (2026)]] argue this is an [[assessment-validity]] problem before it is an integrity one: what agent completion removes is the assumption that the submitted work was produced by the person whose learning is being assessed. The design question for online teaching therefore shifts from how to detect misuse to what evidence of learning an online course can still produce, which is the thread running through the sections below. ## Formats and settings Online teaching and learning takes several related forms that share the medium but differ in reach and structure: - **Blended and hybrid learning.** Models that combine in-person and online components, intentionally integrating digital activities, materials, and interactions with face-to-face teaching. Blended formats ask instructors to decide what is best done synchronously vs. asynchronously and online vs. in person — decisions that [[learning-design]] principles organize and that AI both supports and complicates. In the blended context, AI tools offer opportunities for [[personalized-learning|personalization]] and always-on support while raising integrity and offloading risks that span both the online and in-person portions. [[mejia-domenzain-ml-findings-teachers-blended-2026|Mejia-Domenzain et al. (2026)]] illustrate this in two blended settings — flipped university classrooms and reflective writing in vocational education — where a learning analytics dashboard (DashED) communicated ML-derived [[self-regulated-learning]] profiles to teachers. Adoption concerns diverged by context: flipped-classroom (university) teachers worried most about data anonymization and student opt-out, whereas reflective-writing (vocational) teachers feared misuse of the tool by fellow educators and stressed the need to contextualize data. In use, flipped-classroom teachers followed a sequential exploration and favored course-level adaptation and showing dashboards in class, while vocational teachers revisited summary pages and used the tool mainly for individual coaching sessions — evidence that blended analytics design must be context-aware. - **Synchronous versus asynchronous design.** The distinction matters more than the delivery medium, because the two formats make opposite demands on the learner. Synchronous sessions carry attention, pacing and accountability inside the session itself; asynchronous courses have to design them in, since the learner alone decides when to work and receives no ambient accountability from a room. The research base often blurs this: the knowledge base's review of AI and [[student-engagement|engagement]] in online learning (24 studies) treats engagement alone and explicitly conflates synchronous with asynchronous contexts, so its conclusions should not be read as asynchronous-specific. What is asynchronous-specific is the evidence on how self-paced learners lose focus and pace — self-regulated behaviors such as goal setting, environment structuring and time management coincide most with low digital distraction ([[decreasing-digital-distraction-college-online-learning-2026|Shi et al. 2026]], 530 students), while metacognitive knowledge and [[well-being]] decline across a term in step with clustered assessment deadlines ([[song-genai-learning-partner-srl-over-time-2026|Song et al. 2026]], 75 students). The practical counterpart is the FAQ on [[asynchronous-online-courses-ai|designing and facilitating asynchronous courses when AI can do the work]]. - **Distance education.** Programs designed for learners who study remotely, often at scale and across regions (e.g., the Open University's 200K+ learners). Distance learning is where 24/7, context-embedded AI support and the impossibility of in-person invigilation are most salient. Comparative evidence from South African teacher preparation shows the medium itself is associated with preparedness: [[ai-training-science-teacher-tpack-distance-2026|Mnguni et al. (2026)]] found self-reported TPACK for AI-integrated science teaching higher among final-year student teachers at a campus-based university (64.0%) than at a distance education university (47.4%), with the weakest reported domain in both settings being Pedagogical Knowledge. The pattern warns that distance programs cannot assume that the same AI training produces the same readiness, and that the design of the training, not its presence, is what differs. ## Opportunities and benefits of AI for online teaching and learning - **Scalable personalization.** Traditional MOOCs excel at reach but struggle to adapt — "one video for N students." [[llm]]-driven agent systems ([[mooc-to-maic|MAIC]]) invert this to "N agents for 1 student," using specialized Teacher, Assistant, Classmate, and Analyzer agents to deliver [[adaptive-learning|adaptive instruction]], personalized feedback, and dynamic learning paths at MOOC scale. Systems like [[learnmate2-llm-adaptive-learning|LearnMate²]] address the "personalization gap" in open online learning with personalized study plans, real-time contextual assistance, and [[adaptive-learning|adaptive]] activities. Personalized video is a concrete route to this goal: [[personalized-ai-generated-videos-preference-2026|Tomlinson et al. (2026)]] found students in a large online course preferred AI-generated personalized videos over non-personalized human-recorded ones — a preference whose personalization effect outweighed the value placed on a human presenter — suggesting scalable, [[generative-ai]]-produced personalized media can close the "one video for N students" gap in online instruction. - **Always-on, context-embedded support.** In distance and [[adult-learning|adult learning]] contexts where learners study at work or at home, 24/7 support embedded in the course is a major benefit. The [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of|Open University's AIDA assistant]] found purpose-built, in-environment GenAI support increased [[student-engagement|engagement]] (doubled usage time in an exploratory trial), with 96% of students wanting it in their formal studies. - **Conversational, dialogic tutoring at scale.** [[conversational-ai]] tutors built on proven [[intelligent-tutoring]] technology ([[conversational-ai-tutors-framework|keep/change/center/study framework]]) promise high-quality, dialogue-based tutoring — engaging students' thoughts, questions, and [[misconceptions]] — that is far more scalable than human tutoring. - **Facilitation and analytics.** AI can support [[collaborative-learning|online discussions]] and [[learning-analytics]], forecasting engagement, and helping instructors allocate attention. [[hao-peer-exposure-bridging-social-capital-ai-summaries-2026|Hao & Cukurova (2026)]] add that LLM-generated discussion summaries can act as navigational [[scaffolding|scaffolds]] in large asynchronous forums — broadening students' peer exposure and the network conditions for bridging social capital without burdening students or instructors with the summarizing workload. - **Early-warning analytics for at-risk online learners.** [[zhang-ml-student-progress-programming-2026|Zhang, Jeffries & Koprinska (2025)]] show that interpretable [[reinforcement-learning|machine learning]] on content-interaction logs predicts module-level progress and flags dropout ("No submission") outcomes in large-scale online [[cs-education|programming]] courses up to 7–8 days before module deadlines, giving online instructors a concrete window to [[teacher-role|intervene]] with disengaged students rather than discovering failure only after the fact. - **Affordability and speed.** AI can generate course materials at a fraction of traditional cost — MAIC reduced MOOC course production from ~\$25K/60 hours to under \$2/30 minutes. ## Challenges of online teaching in the AI era The online medium and generative AI combine to intensify a specific cluster of challenges that instructors must confront head-on. Where face-to-face teaching can rely on presence, immediate accountability, and invigilation, online teaching must design explicitly for them. ### Academic integrity and cheating Online courses already present invigilation challenges — in-person proctoring is often unfeasible for distributed, asynchronous learners. Generative AI compounds this by making AI-generated work indistinguishable from student work and by enabling contract-cheating style shortcuts at scale. The knowledge base's evidence on [[academic-integrity]] and [[ai-misuse-learning-harm]] shows that misuse is driven less by AI errors than by students copying answers instead of learning. Because online assessment frequently cannot distinguish assisted from independent work, misuse can inflate immediate grades while eroding durable knowledge — a perceived-vs-actual gap that is especially dangerous at a distance where instructors have less visibility into student process. Detection tools are a partial, contested response ([[ai-detection|AI plagiarism detection]], [[remote-proctoring]]), and the knowledge base's stance favors [[authentic-assessment|authentic, process-revealing assessment]] over detection arms races. ### AI misuse and cognitive offloading The most serious risk is that online learners outsource the very cognitive work that builds understanding. The [[genai-performance-vs-learning|performance–learning gap]] shows generative AI easily boosts immediate performance while bypassing the deep processing required for durable learning. Field evidence is direct: - A causal RCT (~1,000 high-school math students) found unguarded AI assistance raised practice performance **+48%** but reduced unassisted, closed-book exam scores **−17%** — the students who never had AI access outperformed those who did. A [[guardrails|guardrailed]] hint-not-answer tutor eliminated the harm. - Population-scale behavioral data (3.2M ALEKS interactions) found study time on AI-susceptible problems fell **−26.9%** after ChatGPT's release, with a **−25% decline in odds of correct proctored retention items** — an effect that vanished under proctoring, pinning it on off-platform AI use. Online learning is particularly vulnerable: the medium already distances learners from immediate accountability, and self-paced, screen-based work invites the "ask for the answer" shortcut that [[cognitive-offloading]] [[research-methods-aied|research]] identifies as the core harm mechanism. The response is not to ban AI but to apply [[guardrails]] — hint-not-answer scaffolding, knowledge grounding, and [[human-in-the-loop-ai|human oversight]] — so that AI augments rather than replaces learner cognitive work. ### Other challenges - **Over-eager AI facilitation.** LLM facilitators are excessively eager to intervene in online discussions, which can irritate participants and derail good conversation; human caution is the better model ([[llm-facilitation-timing-online-discussions|Tsirmpas et al.]]). - **Motivation erosion.** The perceived availability of an effortless AI shortcut reduces autonomous [[motivation]] and persistence, compounding learning harm. - **Equity and the digital divide.** Access to reliable devices, connectivity, and high-quality AI varies; online learning with AI can widen rather than narrow [[equity-in-ai-education]] gaps ([[digital-divide]]). - **Data privacy and trust.** Online platforms collect rich learner data; AI systems raise transparency and privacy concerns ([[privacy]]), especially for adults balancing work and study. - **Organizational readiness.** The demise of KhanMigo — learners not actually engaging with the chatbot, with limited evidence of gains — cautions that technical capability must be matched with [[governance]] and organizational readiness. ### Assessment validity when a submission can be produced without the learner If invigilation is unfeasible and an agent can complete the work, the response the knowledge base favors is to change what counts as evidence rather than to police harder. Four design responses recur across the recent literature. **Pair the vulnerable task with a twin.** [[roe-assessment-twins-2026|Roe, Perkins and Giray (2026)]] keep the pedagogically valuable but AI-vulnerable assessment — the take-home essay or case analysis — and add a second, less vulnerable task that assesses the *same* [[learning-gains|learning outcomes]], scheduled closely enough for cross-verification and marked interdependently. Their mapping runs across Messick's six strands of validity evidence, and the design process is three steps: identify the vulnerabilities, align outcomes and choose the twin, then develop marking that connects the two. A short case variation, an explanation of one key decision, or a brief oral defense can serve as the twin. **Target ownership rather than authorship.** [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene|Ebrahimzadeh, Shibani and Buckingham Shum]] argue that blended human–AI authorship undermines several forms of validity evidence, and propose **coauthorship integrity** as validity evidence in its own right: violated when a student submits AI-generated content they do not understand. To check understanding at scale they report an **AI Viva**, a conversational agent that runs a hybrid viva voce with comprehension questions of controllable type and complexity, validated by expert educators and assessment specialists. For online courses this is a genuinely scalable form of verification, and it produces evidence about comprehension rather than about who was in the room. **Move the oral exam online.** [[asynchronous-oral-assessment-2026|Pentland, Lowenthal and Krier (2026)]] deliver prompts just in time and have students record brief, time-limited webcam responses they cannot revisit, graded against embedded rubrics with transcripts generated automatically. Across two studies — an intermediate accounting pilot and a data analytics course — students scored higher on these assessments than on in-person multiple-choice exams (significant in the second study, a positive trend in the first), with moderate cross-format correlations supporting convergent validity; students reported preparing differently and using more active study strategies. The format addresses the async problem directly: the thinking is performed live at a time of the student's choosing, at administrative cost that does not scale with cohort size. **Sequence the work so the reasoning is committed first.** [[brcic-effortless-trap-productive-struggle-2026|Brcic and Frljic (2026)]] frame the design question as **placement** rather than permission, and the causal evidence they assemble shows the outcome flipping on placement alone — the same unguarded helper that left high-school students about **17% worse** on an unaided exam did no harm once rebuilt to withhold answers, while a well-engineered [[intelligent-tutoring|tutor]] roughly **doubled** learning. Their diagnostic is the one to keep: *if letting AI in makes the task feel effortless, it is in the wrong place.* Operationally that yields a sequence an online course can write into the assignment — think, commit, use AI, critique, revise, explain — where the commitment step is what makes the rest assessable, because an artifact produced from scratch has no revision history to interrogate. Underneath sits a state this knowledge base names [[metacognitively-discordant-completion-genai-2026|metacognitively discordant completion]]: correct, complete work submitted by a student who knows the understanding never arrived, which in an asynchronous course is indistinguishable from success unless the design asks for something more. Whatever combination an instructor chooses, one constraint should shape it. Online study is frequently the only accessible option — students choose it because of employment, caregiving responsibilities, disabilities or geography — so verification has to be **small and proportionate**: a short recorded explanation, a personalized application, a response to an instructor-selected question, an annotated decision trail, a low-stakes individual check. Detection carries its own equity costs here, since tools that flag non-native writers disproportionately generate false positives ([[ai-detection]], [[digital-divide]]). ## Recommended pedagogical strategies for online teaching and learning - **Active and interactive learning.** Prefer strategies that keep students doing and thinking rather than passively receiving — [[active-learning]], interactive exercises, and [[socratic-method|Socratic]] dialogue. AI that prompts reasoning (rather than supplying answers) preserves the productive struggle and [[desirable-difficulties]] that build durable learning. - **Scaffolded, guided support.** Use [[scaffolding]] that fades as learners progress, and design [[self-regulated-learning]] supports so learners direct their own learning rather than depending on the tool. - **Collaborative and discussion-based learning.** Structure online discussions and group work deliberately; use [[collaborative-learning]] activities and, when AI participates, calibrate its facilitation and its role as a peer. - **Authentic, process-revealing assessment.** Shift toward [[authentic-assessment]] and assessments that capture process — drafts, oral defenses, self-explanation, reflective [[eportfolio|portfolios]] — which are more AI-resistant and reveal genuine understanding. - **Personalized and adaptive paths.** Use AI-enabled [[personalized-learning|personalization]] and [[adaptive-learning|adaptive]] activities to tailor pacing and difficulty, while keeping personalization deep (task sequencing, difficulty calibration) rather than merely surface-level (custom examples). - **Social presence and community-building.** Deliberately cultivate social presence and [[collaborative-learning|community]] — the core of the [[community-of-inquiry]] framework — through companion AI, synchronous check-ins, and peer interaction, since online isolation is a key barrier to [[student-engagement|engagement]] and belonging. In the AI era this means curating the three presences (cognitive, social, teaching) even as machine-generated discourse complicates who is "present" (see [[community-of-inquiry]]). - **Blended [[design-thinking]].** For hybrid formats, apply [[learning-design]] principles to decide what is best done synchronously vs. asynchronously and online vs. in person, and how AI supports each. - **Name the AI's role, and check the student still has one.** "Students may use AI" is too broad to design against. Roles carry different pedagogical consequences — a study of generative AI in marketing education distinguishes **tutor, teammate and tool** and shows each shaping teaching, social and cognitive presence differently ([[genai-marketing-education-roles-2026|GenAI in Marketing Education]]) — so an online activity should be able to state its division of labor: the AI's job here is X, the student's job is Y. If Y contains little thinking, the activity needs redesigning rather than a stricter policy. - **Sequence discussion as position, challenge, reconsideration.** The conventional asynchronous formula of posting once and replying twice is both superficial and agent-completable. A stronger structure asks students to commit to an interpretation, meet a counterexample or critique, then explain how their reasoning moved; what is graded is the movement between ideas rather than the post count. This also relocates teaching presence: replying mechanically to dozens of posts is the least valuable form of it, while synthesizing patterns across the discussion — the assumptions that keep recurring, the disagreements worth naming, the counterexample that unsettles a consensus — is the part no agent in this literature performs. - **Human-in-the-loop governance.** Keep educators and [[teacher-role|instructors]] in the loop over AI tools, grounded in [[tpack|pedagogical content knowledge]], so pedagogical intent — not the tool's default — drives design. ## Implications for online instructors and instructional designers - **Guardrail the AI, don't just supply it.** Use hint-not-answer [[scaffolding]] that keeps learner cognitive work in the loop; the [[guardrails|guardrailed]]-tutor RCT shows this eliminates the exam penalty that unguarded access causes. See the [[guardrails]] concept for the full design layer ([[prompt-engineering|prompting]], [[rag]] grounding, training, QA). - **Design AI-resistant and proctored/unassisted assessments.** Because online grading often can't distinguish assisted from independent work, include closed-book, proctored, or process-revealing assessments to surface and discourage misuse ([[ai-misuse-learning-harm]]). - **Teach AI literacy explicitly.** Help students recognize reliance patterns and calibrate trust ([[ai-literacy]]); build [[self-regulated-learning|self-regulation]] and [[metacognition]] to counter offloading. - **Embed AI in the learning environment, not as an external bolt-on.** Purpose-built, contextually-tuned assistants embedded in the course (like [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of|AIDA]]) outperform generic external chatbots and increase acceptance. - **Calibrate AI facilitation toward human caution.** When using AI to moderate discussions, prefer settings that intervene sparingly ([[llm-facilitation-timing-online-discussions|Tsirmpas et al.]]). - **Design for adult life constraints.** For adult and distance learners, prioritize mobile access, offline capability, and asynchronous availability ([[ai-adult-learning-guidelines-dis2026|AI-ALOE guidelines]]). - **Co-design with students and staff, and build governance.** Participatory development, senior sponsorship, cross-unit collaboration, and robust [[governance]] are enabling factors for responsible GenAI adoption. - **Assume an agent will attempt every unproctored activity, and design from that assumption.** The three-question test is quick and exposes weak activities: could an AI system complete this without the student understanding the material; what cognitive activity is supposed to produce the learning; what evidence will show the student performed it. When the first answer is yes and the other two are hard to answer, the problem is the learning design rather than the AI policy. - **Keep verification proportionate to the risk.** Where competence must be certified, prefer a twin task, an asynchronous oral defense or a short comprehension check over blanket monitoring; keep the rest of the course flexible for the learners who depend on that flexibility. - **Use analytics to support, not replace, teaching.** Leverage [[learning-analytics]] to forecast engagement and target support, but keep [[human-in-the-loop-ai|human oversight]] central. ## Connected Concepts - [[assessment-validity]] — validity of the inference from submitted work to learning - [[agentic-ai]] — autonomous systems that operate tools and platforms, including an LMS - [[community-of-inquiry]] — Community of Inquiry - [[pedagogy]] - [[learning-design]] - [[active-learning]] - [[collaborative-learning]] - [[scaffolding]] - [[self-regulated-learning]] - [[cognitive-offloading]] - [[academic-integrity]] - [[ai-misuse-learning-harm]] - [[ai-literacy]] - [[intelligent-tutoring]] - [[personalized-learning]] - [[adaptive-learning]] - [[student-engagement]] - [[digital-divide]] - [[governance]] - [[teacher-role]] - [[human-in-the-loop-ai]] - [[authentic-assessment]] - [[guardrails]] - [[ai-detection]] ## Connected Articles - [[ai-agents-complete-lms-assessment-validity-2026]] — Autonomous agents completed unproctored LMS assessments end to end: an assessment-validity problem (Hadjisolomou & El-Haddad 2026) - [[roe-assessment-twins-2026]] — Assessment twins: pairing a GenAI-vulnerable task with a closely scheduled, less vulnerable one on the same outcomes - [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene]] — Coauthorship integrity as validity evidence, and the AI Viva as scalable verification - [[asynchronous-oral-assessment-2026]] — Asynchronous oral assessments: time-limited unrevised recordings graded against embedded rubrics - [[brcic-effortless-trap-productive-struggle-2026]] — The effortless trap: placement of AI rather than permission or prohibition - [[metacognitively-discordant-completion-genai-2026]] — Metacognitively discordant completion: correct work submitted without understanding - [[decreasing-digital-distraction-college-online-learning-2026]] — Which self-regulated strategies coincide with low digital distraction (530 students) - [[song-genai-learning-partner-srl-over-time-2026]] — SRL as stable aptitude and fluctuating state; metacognition and well-being declining across a term - [[genai-marketing-education-roles-2026]] — AI as tutor, teammate and tool: roles and their effects on presence - [[kirsanov-beyond-detection-ai-online-assessments-2026]] — Beyond detection: assessment design for online settings - [[reconceptualizing-community-inquiry-generative-ai]] — Reconceptualizing Community of Inquiry in the age of generative AI - [[ai-student-engagement-online-learning-review-2025]] - [[lock-integrating-ai-online-learning-higher-ed-2025]] — Integrating AI in online learning in higher education: a four-theme critical literature review - [[ai-online-education-engagement-satisfaction-2026]] - [[ai-distance-education-systematic-review-2026]] - [[ai-decision-support-online-learning-assessment-2026]] - [[mooc-to-maic]] — From MOOC to MAIC: Reshaping Online Teaching and Learning through LLM-driven Agents - [[learnmate2-llm-adaptive-learning]] — LearnMate²: Personalized and Adaptive Support System for Online Learning - [[llm-facilitation-timing-online-discussions]] — Human and LLM Facilitator Tendencies in Online Discussions - [[elevate-genai-virtual-tutors]] — ELEVATE: Human-Centered GenAI Virtual Tutors - [[conversational-ai-tutors-framework]] — The Path to Conversational AI Tutors - [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of]] — Implementing AIDA at the Open University - [[ai-adult-learning-guidelines-dis2026]] — Guidelines for Designing AI Technologies to Support Adult Learning - [[deeptutor]] — DeepTutor: Toward Agentic Personalized Tutoring - [[educasim-cs1-instructional-practice]] — EducaSim: scalable role play for massive online courses - [[mejia-domenzain-ml-findings-teachers-blended-2026]] — Making ML findings accessible to teachers in blended classrooms - [[zhang-ml-student-progress-programming-2026]] - [[personalized-ai-generated-videos-preference-2026]] — Students prefer personalized AI-generated videos over non-personalized human-recorded ones (Tomlinson et al. 2026) - [[hao-peer-exposure-bridging-social-capital-ai-summaries-2026]] — AI-Generated Summary-Driven Learning Design in Online Discussion Forums - [[ai-training-science-teacher-tpack-distance-2026]] — Campus-based student teachers reported higher TPACK for AI-integrated science teaching than distance education peers (64.0% versus 47.4%) --- ## [Video in Education](https://edtechdev.github.io/aied/concepts/video-education/) > **Video in education** — the use of video as a medium for [[teacher-role|teaching]] and learning, and how [[generative-ai|generative AI]] is reshaping it: AI-generated and AI-[[personalized-learning|personalized]] instructional videos, AI avatars and presenters, adaptive video generation, video-based [[learning-analytics|learning analytics]] and attention/[[student-engagement|engagement]] sensing, and AI support for lecture-video consumption. The knowledge base treats video as both an established online-learning medium and a rapidly evolving site of AI innovation, spanning [[online-teaching-and-learning|online]], hybrid, and in-person teaching. ## Questions to Consider - Educational video has long been a "one-size-fits-all" resource — identical content for every learner. Generative AI now makes per-learner video feasible, and [[research-methods-aied|research]] suggests students value that personalization highly. What does personalization add beyond relevance — and what might it cost? - Students often say they still value a human instructor's presence and authenticity in video. Yet in head-to-head preference, personalized AI video can beat generic human-recorded lectures. What tradeoffs are learners actually making, and how durable are they? - AI avatars cloned from instructors can generate video at scale — but they can also trigger "uncanny valley" discomfort and [[ethics|ethical]] objections (environmental impact, labor, academic integrity). When is an AI presenter acceptable, and when does it cross a line that no technical fix addresses? - Much video research relies on students' preferences and [[self-report-measures|self-report]]. How well do stated preferences predict actual [[learning-gains|learning outcomes]] — and when might a video that "feels good" teach less well than one that does not? - Video analytics can detect attention, engagement, and dropout points. What are the [[pedagogy|pedagogical]] and [[privacy]] implications of instrumenting video learning this closely? ## Introduction Video is a cornerstone of contemporary education — especially [[online-teaching-and-learning|online and hybrid learning]] — prized for its flexibility, scalability, and consistency. Yet conventional instructional video is produced as a one-size-fits-all artifact, presenting identical content to every learner regardless of their interests, background, or [[prior-knowledge|prior knowledge]]. Generative AI is shifting video from a static broadcast medium to a dynamic, individually tailored one, and is also generating new questions about presence, [[trust]], [[privacy]], and measurement. ### How the knowledge base's research clusters - **AI-generated and personalized instructional video.** A central thread asks whether students accept AI-produced video and how it compares to human-recorded content. [[ai-generated-instructional-videos-computing-ed|Student surveys in computing education]] probe perceptions and preferences for AI-generated instructional video. In a large field deployment, [[personalized-ai-generated-videos-preference-2026|Tomlinson et al. (2026)]] found that students preferred AI-generated *personalized* videos over non-personalized human-recorded lectures — a preference in which the personalization effect outweighed the value placed on a human presenter. [[ai-video-dual-gatekeeping-2026|Dual gatekeeping research]] shows how instructor oversight ("gatekeeping") across two stages of AI video production yields more pedagogically grounded output, connecting to [[human-in-the-loop-ai|human-in-the-loop]] design. - **Adaptive and structured video generation.** [[courseblueprint-adaptive-video-generation|CourseBlueprint]] offers a structured pipeline that generates adaptive pedagogical video grounded in course corpora, showing that explicit pedagogical structure — not just [[ai-literacy|AI fluency]] — drives effective AI video. [[bespoke-industry-personalized-lecture-videos-2026|Bespoke]] applies the same logic at whole-lecture scale: from 31 graduate lectures it generated 209 videos for healthcare, finance, energy, and a generic audience, and 25 domain-matched experts rating 92 of them placed 87% at or above the midpoint written as "a standard MOOC lecture's quality" (mean 3.42 out of 5), at about \$0.22 in API cost per minute — with voice, slide timing, and layout the recurring defects. - **Video learning analytics and attention.** Instrumenting video reveals how learners engage. [[engagement-assessment-video|Engagement assessment in video learning]] and [[savvy-student-attention-video-learning|SAVVY]] visualize student attention during video-based learning, supporting [[learning-analytics|learning analytics]], [[self-regulated-learning|self-regulation]], and early-warning for disengagement. Segmentation work (e.g., [[adhd-video-segmentation-computing-education|temporal video segmentation]]) tailors video to individual differences. - **Avatars and presence.** AI avatars — virtual presenters and pedagogical agents — raise questions about identity, [[community-of-inquiry|social presence]], and trust. [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning|Avatar identity and epistemic trust]] examines how a presenter's apparent identity shapes learners' trust, while [[ai-psychotherapy-training-avatars|AI avatars in training]] extend the pattern to professional practice. - **In-video scaffolding comments — and what AI still gets wrong.** [[wang-chatgpt-comments-video-learning-scaffolding-2026|Wang, Du and Jin (2026)]] generate *i-Comments*, [[scaffolding]] messages rendered inside the video frame and synchronized to the content, by computing frame-level entropy and inserting support only in low-information intervals. Benchmarked against 120 comments from experienced instructors, ChatGPT's 1,000 comments were denser, far less structurally varied (POS 3-gram diversity 5.5–6.2% vs. 25.1–38.5%), harder to read on every readability index, and less topically aligned (emotional-support BERTScore 0.317 vs. 0.574); 40 learners rated the human comments significantly higher on timing and helpfulness, although a newer model narrowed that gap. The argument is a design argument as much as an automation one: support embedded in the media avoids the attention and cognitive cost of pausing to query a separate [[conversational-ai|chatbot]] ([[ai-feedback-quality|feedback quality]], [[social-emotional-learning|emotional support]]). - **AI support for lecture-video consumption.** Beyond generation, AI helps learners and teachers work with existing video: [[bilingual-llm-lecture-companion-srl-2026|bilingual LLM lecture companions]] support self-regulated learning with recorded lectures, and [[gemini-lualatex-physics-video-transcription-2026|transcription pipelines]] convert lecture video into accessible text. ### Personalization versus human presence A recurring tension is whether the value of [[personalized-learning|personalization]] can outweigh the value of a visible human instructor. [[personalized-ai-generated-videos-preference-2026|Tomlinson et al. (2026)]] frame personalization and social presence as *partially substitutable signals of instructional care*: human delivery enhances [[affective-computing|affective]] experience and authenticity, while personalization enhances relevance — and students are willing to trade one for the other. Their large-course ranking data (88.4% preferred some personalized video; only 73.8% preferred human-recorded) suggest personalization is now often the more influential factor, pointing toward a complementary model where human instructors supply expertise and social connection while AI extends their reach with individually tailored media. ### Design, ethics, and measurement Producing effective AI video requires pedagogical structure and human oversight, and it raises distinct concerns: AI presenters may evoke discomfort or distrust (the "uncanny valley"); generative video risks factual inaccuracy that learners may not catch; scaling personalization requires collecting or inferring learner attributes, with attendant [[privacy]], bias, and [[governance]] concerns; and a subset of learners object to AI-generated instruction on principled grounds (environmental impact, labor, automation, [[academic-integrity|academic integrity]]). Measurement is likewise in flux — much evidence rests on stated preference and perceived value rather than objective learning outcomes, so preference data must be read alongside (often forthcoming) outcome data. ## Connected Concepts - [[online-teaching-and-learning]] — video as a core medium of online and hybrid instruction - [[generative-ai]] — the engine of AI-generated and personalized video - [[personalized-learning]] — personalization as the driver of AI video's appeal - [[adaptive-learning]] — adaptive video generation and pacing - [[multimodal]] — video combining visual, audio, and textual modalities - [[learning-analytics]] — analytics on video engagement and attention - [[student-engagement]] — the engagement that video personalization aims to boost - [[pedagogical-agent]] — AI avatars/presenters as virtual pedagogical agents - [[llm]] — large language models underlying script and video generation - [[trust]] — learner trust in AI presenters and content ## Connected Articles - [[bespoke-industry-personalized-lecture-videos-2026]] — Bespoke: generating MOOC-quality industry-personalized lecture videos at scale (Puech et al. 2026) - [[wang-chatgpt-comments-video-learning-scaffolding-2026]] — ChatGPT-generated in-video comments: entropy timing, quality gaps vs. human comments (Wang, Du & Jin 2026) - [[personalized-ai-generated-videos-preference-2026]] — Students prefer personalized AI-generated videos over non-personalized human-recorded ones (Tomlinson et al. 2026) - [[ai-generated-instructional-videos-computing-ed]] — Student perceptions/preferences of AI-generated instructional video in computing education - [[ai-video-dual-gatekeeping-2026]] — Dual gatekeeping for pedagogically grounded AI video creation - [[courseblueprint-adaptive-video-generation]] — CourseBlueprint: adaptive pedagogical video generation - [[engagement-assessment-video]] — Engagement assessment in video learning - [[savvy-student-attention-video-learning]] — Student attention visualization for video-based learning - [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning]] — How avatar identity shapes epistemic trust in AI-mediated learning - [[bilingual-llm-lecture-companion-srl-2026]] — Bilingual LLM lecture companions for self-regulated learning - [[adhd-video-segmentation-computing-education]] — Temporal video segmentation for individual differences - [[ai-psychotherapy-training-avatars]] — AI avatars in psychotherapy training - [[gemini-lualatex-physics-video-transcription-2026]] — Transcribing physics lecture video into accessible text - [[pivot-generative-video-tutors-stem-2026]] — From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning --- ## [Learning Theories](https://edtechdev.github.io/aied/concepts/learning-theories/) > **Learning Theories** — the family of frameworks that explain how learning happens, and the umbrella concept for the knowledge base's theory-related ideas. In [[ai-education|AI in education]], learning theories shape both how AI systems are designed (the pedagogy they embody) and how the field interprets whether AI "works": the same tool can be a scaffold under [[constructivist]] assumptions, a reinforcement engine under [[behaviorism]], or a cognitive-load hazard under Cognitive Load Theory. ## Questions to Consider - Think about an AI tutor or adaptive system you have used or seen. What assumptions did it make about how people learn — did it reward right answers (behaviorism), build understanding (constructivism), or manage mental effort (cognitive load)? Did its makers ever state those assumptions? - The page describes a recurring gap: discourse espouses constructivism while AI implementations default to drill-and-feedback mechanics. Where have you seen a tool claim to support deep learning but actually just reinforce surface responses? - The same AI tool can look like success under one theory and failure under another — inflated homework scores read as learning under behaviorism but as a failure to build durable understanding under constructivism. Which lens is fairer for judging whether students actually learned? - Because every AI tutor embeds a theory whether its designers say so or not, 'does it work?' may be the wrong question. What is the better question to ask about an educational AI, given the theories it could embody? - Generative AI is prompting educators to consider new theories — like learning as iterative co-construction between humans and AI, or AI as a cognitive partner across the lifespan. Does the rise of AI genuinely require new learning theories, or do existing ones still suffice? - Learning-gains measurement is itself 'theory-laden': an instrument built on one theory may not capture the gains another predicts. How might two researchers with different theoretical commitments look at the same data and reach opposite conclusions about whether an AI tool worked? ## Introduction This is the umbrella concept for the knowledge base's learning-theory strand. Learning theories sit at the heart of AI in education because every AI tutor, adaptive system, and feedback tool embeds assumptions about how people learn — whether the designers state them or not. The knowledge base documents these theories individually and treats them as the conceptual lens through which AI's design and effects are evaluated. The research field that tests those frameworks is a separate page: [[learning-sciences|the learning sciences]] study learning empirically — experiments, classroom trials and design-based research — and treat a theory as something to be confirmed or falsified, whereas this page collects the frameworks themselves; the two phrases are kept distinct in this knowledge base, the singular "learning science" naming this theory strand. ### The learning-theory landscape The knowledge base documents several families of learning theory, each with its own concepts: - **Classical learning theories.** [[behaviorism]] (learning as observable behavioral change through reinforcement and drill-and-practice) and [[constructivist|constructivism]] (learning as active knowledge construction) are the two poles that recur most often in AI research. The field frequently exhibits a "constructivism in name, behaviorism in practice" gap, where discourse espouses construction but AI implementations default to drill-and-feedback mechanics.([[ai-vocational-education-training-review]]) [[cognitive-psychology|cognitivism]] is the third classical pole — learning as a change in internal mental representations — and the theory most responsible for AIED's signature contributions ([[knowledge-tracing]], [[cognitive-diagnosis]], [[student-modeling]], [[intelligent-tutoring]]). - **Sociocultural and developmental theories.** [[sociocultural-learning]] holds that learning and development arise through social participation and are mediated by cultural tools and more knowledgeable others — spanning the [[sociocultural-learning|Zone of Proximal Development]], [[scaffolding]], apprenticeship, communities of practice, and [[distributed-cognition|distributed cognition]]. In the AI age, [[generative-ai|generative AI]] is increasingly framed as a *mediational agent* that both mediates activity and generates contingent contributions to interaction.([[generative-ai-mediational-agent-sociocultural-2026]]) - **Cognition and cognitive architecture.** [[cognitive-psychology|Cognitive psychology / cognitivism]] is the umbrella for this family: Cognitive Load Theory (how working-memory limits shape instruction), Dual-Process Theory (fast intuitive vs. slow deliberative processing), and [[metacognition]] (monitoring and regulating one's own learning) explain the *internal* mechanisms that AI tools engage or bypass. - **Motivation and self-direction.** [[self-determination-theory]] (autonomy, competence, relatedness), [[self-efficacy]] (confidence in one's capability), [[self-regulated-learning]] (goal-setting, monitoring, and adjustment), and [[motivation]] explain why learners engage with AI the way they do. - **Learning context and activity.** [[experiential-learning]], [[active-learning]], [[project-based-learning]], [[collaborative-learning]], [[transfer-of-learning]], [[desirable-difficulties]], and [[embodied-learning]] describe the kinds of activity and context that produce durable learning. ### Learning theories and learning gains Learning theories are ultimately evaluated by their outcomes, and the knowledge base's [[learning-gains]] concept is where theory meets evidence. Each theory makes a different prediction about what *counts* as learning and how to measure it: behaviorism predicts observable performance gains on drill-and-feedback tasks; constructivism predicts deeper understanding that transfers to novel problems; sociocultural theory predicts gains in participation and mediated problem-solving; and cognitive-load theory predicts gains only when instruction respects working-memory limits. This is why [[learning-gains]] measurement is theory-laden — an instrument built on one theory may not capture the gains another predicts. In the AI era, the sharp divergence between AI-assisted performance and unassisted [[learning-gains|learning gains]] (see [[generative-ai-reduced-study-time-math]], [[stromberg-generative-ai-learning-penalty-secondary-2026]]) can be read through this lens: a behaviorist reading sees inflated homework scores as success, while a constructivist reading, focused on durable understanding, sees the same evidence as a failure to learn. Connecting theories to [[ai-ed-evaluation|evaluation]] and [[learning-gains|measured outcomes]] is thus essential to deciding which theoretical lens a given AI system actually satisfies. ### Why learning theories matter for AI in education Learning theories matter for three reasons: - **They predict AI's effects.** Whether an AI tool improves or harms learning depends on which mechanism it activates. A tutor that gives away answers harms under a constructivist lens (it bypasses construction), is neutral under behaviorism (it reinforces), and raises cognitive-load concerns (it offloads rather than builds). The knowledge base's [[cognitive-offloading]] and [[cognitive-offloading|Over-Reliance]] concepts capture the risk side of this. - **They expose the theory-practice gap.** Empirical work repeatedly finds that AI implementations embody different theories than the discourse claims — most notably constructivist language paired with behaviorist drill-and-practice mechanics.([[ai-vocational-education-training-review]]) Evaluating AI therefore requires asking *which* theory a system actually embodies, not just whether it "works." - **They are being actively rethought.** Generative AI is [[prompt-engineering|prompting]] educators to revisit whether the classical theories suffice. [[generativism-learning-theory|Generativism]] proposes that learning in the AI age increasingly occurs through iterative co-construction between human learners and AI systems, extending rather than replacing the classical four (behaviorism, cognitivism, constructivism, connectivism).([[generativism-learning-theory]]) ### New theoretical directions from recent AIED work Recent theoretical work extends the classical strand in several directions, each re-centering the human–AI relationship rather than treating AI as a neutral tool: - **Human–AI co-[[regulation]].** A developmental framework positions AI not as an external instrument but as a **cognitive partner** that co-regulates thinking, learning, and self-control across the lifespan.([[ai-cognitive-partner-co-regulation-learning]]) Drawing on executive function, [[metacognition]], distributed cognition, and sociocultural development, it casts AI in four roles — scaffold, metacognitive support, external memory / [[cognitive-offloading]] system, and decision partner — with the framework most relevant from middle childhood onward. - **Ensemble Cognition.** A [[philosophy-of-ai-in-education|philosophical]] framework reconceptualizes thinking as emerging from dynamic interactions between human and artificial agents rather than residing solely in individual minds.([[ensemble-cognition-philosophy-ai-education]]) It challenges the "consciousness paradigm" (the autonomy, consciousness, and stability assumptions) and articulates five features — distributed agency, dynamic centrality, cognitive orchestration, multi-representational integration, and context-sensitive switching — while distinguishing AI's **functional agency** from moral responsibility. - **[[self-directed-learning|Self-Directed]] Growth / A2PL.** An extension of [[self-regulated-learning|self-directed learning]] integrates Generative AI with [[learning-analytics|learning analytics]] to cultivate **Self-Directed Growth**, operationalized through the Aspire to Potentials for Learners (A2PL) model.([[self-directed-growth-generative-ai-learning-analytics]]) It reconfigures learner aspirations (humanistic), complex thinking (constructivist), and self-assessment (pragmatic) into a single competency, positioning GAI as a non-prescriptive collaborative scaffold rather than a content provider. - **Deceptive overgeneralization.** [[deceptive-overgeneralization-adaptive-learning-2026|An, McLaren, and Stamper (2026)]] extend the ACT-R / Knowledge-Learning-Instruction tradition by theorizing when observed correctness masks incomplete conditional understanding: learners compile an overgeneralized production that omits an application constraint yet still performs correctly — a failure mode that adaptive mastery systems, and even traditional instruction, can miss unless they test *when to withhold* an action. - **Executable KLI theory.** [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit, Koedinger & Carvalho (2025)]] ground the Knowledge-Learning-Instruction framework in an executable computational model (the Apprentice Learner framework with an ACT-R-style memory mechanism) that reproduces a cross-over interaction in human data: pure practice aids verbatim fact memory (by delaying forgetting) while example-integrated practice aids generalizable skill induction. Because KLI ties constant (fact) knowledge to memory processes and variable (skill) knowledge to induction, the result is a predicted content–treatment interaction rather than a contradiction between testing and worked-example recommendations — and the model's success only when a memory mechanism is present demonstrates that practice and examples play distinct, complementary roles. ### How the knowledge base organizes this strand Rather than treating learning theories as abstract philosophy, the knowledge base grounds each in the AI-in-education research that uses it. The [[constructivist]] and [[behaviorism]] pages document how AI designs embody (or betray) each theory; Cognitive Load Theory, [[self-regulated-learning]], [[metacognition]], and [[transfer-of-learning]] connect theory to specific AI mechanisms and outcomes. This mirrors how the knowledge base treats other umbrella domains like [[feedback]] and [[assessment]] — a coherent system of interacting concepts rather than isolated pages. ### Learning theories and "education about AI" Learning theories also appear as content in [[ai-literacy|AI literacy]] curricula: learners study behaviorism, cognitivism, constructivism, and connectivism to understand the [[pedagogy|pedagogical]] assumptions behind the tools they use.([[generativism-learning-theory]]) Teaching this strand gives students (and educators) the vocabulary to critique why an AI product is built the way it is — and whether its mechanics serve the learning goal at hand. - **The mediational agent.** Warschauer, Tate, and Ritchie (2026) argue generative AI breaks the sociocultural distinction between mediational means and social interaction, proposing the *mediational agent* — a system that both mediates action and generates contingent, non-accountable contributions, occupying a hybrid space between a tool and a social partner. This yields five human-first habits of participation (primacy of human cognition, purposeful [[student-engagement|engagement]], supervisory agency, epistemic vigilance, reflective self-regulation).([[generative-ai-mediational-agent-sociocultural-2026]]) ### Theories proposed for the AI era Alongside the classical families, the knowledge base documents theories written specifically for learning with AI systems, and these carry the design implications that the older theories leave open. [[yan-agentivism-learning-theory-ai-2026|Agentivism (Yan and Gašević 2026)]] is a mid-range example: it defines learning as durable growth in human capability rather than successful task completion, names four mechanisms (delegated agency, epistemic monitoring and verification, reconstructive internalization, and transfer under reduced support), and states six testable propositions, among them that AI support preserving learner responsibility for problem framing, criteria setting and justification produces stronger learning than support that delivers answers. ## Connected Concepts - [[behaviorism]] - [[cognitive-psychology]] - [[constructivist]] - [[metacognition]] - [[distributed-cognition]] - [[self-regulated-learning]] - [[self-determination-theory]] - [[self-efficacy]] - [[motivation]] - [[sociocultural-learning]] - [[scaffolding]] - [[transfer-of-learning]] - [[learning-gains]] - [[desirable-difficulties]] - [[active-learning]] - [[experiential-learning]] - [[collaborative-learning]] - [[embodied-learning]] - [[cognitive-offloading]] - [[learning-design]] - [[learning-sciences]] - [[philosophy-of-ai-in-education]] - [[ai-education]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education ## Connected Articles - [[yan-agentivism-learning-theory-ai-2026]] — A mid-range learning theory for human-AI interaction, with four mechanisms and six testable propositions (Yan and Gašević 2026) - [[powerful-learning-with-emerging-technology-2025]] — Three design principles for emerging technology: evidence-based, learner-centered, skill-building - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[airis-hybrid-human-ai-cognition-2026]] — AI-Augmented Inquiry and Regulation in Hybrid Systems (AIRIS) - [[zhu-e3-hot-embodied-intelligence-sustainable-learning]] — Fostering Sustainable Learning via Embodied Intelligence (E3-HOT) - [[voicu-ai-interpretive-cognition-ssh-2026]] - [[ai-cognitive-partner-co-regulation-learning]] — Positions AI as a cognitive partner in human-AI co-regulation; developmental framework across the lifespan - [[ensemble-cognition-philosophy-ai-education]] — Ensemble Cognition: a philosophical framework reconceptualizing thinking as human–AI interaction - [[self-directed-growth-generative-ai-learning-analytics]] — Self-Directed Growth and the A2PL model extending self-directed learning with GenAI - [[generativism-learning-theory]] — Proposes a new learning theory for the generative AI age, revisiting the classical four - [[ai-vocational-education-training-review]] — Documented the constructivism/behaviorism theory-practice gap in AI for VET - [[genai-educational-outcomes-meta-analysis]] - [[vargas-situated-learning-ai-review-2024]] - [[raffaghelli-situated-ai-ethics-2026]] - [[elsayed-pedagogical-symbiosis-posthuman-learner]] - [[niari-ai-pedagogical-mediator-collaborative-learning]] - [[videla-embodied-ai-education-choreography]] - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a Mediational Agent - [[strydom-human-gai-paradigms-2026]] — Framing human-AI dynamics: seven GAI engagement paradigms (Strydom 2026) - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[lukesova-clue-before-correction-2026]] — Clue Before Correction: ChatGPT for Autonomous Language Learning - [[rachatasumrit-example-problem-ratio-2026]] --- ## [Behaviorism](https://edtechdev.github.io/aied/concepts/behaviorism/) > **Behaviorism** — the learning theory that treats learning as a change in observable behavior produced by stimulus–response associations and reinforcement, rather than by changes in internal mental states. In [[ai-education|AI in education]], behaviorist principles underlie the drill-and-practice, immediate-feedback, and adaptive-pacing designs that dominate many [[intelligent-tutoring]] and [[adaptive-learning]] systems.([[ai-vocational-education-training-review]]) ## Questions to Consider - Behaviorism treats learning as a change in observable behavior driven by stimulus-response and reinforcement — not by internal mental states. Before reading, which educational experiences in your own past were built on reward, repetition, and immediate feedback? What did they succeed at, and what might they have missed? - A surprising finding on this page is that even where educational discourse espouses rich constructivist theories, actual AI implementations are predominantly behaviorist — drill-and-practice, immediate feedback, adaptive pacing. Why do you think behaviorist mechanics dominate in practice despite being out of fashion in theory? - The page warns of an educational 'Turing Trap' — using AI to replicate rather than augment human instruction. If an AI system optimizes for correct responses and efficiency, what might it quietly optimize away in a learner's active knowledge construction? - Behaviorist designs are described as powerful for foundational fluency — vocabulary, arithmetic, code syntax — but inadequate for higher-order, conceptual, or agentic learning on their own. Where in your own learning would a drill-and-practice AI help, and where would it actively hurt? - The design question posed is not whether behaviorism is 'right' but whether a given system's mechanics serve the learning goal. How would you tell whether an AI tutor's immediate-feedback, adaptive-pacing design is building genuine transferable understanding or just making observable performance look good? ## Introduction Behaviorism holds that learning is the strengthening or weakening of stimulus–response connections through reinforcement, and that unobservable mental constructs are poor explanations of learning. Its applied legacy in education is **programmed instruction and drill-and-practice**: presenting content in small steps, eliciting a response, and immediately reinforcing correct answers. These principles map cleanly onto the mechanics of [[adaptive-learning]] and [[intelligent-tutoring]] systems, which adapt pacing and difficulty to student responses and provide immediate feedback. ## Core ideas - **Learning is behavioral change.** The target is a measurable change in performance, not an internalized understanding. This makes behaviorist designs natural for observable outcomes like fluency, speed, and accuracy. - **Reinforcement drives learning.** Correct responses are reinforced and errors corrected, typically with immediate feedback — a design pattern ubiquitous in AI tutoring and drill systems.([[ai-vocational-education-training-review]]) - **Small steps and scaffolding by pacing.** Instruction is broken into incremental units with feedback at each step, analogous to the way adaptive systems sequence practice. - **The learner is largely passive in knowledge construction.** The environment (or system) structures and rewards responses; the learner responds rather than constructs meaning — the direct opposite of [[constructivist]] assumptions. ## Behaviorism and AI in education ### Behaviorist designs dominate practice Empirical work repeatedly finds that actual AI implementations are predominantly **behaviorist or cognitively oriented** — emphasizing drill-and-practice, immediate [[feedback]], and adaptive pacing — even where discourse espouses richer theories. A [[meta-analysis-systematic-review|systematic review]] of AI in vocational education and training (VET) concluded that constructivist theories are espoused in VET discourse while **behaviorist AI implementations dominate in practice**, and warned of an educational "Turing Trap" — using AI to replicate rather than augment human instruction.([[ai-vocational-education-training-review]]) ### The tension with constructivism and agency The behaviorist emphasis on response-and-reinforcement sits in direct tension with [[constructivist]], [[self-regulated-learning]], and [[agency]] goals. When AI systems optimize for correct responses and efficiency, they can under-serve the learner's active knowledge construction, critical reflection, and autonomous decision-making. This is the same gap flagged in the [[constructivist]] "constructivism in name, behaviorism in practice" pattern — and it connects behaviorism to debates about [[cognitive-offloading]] and [[cognitive-offloading|Over-Reliance]] when AI does the cognitive work for students. ### Where behaviorist designs still fit Behaviorist principles remain well suited to: - **Foundational skill and fluency building** — where repetition and immediate feedback measurably improve automaticity (e.g., vocabulary, arithmetic, code syntax). - **[[adaptive-learning]] and [[intelligent-tutoring]]** — which rely on step-wise practice, response-driven pacing, and immediate feedback.([[ai-vocational-education-training-review]]) - **Low-stakes [[formative-assessment]]** and drill in well-defined domains where the target outcome is observable and the path to it is largely procedural. The design question is not whether behaviorism is "right" but whether a given AI system's behaviorist mechanics serve the *learning goal* — for procedural fluency they can be powerful; for higher-order, conceptual, or agentic learning they are inadequate on their own. ## Behaviorism and "education about AI" Behaviorism also appears in how learners encounter AI as a topic. The theory is one of the four dominant [[learning-theories|learning theories]] — behaviorism, cognitivism, constructivism, and connectivism — that [[generative-ai|generative AI]] is [[prompt-engineering|prompting]] educators to revisit.([[generativism-learning-theory]]) It is also referenced in cooperative-learning and design contexts as part of the theoretical backdrop learners are taught.([[ccct-cooperative-learning-technique]]) Understanding behaviorism helps learners see why many AI tools (and the products built on them) are designed for response-and-reinforcement rather than for deeper construction. ## Implications for design and research 1. **Match mechanics to goals.** Behaviorist drill-and-feedback designs suit procedural fluency and observable outcomes; they are a poor fit for conceptual, transferable, or agentic learning goals on their own. 2. **Watch the theory-practice gap.** Researchers should check whether an AI implementation's behaviorist mechanics are serving the espoused learning goal or quietly replicating the "Turing Trap" of AI as an answer machine.([[ai-vocational-education-training-review]]) 3. **Pair behaviorism with richer scaffolds.** Immediate-feedback designs are most effective when embedded in a broader [[scaffolding]] and [[self-regulated-learning]] context, rather than standing alone as pure drill. 4. **Evaluate observable *and* transferable outcomes.** Behaviorist success criteria (speed, accuracy) should be complemented by measures of whether learning transfers and generalizes, per [[transfer-of-learning]] and [[research-methods-aied]]. ## Connected Concepts - [[constructivist]] - [[cognitive-psychology]] — Cognitivism, the third classical pole of learning theory - [[learning-design]] - [[adaptive-learning]] - [[intelligent-tutoring]] - [[feedback]] - [[formative-assessment]] - [[self-regulated-learning]] - [[agency]] - [[cognitive-offloading]] - [[learning-theories]] ## Connected Articles - [[ai-vocational-education-training-review]] — Behaviorist AI designs dominate VET practice despite espoused constructivism; the "Turing Trap" - [[generativism-learning-theory]] — Behaviorism among the four dominant theories generative AI is prompting a rethink of - [[ccct-cooperative-learning-technique]] — Behaviorism cited in cooperative-learning design for higher education - [[wang-multi-agent-systems-learning-designers-2025]] — Behaviorist persona among collaborative multi-agent design approaches --- ## [Constructivism](https://edtechdev.github.io/aied/concepts/constructivist/) > **Constructivism** — the [[learning-theories|learning theory]] that knowledge is actively built by the learner through experience, reflection, and interaction, rather than passively received from an instructor or system. In [[ai-education|AI in education]], constructivism underlies the design commitment that AI tools should support learners' own knowledge construction — [[prompt-engineering|prompting]], questioning, and [[scaffolding]] — rather than perform the [[cognitive-offloading|cognitive work]] for them.([[ai-vocational-education-training-review]])([[genai-mindtool-generative-learning]]) ## Questions to Consider - Have you ever 'learned' something in class only to realize you couldn't actually explain or use it later? What was missing — and what does that tell you about how real understanding forms? - Constructivism claims knowledge is built, not transmitted. If that's true, what happens when an AI tutor simply supplies the correct answer? - The phrase 'constructivism in name, [[behaviorism]] in practice' describes AI tools that claim to support [[active-learning|active learning]] but actually run drill-and-practice. Have you seen this gap? How would you detect it in a tool you're evaluating? - Papert's constructionism says we learn most powerfully by building shareable artifacts. In the AI era, one framework puts it as: 'the AI writes the code, but the student writes the model.' What is a student actually constructing when AI handles the mechanics? - Some AI tools practice 'generative refusal' — withholding answers and posing questions instead. When would deliberately withholding help be more pedagogically valuable than providing it? - If knowledge is constructed, then AI literacy isn't learned by hearing lectures about AI — it's learned by using, critiquing, and building with AI. What does that imply about how AI literacy should be taught to you or your students? ## Introduction Constructivism is a family of theories rather than a single doctrine, but its core claim is shared: learners do not absorb meaning; they construct it. Understanding in this view is not the accumulation of transmitted facts but the active organization of experience into mental models. This has direct implications for how AI in education should be designed, evaluated, and taught — and it helps explain both the promise and the risk of [[generative-ai|generative AI]] in the classroom. **[[mishra-control-vs-agency-history-2025|Mishra et al.]]** contrast Papert's constructionism (Logo, microworlds, debugging-as-learning) with Anderson's cognitive tutors as competing visions of creative agency vs. systematic control in AIED history. ## Core ideas - **Knowledge is constructed, not transmitted.** Learners build understanding by acting on the world, reconciling new information with [[prior-knowledge|prior knowledge]], and reflecting on the results. An AI tutor that simply supplies correct answers bypasses the constructive activity that produces durable understanding.([[generative-refusal-ai-tools-for-thought]]) - **Prior knowledge shapes new learning.** New ideas are interpreted through the learner's existing mental models, so instruction must surface and build on what learners already know — a principle directly relevant to [[misconceptions]] and to AI tutors that adapt to the learner. - **Social interaction supports construction.** A major strand — social constructivism — holds that meaning is co-constructed through dialogue, collaboration, and culturally [[situated-learning|situated]] activity. This connects constructivism to [[collaborative-learning]] and to [[socratic-method]] approaches in which AI prompts rather than dictates.([[ai-agents-constructive-conflict-design-education-2026]]) - **Construction is visible in activity.** Learners reveal (and consolidate) their understanding by generating, explaining, and producing — which is why the [[icap-framework|ICAP framework]] ranks "constructive" and "interactive" [[student-engagement|engagement]] above "active" and "passive" modes.([[hingle-collaborative-ai-literacy-2025]])([[icap-cognitive-engagement-llm-agents]]) ## Constructionism **Constructionism** is the branch of constructivism associated with Seymour Papert that adds a specific claim: learning happens most powerfully when learners construct *external, shareable artifacts* — physical or digital objects they design, build, and debug. Where Piagetian constructivism focuses on the internal mental construction of knowledge, constructionism holds that this construction is best supported and made visible through making something tangible (Harel & Papert, 1991). In [[history-of-aied|AIED history]], constructionism stands as the "agency" pole of the field's central control-vs-agency tension, set against Anderson's structured cognitive tutors. - **Logo and microworlds.** Papert co-developed Logo (1967) with its iconic "turtle" — a programming microworld where children explore geometry and other powerful ideas by commanding and debugging a visible agent. Debugging is reframed as a natural, valuable part of learning, not failure.([[mishra-control-vs-agency-history-2025]]) - **Construction over instruction.** Constructionism critiques "instructionism" — the assumption that [[teacher-role|teaching]] is the efficient transfer of knowledge — and instead positions learners as [[agentic-ai|autonomous agents]] who construct understanding through projects and experimentation (Papert, 1980, *Mindstorms*). - **Lineage into modern [[edtech-platform|edtech]].** Logo's emphasis on creative, hands-on construction underpins [[game-based-learning]], [[project-based-learning]], [[educational-robotics|robotics]] (LEGO Mindstorms, [[cs-education|Scratch]], programmable bricks), and the broader maker movement. - **The constructionist legacy in AI.** Constructionism implies AI tools should serve as **materials to build with** — thinking tools and creative co-constructors that the learner directs — rather than as answer-providing instructors. This is the direct ancestor of the knowledge base's [[genai-mindtool-generative-learning|mindtool]] framing of generative AI and of design commitments that preserve [[agency|learner agency]] over the learning process.([[educational-robotics-pathways-2026]]) - **Constructionism in the GenAI era: learn by writing the model, not the code.** The arrival of code-generating AI has *renewed* constructionism as a design response rather than weakening it. Gousopoulos's **Code-to-Learn with Generative AI (CtL-GenAI)** framework synthesizes constructionism, cognitive-load theory, [[self-regulated-learning]], the [[icap-framework|ICAP]] engagement model, [[productive-failure|productive failure]], and [[sociocultural-learning|sociocultural]] scaffolding for upper-secondary students building software with AI. Its organizing claim — **"the AI writes the code, but the student writes the model"** — reframes the construction target: when GenAI does the syntactic work of writing code, the learner's construction shifts to building and debugging the *conceptual model* the code expresses. CtL-GenAI defines model authorship as a construct with four facets and ordered levels carrying observable indicators, and formalizes a partial-credit, falsifiable measurement model to test whether such learning actually occurs.([[code-to-learn-genai-artifact-construction-2026]])([[ai-writes-code-student-writes-model-2026]]) This is constructionism's classic "make something shareable and debug it" updated so that the artifact the student makes and reflects on is a mental model made visible, not merely source code — and it couples the theory to an explicit measurement program so the claim becomes empirically testable. Constructionism is thus both a learning theory and a critique: it insists that the purpose of education is not to reproduce existing knowledge structures but to empower learners to construct and transform them — a stance with clear implications for whether AI in education reinforces or challenges established hierarchies. ## Constructivism and AI in education ### AI for constructivist learning Well-designed AI can enable construction at scale. [[intelligent-tutoring]] and [[intelligent-tutoring|AI Tutoring]] systems can pose problems and guide [[help-seeking]] instead of giving away answers; [[simulation]] and [[game-based-learning]] environments let learners build and test mental models; and [[project-based-learning]] and [[experiential-learning]] activities supported by AI give learners authentic construction tasks. The central design pattern is **[[scaffolding]]** — calibrated support that fades as competence grows — rather than completion.([[conversational-ai-tutors-framework]])([[embodied-inquiry-ai-facilitator-physics-2026]]) Classifying the *questions learners ask* is one way to see construction happening, and to act on it. [[lee-learner-question-types-ai-education-2026|Lee, Atif & Kang (2026)]] sort 434 authentic learner questions from 11 IT students across 12 courses into three constructivist instructional roles — knowledge transmitter, facilitator, and co-learner — and train four transformers to recognize them. DeBERTa classified factual knowledge-transmitter questions at 96.67% precision but facilitator questions at only 78.79%, and every model confused the two higher-order roles most often: detecting dialogic, exploratory inquiry is far harder than detecting information-seeking. Because the typology treats questions as diagnostic evidence of epistemic engagement rather than as mere inputs, it supports a distinctly constructivist design move — when a learner repeatedly asks only factual questions, the system can prompt reflective, exploratory questioning that develops [[metacognition]] and critical inquiry rather than answering at whatever depth the learner's question implies. ### The risk of "constructivism in name, behaviorism in practice" Empirical work repeatedly finds a gap between espoused constructivist goals and actual AI implementations. A [[meta-analysis-systematic-review|systematic review]] of AI in vocational education, for instance, found that constructivist theories are espoused in VET discourse while **behaviorist drill-and-practice designs dominate in practice**, and warned of an educational "Turing Trap" — using AI to replicate rather than augment human instruction.([[ai-vocational-education-training-review]]) This pattern generalizes across the field: - When generative AI completes writing, reasoning, or code for students, the learner loses the constructive thought process the task was designed to build — the concern central to [[cognitive-offloading]] and [[cognitive-offloading|Over-Reliance]].([[generative-refusal-ai-tools-for-thought]]) - AI implementations that emphasize adaptive feedback and efficiency frequently under-serve the learner-agency, critical-reflection, and autonomous-decision goals that constructivism implies.([[ai-vocational-education-training-review]]) ### The naming trap: generative AI output is not generative learning A recurring confusion in the field turns on a name collision. **Generative AI** names a class of *technology* — models that generate text, images, or code. **Generative learning** (Wittrock's generative-learning theory) names a *learner activity* — the learner actively making meaning by constructing connections between new information and prior knowledge, through strategies such as summarizing, mapping, drawing, self-testing, and self-explaining. The two are not the same thing, and conflating them has real [[pedagogy|pedagogical]] consequences: an AI *producing* a summary or a map for the student is the opposite of the student *performing* the generative-learning act. [[genai-mindtool-generative-learning|Dabbagh & Fake (2026)]] build directly on this distinction, arguing that a GenAI mindtool supports generative learning only when the *learner* drives the constructive activity — generating a mind map with AI assistance is generative learning; having the AI generate the map wholesale is not, however fluent or correct the output. The deciding question is **who performs the meaning-making**: - Does the student construct an explanation, or merely receive one? - Does AI prompt the learner to connect ideas, or supply the connections for them? - Is the artifact (summary, map, code, model) the *product* of the learner's construction, or a substitute for it? This mirrors the [[icap-framework|ICAP]] hierarchy — constructive and interactive engagement outrank active and passive — but sharpens it: a tool can produce visibly "constructive-looking" output while the learner sits in a *passive* or *active* mode. Evaluating a GenAI tool for generative learning therefore means inspecting where the constructive effort actually happens, not whether generative output is present. This is the same constructivist-in-name / behaviorist-in-practice trap, applied to the specific case of generation: [[ai-writes-code-student-writes-model-2026|model authorship]] (the AI writes the code, the student writes the model) is one concrete resolution — the learner constructs the *conceptual model* even when AI supplies the surface artifact. ### Design responses grounded in constructivism - **Generative Refusal** — AI tools that strategically withhold generated text and pose questions instead, returning [[desirable-difficulties|cognitive friction]] to the user so that the labor of articulation itself builds understanding.([[generative-refusal-ai-tools-for-thought]]) - **Thinking tools over answer machines** — using GenAI as a [[genai-mindtool-generative-learning]] in which the learner drives the tool, rather than the tool replacing the learner.([[genai-mindtool-generative-learning]]) - **Constructive conflict** — adversarial AI agents that challenge a learner's design or reasoning, prompting reconsideration and deeper construction of alternatives, in the tradition of Socratic tutoring.([[ai-agents-constructive-conflict-design-education-2026]]) - **Internal feedback via comparison** — having learners compare their own work against AI-generated exemplars so that the act of comparison itself generates learning.([[ai-internal-feedback-evaluative-judgments]]) - **Question-type-aware prompting** — classifying learner questions into constructivist roles so the system can deliberately escalate a student from information-seeking toward exploratory, dialogic inquiry instead of mirroring whatever cognitive depth the question implies. Because facilitator and co-learner intent remain confusable for automated classifiers, this design keeps a human validating the categorization before it drives [[feedback]] or [[scaffolding]].([[lee-learner-question-types-ai-education-2026]]) ## Constructivism and "education about AI" Constructivism also shapes how AI literacy itself is taught. If knowledge is constructed, then AI literacy is not acquired by lecturing about models but by actively using, critiquing, and building with AI — generating artifacts, interrogating outputs, and reflecting on the interaction.([[hingle-collaborative-ai-literacy-2025]]) This positions [[ai-literacy]] as an active, participatory competency rather than a body of passive knowledge, and it connects constructivism to [[critical-thinking]] and to [[agency]] in learners' encounters with AI. ## Implications for design and research 1. **Preserve the constructive activity.** AI should scaffold the learner's own thinking — prompt, question, and support — rather than perform it. Designers should ask whether the tool increases or replaces the learner's constructive effort.([[generative-refusal-ai-tools-for-thought]]) 2. **Use the [[icap-framework|ICAP]] lens.** ICAP classifies engagement into constructive, interactive, active, and passive modes — use it to evaluate whether AI interactions actually elicit constructive and interactive modes rather than passive consumption. Designers should favor the deeper (constructive and interactive) modes where the learning goal warrants.([[hingle-collaborative-ai-literacy-2025]]) 3. **Align theory and implementation.** Researchers should look beyond whether AI "works" to *how* it embodies a learning theory, checking for the constructivist-in-name, behaviorist-in-practice gap.([[ai-vocational-education-training-review]]) 4. **Study learner agency and transfer.** Constructivist commitments imply evaluating not just immediate test gains but whether learners can transfer and independently apply their constructed understanding.([[research-methods-aied]]) ## Connected Concepts - [[community-of-inquiry]] — Community of Inquiry (grounded in constructivist/Deweyan pragmatism) - [[cognitive-psychology]] — Cognitivism, the third classical pole of learning theory - [[active-learning]] - [[learning-by-teaching]] - [[scaffolding]] - [[self-regulated-learning]] - [[collaborative-learning]] - [[experiential-learning]] - [[project-based-learning]] - [[embodied-learning]] - [[learning-design]] - [[generative-ai]] - [[intelligent-tutoring]] - [[cognitive-offloading]] - [[agency]] - [[critical-thinking]] - [[ai-literacy]] - [[misconceptions]] - [[learning-theories]] - [[behaviorism]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[theory-development-aied]] — Theory Development in AI in Education - [[productive-failure]] ## Connected Articles - [[lee-learner-question-types-ai-education-2026]] — Learner questions classified into three constructivist roles: transmitter, facilitator, co-learner (Lee, Atif & Kang 2026) - [[mishra-control-vs-agency-history-2025]] — Positions constructionism (Papert) against cognitive tutors in AIED history - [[code-to-learn-genai-artifact-construction-2026]] — Code-to-Learn with GenAI: constructionism framework for artifact construction - [[ai-writes-code-student-writes-model-2026]] — Model authorship: theory and measurement program for learning-by-construction with GenAI - [[rewriting-curriculum-genai-pedagogy-2026]] — Rewriting the curriculum: GenAI-driven pedagogical change - [[zhu-e3-hot-embodied-intelligence-sustainable-learning]] — Fostering Sustainable Learning via Embodied Intelligence (E3-HOT) - [[ai-vocational-education-training-review]] — Constructivism espoused but behaviorist AI dominates in VET; the "Turing Trap" - [[generative-refusal-ai-tools-for-thought]] — AI tools that withhold generation to protect constructive thought - [[genai-mindtool-generative-learning]] — GenAI as a thinking tool supporting learner construction - [[ai-agents-constructive-conflict-design-education-2026]] — Adversarial AI agents prompting constructive reconsideration - [[hingle-collaborative-ai-literacy-2025]] — Collaborative AI literacy and the ICAP engagement framework - [[ai-internal-feedback-evaluative-judgments]] — AI-supported comparison generating evaluative judgments - [[icap-cognitive-engagement-llm-agents]] — ICAP and cognitive engagement with LLM agents - [[conversational-ai-tutors-framework]] — Scaffolding dialogue in AI tutors - [[embodied-inquiry-ai-facilitator-physics-2026]] — Embodied inquiry with an AI facilitator - [[beyond-detection-authentic-assessment-ai-2025]] — Authentic assessment and knowledge construction - [[teacher-ai-teaming-five-levels]] — Levels of teacher–AI collaboration in design - [[ccct-cooperative-learning-technique]] — Cooperative learning framed through constructivist theories - [[learning-with-machines-toward-a-theory-of-epistemic-co-agency]] — Epistemic co-agency between learner and machine - [[ensemble-cognition-philosophy-ai-education]] - [[vargas-situated-learning-ai-review-2024]] - [[li-ai-science-situated-learning-teachers-2025]] - [[ojeda-ramirez-community-based-ai-learning]] - [[vargas-ai-catalyst-situated-learning-2026]] - [[elsayed-pedagogical-symbiosis-posthuman-learner]] - [[niari-ai-pedagogical-mediator-collaborative-learning]] - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a Mediational Agent - [[context-based-ai-secondary-chemistry-2026]] — Context-based 7E + AI instruction in secondary chemistry - [[educational-robotics-pathways-2026]] — Pathways to Learning AI-Powered Educational Robotics (2026) - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[genai-integration-constructivist-higher-ed-bangladesh-2026]] — GenAI integration in Bangladeshi higher ed through constructivism (Alam et al. 2026) --- ## [Cognitive Psychology](https://edtechdev.github.io/aied/concepts/cognitive-psychology/) > **Cognitive psychology / cognitivism** — the family of theories that explain learning through internal mental processes — attention, perception, memory, reasoning, and metacognition — rather than through observable behavior alone. In [[ai-education|AI in education]], cognitivist assumptions underpin the field's most distinctive contributions: [[intelligent-tutoring]] systems that model learner knowledge, [[knowledge-tracing]] and [[cognitive-diagnosis]] that track what a learner knows, [[feedback]] designs grounded in error diagnosis, and the whole [[student-modeling|learner modeling and adaptive instruction]] family. Cognitivism is the middle ground between [[behaviorism]] (learning as behavioral change) and [[constructivist|constructivism]] (learning as active meaning-making), and it is the theoretical lens most closely tied to the computer metaphor of the mind that animated early AIED. ## Questions to Consider - When you think about 'learning,' do you picture a change in what someone does, or a change in what they know and can retrieve? How might that distinction change how you judge whether an AI tutoring tool actually works? - AI tutoring is built on a 'computer metaphor' — treating the mind as an information-processing system with memory limits. Where does that metaphor feel powerful, and where might it miss something important about how humans learn? - An AI tool makes a task feel effortless: it explains the next step, reduces friction, and the learner performs beautifully while using it. Does that count as successful [[teacher-role|teaching]]? How would you know whether the learner can now do it without the tool? - Cognitive Load Theory distinguishes intrinsic, extraneous, and germane load. If you were designing an AI assistant, which type of load would you deliberately try to reduce, and which would you be careful NOT to remove? - If a learner knows they can offload memory and reasoning to an AI, when is that a smart strategy and when is it a shortcut that quietly prevents learning? What determines the difference? - Cognitivism assumes knowledge can be broken into components and tracked over time. What might get lost when we reduce a learner's understanding to a set of traceable knowledge components? ## Introduction Cognitive psychology is the learning-theory tradition that treats learning as a change in internal mental representations — concepts, schemas and procedures held in memory — rather than as a change in observable behavior. Its information-processing vocabulary (attention, encoding, retrieval, bounded working memory) supplied both the diagnostic language for learning difficulty and the architecture behind [[intelligent-tutoring]] and [[knowledge-tracing]]: systems that infer a learner's internal state and adapt to it. It remains the reference frame for [[metacognition]], [[desirable-difficulties]] and [[self-regulated-learning]] across this knowledge base. ## Core ideas - **Learning is a change in internal mental representations.** Cognitivism holds that learning involves the acquisition, storage, and reorganization of knowledge in memory — concepts, schemas, and procedures — rather than just a change in observable response. What a learner *knows and can retrieve* matters, not just what they do. - **The information-processing (computer) metaphor.** The mind is treated as an information-processing system with capacities and bottlenecks — [[item-response-theory|measurement]] of latent ability, working-memory limits, encoding and retrieval — which is precisely the model that made AI tutoring (a computer program that models and adapts to learner cognition) a natural fit. - **Attention and memory are bounded.** Working memory has limited capacity; durable learning requires encoding into long-term memory through rehearsal, elaboration, and [[retrieval-spacing-interleaving|retrieval practice]]. This connects cognitivism to [[research-methods-aied|research]] on [[cognitive-offloading]] (delegating memory/processing to external tools) and to the "performance–learning gap" when AI bypasses retrieval and practice. - **Metacognition regulates cognition.** [[metacognition]] — monitoring and controlling one's own thinking — is a distinctly cognitivist construct, and it explains why learners' calibration of when to rely on AI matters for learning (see [[cognitive-offloading]] and [[self-regulated-learning]]). - **Knowledge is decomposable and traceable.** Cognitivist AIED assumes learner knowledge can be represented as components and tracked over time — the foundation of [[knowledge-tracing]], [[cognitive-diagnosis]], and [[item-response-theory]]. ## Cognitivism and AI in education ### The cognitivist lineage of AIED Cognitivism is arguably the theory most responsible for AI in education existing at all. The early cognitive tutors (e.g., Anderson's ACT-R-based tutors) [[embodied-learning|embodied]] the assumption that learning could be modeled as production rules and that a system could trace which rules a learner had mastered. This produced the canonical architecture that still defines the field: a domain model, a [[student-modeling|student model]] that tracks the learner's knowledge state, and a [[pedagogy|pedagogical]] model that adapts instruction — all cognitivist in origin. Modern [[knowledge-tracing]] (Bayesian, deep-learning, and IRT-based) and [[cognitive-diagnosis]] continue this tradition. The same assumption underwrites [[intelligent-tutoring]], [[adaptive-learning]], and [[personalized-learning]], which are grouped in the knowledge base under the [[student-modeling|Learner Modeling and Adaptive Instruction]] umbrella. ### Cognitive load and the design of instruction Cognitive Load Theory (CLT) is the most widely applied cognitivist framework in [[learning-design|instructional design]]: it distinguishes intrinsic load (task complexity), extraneous load (presentation friction), and germane load (schema-building effort). Well-designed AI should reduce extraneous load while preserving germane processing; poorly integrated AI reduces all three, leaving completed tasks with empty learning. CLT's working-memory framing is also central to debates about [[cognitive-offloading]] — whether AI reduces harmful extraneous load or short-circuits the germane processing that produces learning. ### Cognitivism vs. behaviorism and constructivism - **vs. [[behaviorism]]:** Behaviorism explains learning as observable behavioral change through reinforcement and drill; cognitivism insists on internal representations and traces mental states. AI practice often shows a "constructivism in name, behaviorism in practice" gap, but cognitivist designs (student modeling, knowledge tracing) are distinct from pure behaviorist drill-and-feedback because they *represent and adapt to the learner's inferred knowledge* rather than merely reinforcing responses. - **vs. [[constructivist|constructivism]]:** Constructivism holds that [[learners]] actively construct meaning through experience; cognitivism emphasizes accurate encoding of (often pre-structured) knowledge and skill. AIED's cognitivist lineage (structured domains, explicit knowledge components) is sometimes critiqued as too behaviorist or too transmission-oriented by constructivists, while cognitivism counters that representing and tracing knowledge is what enables genuinely adaptive instruction. - **vs. the [[learning-sciences|learning sciences]]:** Cognitivism supplies the mechanisms that field designs with — working memory, encoding, retrieval, decomposable knowledge components — but is not itself design-oriented. It explains how learning happens; the learning sciences ask how to build environments in which it happens and hold those designs to empirical test. ### The AI-era tension: cognitivism's boundary is under pressure [[generative-ai|Generative AI]] both extends and challenges cognitivism. It extends it by making knowledge representations more powerful (LLMs as knowledge engines that can be traced via [[knowledge-tracing]] and adapted via [[student-modeling]]). It challenges it by complicating where cognition "is": when AI performs reasoning, memory, and even metacognitive-like functions, the cognitivist assumption that learning is internal processing in the individual mind is unsettled — as [[distributed-cognition]], [[ai-cognitive-partner-co-regulation-learning|co-regulation]], and post-human framings argue cognition can be distributed across human and artificial systems. Yet the cognitivist question remains the field's central one: *does the learner internalize the knowledge, or does the tool hold it?* This is the cognitive-offloading and performance–learning gap question in its purest form. ## Implications for design and research 1. **Design for internalization, not just performance.** Cognitivist AIED should be evaluated on whether the learner can retrieve and apply knowledge *without* the tool — not on assisted performance. This is the [[ai-misuse-learning-harm|performance–learning gap]] and the rationale for measuring unassisted [[transfer-of-learning|transfer]]. 2. **Represent the learner, don't just respond.** Attach structured [[student-modeling]] and [[knowledge-tracing]] to AI dialogue so the system adapts to inferred knowledge rather than responding fluently but blindly.([[educlaw-bench-pedagogical-llm-agents-2026]]) 3. **Respect working-memory limits.** Apply Cognitive Load Theory to AI UX: reduce extraneous load (friction, overloaded interfaces) while preserving germane processing ([[desirable-difficulties|productive struggle]], retrieval practice) rather than minimizing all cognitive demand. 4. **Calibrate metacognition.** Because [[metacognition]] governs when learners choose to offload, teaching calibration (knowing what one can actually do unaided) is a cognitivist answer to over-reliance (see [[cognitive-offloading]]). ## Connected Concepts - [[behaviorism]] - [[constructivist]] - [[learning-theories]] - [[metacognition]] - [[cognitive-offloading]] - [[knowledge-tracing]] - [[cognitive-diagnosis]] - [[student-modeling]] - [[intelligent-tutoring]] - [[adaptive-learning]] - [[personalized-learning]] - [[item-response-theory]] - [[distributed-cognition]] - [[icap-framework]] - [[transfer-of-learning]] - [[self-regulated-learning]] - [[ai-education]] - [[learning-sciences]] - [[retrieval-spacing-interleaving]] — the retention findings this practice family rests on ## Connected Articles - [[cognitive-shift-ai-education]] — The cognitive shift in AI education - [[cogtax-cognitive-taxonomy]] — A cognitive taxonomy for AI use - [[educlaw-bench-pedagogical-llm-agents-2026]] — Pedagogical LLM agents grounded in knowledge tracing - [[nie-personavlm-long-term-personalization-2026]] — LLM student modeling and memory - [[ai-cognitive-partner-co-regulation-learning]] — AI as a cognitive partner in co-regulated learning - [[ensemble-cognition-philosophy-ai-education]] — Ensemble Cognition: thinking as human–AI interaction --- ## [Sociocultural Learning](https://edtechdev.github.io/aied/concepts/sociocultural-learning/) > **Sociocultural learning** — the family of theories, rooted in Vygotsky, that holds learning and development arise through social participation and are mediated by cultural tools, language, and interaction with more knowledgeable others. Cognition is distributed across people, artifacts, and environments rather than residing solely in individuals. In [[ai-education|AI in education]], sociocultural theory frames how [[generative-ai|generative AI]] functions as a new kind of *mediational agent* — a tool that both mediates activity and generates contingent contributions to interaction — and frames the design of [[scaffolding]], the Zone of Proximal Development (ZPD), apprenticeship, and communities of practice. See [[generative-ai-mediational-agent-sociocultural-2026]]. ## Questions to Consider - Vygotsky's Zone of Proximal Development is the gap between what you can do alone and what you can do with help. Recall a time a well-timed hint let you accomplish something you couldn't alone — what made that help effective, and when might it have given too much away? - The page claims human thinking is 'mediated' by cultural tools like language and writing that reorganize how we reason. If that's true, how should we think about an AI [[conversational-ai|chatbot]] as a new kind of thinking tool — and how might it change what 'knowing' means? - Sociocultural theory says cognition is distributed across people, artifacts, and environments rather than inside individual heads. Does that match your experience of how you actually get things done, and what would it mean for designing learning if it's right? - If learning happens first between people and only later within the individual, what are the risks of an AI tutor that lets a student interact mostly with a machine rather than with peers or a more knowledgeable human? - How would you decide how much support an AI tutor should give so that a learner advances without the answer simply being handed over? ## Introduction ### The concept Sociocultural theory (Vygotsky, 1978; Luria; Leontiev) holds that higher mental functions develop through participation in culturally organized activity. Unlike accounts that locate learning solely in the individual's information processing, the sociocultural view emphasizes that: - **Mediation is fundamental.** Humans think with and through cultural tools — language, writing, diagrams, [[ai-technologies|technologies]] — which reorganize how they reason, remember, and solve problems (Wertsch, 1991). These tools do not merely transmit information; they reshape cognition and participation. - **Learning is social.** Higher mental functions appear first between people (intersubjectively, in interaction) and only later within the individual. Learning arises through participation with [[teacher-role|teachers]], peers, and communities — in processes like [[scaffolding]], apprenticeship, and movement through the ZPD. - **[[distributed-cognition|Cognition is distributed]].** Cognitive work is spread across people, artifacts, and environments (Hutchins; Clark & Chalmers; Pea), rather than contained in the individual mind. [[distributed-cognition|Distributed cognition]], [[situated-learning|situated learning]], and communities of practice extend the sociocultural strand. ### The Zone of Proximal Development (ZPD) The ZPD (Vygotsky) is the sociocultural concept most widely applied in [[intelligent-tutoring|AI tutoring]]: the space between what a learner can do alone and what they can do with assistance. Learning happens most effectively when instruction targets this zone — challenging enough to push development, supported enough to make progress. It is the theoretical foundation of [[scaffolding]]: temporary, adjustable support withdrawn as competence grows. In AI in education, ZPD frames the central design question of how much support an [[intelligent-tutoring|AI tutor]] should provide so learning advances without being given away — see [[stanford-evidence-base-ai-k12-2026]]. ### Sociocultural learning in AI education Sociocultural theory shapes AIED [[research-methods-aied|research]] in several distinct ways: - **AI as a mediational agent.** Generative AI complicates the sociocultural distinction between mediational means and social interaction: it both mediates activity *and* generates context-sensitive, contingent contributions that shape interaction, without possessing intentionality, social membership, or accountability. Warschauer, Tate, and Ritchie (2026) propose the *mediational agent* as a hybrid category, and derive human-first habits of participation (primacy of human cognition, purposeful [[student-engagement|engagement]], supervisory agency, epistemic vigilance, reflective [[self-regulated-learning|self-regulation]]) to preserve [[agency|learner agency]].([[generative-ai-mediational-agent-sociocultural-2026]]) - **ZPD-calibrated scaffolding.** [[intelligent-tutoring|AI tutors]] should dynamically calibrate help to sit within each learner's zone. [[stanford-evidence-base-ai-k12-2026]] shows how tutors tuned to a learner's level outperform generic assistance; [[adaptive-learning]] and [[golrang-propact-pair-programming-2026]] operationalize ZPD by adjusting difficulty and hints; and principled frameworks like [[finkelstein-principled-ai-education-2025]] argue support should be withdrawn as competence grows. - **A fourth zone: what the model knows.** [[scan-framework-task-assignment-generative-ai-2025|Tsim and Gutoreva (2025)]] extend Vygotsky's diagram rather than the tutoring loop, adding a *known to [[generative-ai|GenAI]]* zone to the ZPD and reading off four sub-zones that classify what a task should be assigned to: Substitute (no task-specific knowledge, so the model's general competence carries it), Aid (partial knowledge, augmented), Complement (enough knowledge to supervise the model's output), and Non-negotiable (enough to do it unaided, so delegation adds little). [[scaffolding|Scaffolding]] is expressed as task assignment rather than hint delivery, and a [[metacognition|metacognitive]] loop of real-time evaluation, reflection and learning is what moves a task between sub-zones over time. - **Apprenticeship and community.** Sociocultural ideas underpin cognitive apprenticeship, modeling, coaching, and fading; communities of practice frame learning as movement toward fuller participation in a community's practices. - **Cultural and [[governance|institutional]] context.** The [[constructivist|constructivism]]-adjacent sociocultural strand stresses that the cultural dimension shapes what counts as knowing, who is an authority, and what effort means — see the [[young-people-learning-generative-ai-rapid-review-2026|Sydney PreK-12 rapid review's]] learners–contexts–cultures framing. ### Connection to cognitive load and metacognition The sociocultural strand is tightly coupled to [[cognitive-offloading|Cognitive Load]] Theory (support should manage load without eliminating productive effort) and to [[metacognition]] (learners in the zone are actively monitoring and regulating their understanding). [[stanford-evidence-base-ai-k12-2026]] synthesizes [[k-12|K-12]] evidence that AI tools work best when they keep learners in the ZPD rather than answering for them, and [[human-in-the-loop-ai]] research addresses how human and AI support jointly define the learner's zone. - **Generative AI as a mediational agent (2026):** Drawing on Vygotskian mediation, a theory paper proposes reframing generative AI not merely as a tool/mediational means but as a *mediational agent* that actively participates in learning activity, blurring the tool-vs-social-interaction boundary central to sociocultural theory ([[generative-ai-mediational-agent-sociocultural-2026]]). This positions generative models as co-participants rather than passive instruments, with implications for how mediation, [[agency]], and the learner–AI relationship are theorized in [[learning-sciences|the learning sciences]]. ## Connected Concepts - [[scaffolding]] - [[constructivist]] - [[learning-theories]] - [[activity-theory-aied]] - [[situated-learning]] - [[distributed-cognition]] - [[metacognition]] - [[agency]] - [[generative-ai]] - [[human-ai-collaboration]] - [[desirable-difficulties]] - [[adaptive-learning]] - [[human-in-the-loop-ai]] - [[k-12]] - [[intelligent-tutoring]] ## Connected Articles - [[ai-teammate-task-distribution-medical-training-2026]] — SCAN framework: rethinking AI task distribution in medical training (Tsim et al. 2026) - [[youth-enter-chat-llm-student-talk-2026]] — When Youth Enter The Chat: Validation of LLM-Based Measures of Student Talk - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a Mediational Agent - [[golrang-propact-pair-programming-2026]] — Collaborative AI tutoring - [[finkelstein-principled-ai-education-2025]] — Principled AI education frameworks - [[stanford-evidence-base-ai-k12-2026]] — Stanford evidence base for AI in K-12 - [[text-simplification-its]] — Text simplification in ITS - [[young-people-learning-generative-ai-rapid-review-2026]] — Sydney rapid review of GenAI in PreK-12 - [[ai-cognitive-partner-co-regulation-learning]] — AI as cognitive partner and co-regulation - [[trikonet-trivalence-co-creativity-2026]] — TriKoNet: co-creativity as network effect via Actor-Network Theory --- ## [Distributed Cognition](https://edtechdev.github.io/aied/concepts/distributed-cognition/) > **Distributed Cognition** — the theoretical perspective that cognition is not confined to an individual mind but is distributed across people, tools, artifacts, and environments. In AI-in-education [[research-methods-aied|research]], this framework has become central for understanding [[human-ai-collaboration|human–AI collaboration]]: rather than viewing AI as a neutral tool that supports an otherwise self-contained learner, distributed cognition treats thinking as emerging from the interplay between learners, AI systems, peers, and their shared context. It reframes questions of [[agency]], responsibility, and [[learning-gains|learning outcomes]] in terms of how cognitive work is apportioned across a human–AI system. ## Questions to Consider - Distributed cognition says thinking isn't confined to a single mind but is spread across people, tools, and environments. When you use a calculator, a notes app, or an AI assistant, where does the 'thinking' actually happen? - If cognition is distributed across a human and an AI, who is responsible when the result is wrong — and who is accountable for learning? - One framework distinguishes AI's 'functional agency' (it can influence outcomes) from 'moral responsibility' (which remains human). Does that distinction hold up in practice, or does responsibility blur when humans can't understand what the system did? - Human-AI systems that distribute reasoning most efficiently tend to produce the highest task performance but the least [[self-regulated-learning|self-regulated learning]]. Why would making a system 'smarter' at the group level make the individual learner weaker? - The 'Cognitive Commons' argument holds that distributed mastery depends on internalized expertise — you can't effectively oversee an AI system you don't deeply understand. How does that challenge the idea that AI lets us skip the hard work of learning? ## Introduction This is a learning-theory concept within the knowledge base's [[learning-theories]] strand, closely related to [[embodied-learning]], [[situated-learning]], and [[cognitive-offloading]]. Distributed cognition (DCog), originating in the work of Edwin Hutchins and colleagues, describes how cognitive processes such as memory, reasoning, and [[problem-solving]] are spread across multiple agents and material systems rather than residing in a single head. In [[ai-education|AI education]] this matters because AI systems increasingly function as genuine cognitive partners that carry part of the thinking load, raising the question of *where* learning actually happens and *who* is responsible. ### How AI shifts the distribution of cognition Generative and interactive AI systems redistribute cognitive work in ways earlier tools did not. The knowledge base documents several dimensions of this shift: - **AI as a cognitive partner and co-regulator.** AI can act as a [[scaffolding|scaffold]], metacognitive support, external memory system, and decision partner in the co-[[regulation]] of thinking and learning.([[ai-cognitive-partner-co-regulation-learning]]) This positions cognition as co-regulated between learner and system rather than individually managed. - **Agency and responsibility redistribution.** Frameworks such as **ensemble cognition** reconceptualize thinking as emerging from dynamic human–[[student-ai-interaction|AI interaction]], distinguishing AI's *functional agency* (its capacity to influence outcomes without consciousness) from *moral responsibility* (which remains human).([[ensemble-cognition-philosophy-ai-education]]) Distributed cognition thus reframes who is accountable for learning. - **The efficiency–regulation tension.** Empirical work shows a trade-off: human–AI systems that distribute reasoning most efficiently (e.g., via delegated reasoning) achieve higher task performance but may reduce learners' self-regulatory [[student-engagement|engagement]].([[hao-human-ai-collaborative-problem-solving-cognition]]) This is the empirical face of the [[cognitive-offloading]] and [[cognitive-offloading|Over-Reliance]] risks the knowledge base documents. - **Mediation in collaboration.** AI can act as a *[[pedagogy|pedagogical]] mediator* that orchestrates interaction, epistemic sense-making, and regulatory processes in [[collaborative-learning|collaborative learning]], redistributing agency, authority, and responsibility across human and non-human actors.([[niari-ai-pedagogical-mediator-collaborative-learning]]) - **Access configuration distributes cognition within the group.** [[xu-genai-collaborative-space-2026|Xu et al. (2026)]] show that *how* a team shares GenAI determines the distribution of cognition: synchronous work on a single shared interface sustains a common cognitive model (collective prompts, shared external memory), whereas asynchronous private use fragments it, with outputs selectively re-labeled before sharing. GenAI thereby functions as both a distributed cognitive participant and an interactive collaborative space whose permeability must be designed (shared context flows in, private insights do not auto-flow back). - **Teachers too experience AI-dominant versus complementary distribution.** The same efficiency–regulation logic applies to the teacher side of lesson design: [[choi-teacher-ai-interaction-lesson-design-2026|Choi et al. (2026)]] found that novice teachers delegate a large share of instructional-design cognitive load to the AI system (an AI-dominant distribution, largely accepting responses), whereas experienced, AI-proficient teachers reach a complementary distribution in which pedagogical expertise and AI's computational support mutually reinforce — and because [[generative-ai|GenAI]] generates and co-constructs rather than merely stores information, they frame this as *participatory shared cognition*, not mere tool use. ### Distributed cognition and related perspectives The knowledge base treats distributed cognition alongside its neighboring theoretical traditions, which overlap but are not identical: - **[[situated-learning|Situated learning / situated cognition]]** emphasizes that cognition and learning are inseparable from the authentic context and communities of practice in which they occur — a complement to DCog's focus on cognitive *distribution* across systems. - **[[embodied-learning|Embodied cognition]]** stresses the role of the body and sensorimotor interaction, arguing that thinking is grounded in bodily engagement rather than abstract symbol manipulation. - **[[cognitive-offloading]]** is the practical mechanism by which cognitive work is handed off to external systems (including AI), and is the risk-laden counterpart to DCog's descriptive account. - **Extended cognition / the extended mind thesis** (used in posthumanist work) holds that external artifacts can be *constitutive* parts of a cognitive system, not merely instruments — a stronger claim that AI is part of the learner's mind itself.([[elsayed-pedagogical-symbiosis-posthuman-learner]]) ### Why it matters for AI design and evaluation Distributed cognition provides both a design lens and an evaluation lens. For design, it asks how to apportion cognitive work between learners and AI to preserve (not erode) the human learner's agency, [[metacognition]], and self-regulation. For evaluation, it reframes success metrics: instead of asking only "did performance improve?", DCog asks whether the distribution of cognition supports durable learning, epistemic agency, and educational justice — a perspective that connects to the knowledge base's [[ai-ed-evaluation]] and [[learning-theories]] concerns. **Internalized vs. distributed mastery.** The Cognitive Commons framework ([[cognitive-commons-ai-expertise-regeneration|Lovett 2026]]) distinguishes Internalized Mastery (deep domain knowledge in individual minds) from Distributed Mastery (orchestrating human–AI systems) and argues the latter depends on the former via a "Validation Tether": effective oversight of distributed/AI systems presupposes the internalized expertise those systems may undermine. This sharpens the DCog design question — the distribution of cognition must not come at the cost of the expertise that validates it. ## Connected Concepts - [[learning-theories]] - [[cognitive-offloading]] - [[embodied-learning]] - [[situated-learning]] - [[human-ai-collaboration]] - [[agency]] - [[metacognition]] - [[self-regulated-learning]] - [[constructivist]] - [[collaborative-learning]] - [[ai-education]] ## Connected Articles - [[hao-human-ai-collaborative-problem-solving-cognition]] — Empirical study of human–AI collaborative problem-solving through a distributed cognition + co-regulation lens - [[niari-ai-pedagogical-mediator-collaborative-learning]] — AI as a pedagogical mediator redistributing cognition across human and non-human actors - [[ai-cognitive-partner-co-regulation-learning]] — AI as a cognitive partner in co-regulated thinking - [[ensemble-cognition-philosophy-ai-education]] — Ensemble Cognition framework reconceptualizing thinking as human–AI interaction - [[elsayed-pedagogical-symbiosis-posthuman-learner]] — Posthuman learner with cognition distributed across biological and artificial systems - [[fowlin-operationalizing-learning-principles-ai]] — Operationalizing distributed cognition alongside experiential and situated learning - [[learning-with-machines-toward-a-theory-of-epistemic-co-agency]] — Epistemic co-agency as a distributed-cognition-inspired theory of learning with machines - [[cognitive-commons-ai-expertise-regeneration]] — The tragedy of the cognitive commons: AI and expertise regeneration - [[lodge-adaptive-capabilities-genai-future-2026]] — Adaptive capabilities for assuring quality learning in a gen AI-integrated future (Lodge et al. 2026) - [[choi-teacher-ai-interaction-lesson-design-2026]] — Teacher-AI interaction in lesson design: AI-dominant vs complementary distributed cognition by experience and proficiency (Choi et al. 2026) - [[generative-ai-mediational-agent-sociocultural-2026]] — Generative AI as a mediational agent - [[xu-genai-collaborative-space-2026]] — GenAI as agent and collaborative space: how access configuration distributes group cognition (Xu et al. 2026) --- ## [Situated Learning](https://edtechdev.github.io/aied/concepts/situated-learning/) > **Situated Learning** — the theory, rooted in the early-1990s work of Lave and Wenger (1991), that learning is not an isolated, decontextualized act but occurs through participation in authentic activities, contexts, and cultures. Knowledge is co-constructed by learners and peers within [[collaborative-learning|communities of practice]], and novices learn through legitimate peripheral participation — absorbing the culture, language, and practices of expert members as they move from the periphery to the center of a community. Emphasis falls on learning by doing in real-world situations, where [[assessment]] emerges from the task itself rather than being separated from it. ## Questions to Consider - Think of something you genuinely know how to do well — a skill, a trade, a craft. Did you learn it mostly from abstract instruction or from participating in a real community that did that thing? What does that suggest about where knowledge actually lives? - Lave and Wenger describe novices learning through 'legitimate peripheral participation' — starting at the edge of a community of practice and moving inward. Where have you watched (or been) such a newcomer, and what let them move from the periphery to the center? - If knowledge is 'situated' in authentic contexts, what does that imply for the traditional classroom, which deliberately separates learning from real-world situations — and for AI tools designed to deliver decontextualized content? - The page treats situated learning as a design lens for AI. How might an adaptive system or simulation ground learning in authentic practice rather than pulling it out of context? - What obstacles does the [[research-methods-aied|research]] say stand in the way of situating AI-driven learning in real contexts, and which of those have you seen in your own institution? ## Introduction Situated learning is one of the activity-and-context theories within the knowledge base's [[learning-theories]] strand. It takes up Vygotskian themes of social construction but adds a strong emphasis on the intimate integration of "doing" and "learning" and on the importance of communities of practice. As an educational stance it confronts traditional, standardized schooling by foregrounding the learner's [[sociocultural-learning|sociocultural]] context as a key element for acquiring skills and appropriating knowledge relevant to their reality. In the [[ai-education|AI-in-education]] literature, situated learning matters because it provides a design lens for AI: [[adaptive-learning|adaptive systems]], [[intelligent-tutoring|intelligent tutoring]] in authentic scenarios, and [[virtual-and-augmented-reality|immersive]] [[simulation|simulations]] can ground AI-driven education in real-world contexts, while situated learning in turn offers AI a meaningful anchor in authentic practice and complexity. The two are widely treated as complementary, with human guidance remaining essential for [[ethics|ethical]] grounding. ### Situated learning as a design lens for AI The knowledge base's research treats situated learning not merely as an abstract theory but as a concrete design and evaluation framework for AI in education: - **Opportunities and obstacles.** A PRISMA [[meta-analysis-systematic-review|systematic review]] of 60 articles (three decades) finds that AI can augment situated learning — through adaptive systems tailored to students' evolving needs, intelligent tutoring situated in authentic scenarios, automation of administrative tasks, and data-driven teacher support — while the main obstacles are the traditional school's one-way passive learning, an over-emphasis on predefined outcomes, and teachers' limited contextual knowledge. Human guidance remains essential for ethical grounding.([[vargas-situated-learning-ai-review-2024]]) - **AI as a catalyst connecting education to reality.** AI can act as a *catalyst* for situated learning by connecting education with reality and authentic contexts, enabling learning grounded in real-world scenarios.([[vargas-ai-catalyst-situated-learning-2026]]) - **Mediational artifacts in authentic inquiry.** In [[science-education|science learning]], AI tools (virtual labs, simulations, intelligent tutoring) function as "mediational artifacts" that extend situated learning by enabling digital communities of practice and boundary-crossing between school, real-world, and interdisciplinary contexts — transforming students from "knowledge learners" into "scientific practitioners."([[li-ai-science-situated-learning-teachers-2025]]) - **Situated [[curriculum-design|curriculum]] devices.** [[ai-literacy|AI literacy]] can be developed through *Episodes of Situated Learning* — active [[teacher-role|teaching]] instruments (anticipate, produce, reflect) that build AI competencies through real-world [[problem-solving]] rather than abstract instruction.([[panciroli-ai-literacy-episodes-situated-learning]]) - **Situated evaluation in [[design-based-research|design-based]] learning.** [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] grounded their study in situated-learning theory and iterative design [[pedagogy]], evaluating 80 student design posters across instructor, peer-reviewer, and grant-reviewer roles. Role-aware [[prompt-engineering|prompting]] produced qualitatively different evaluative feedback — instructors encouraging and process-oriented, peers supportive and conversational, grant reviewers formal and outcomes-oriented — differences that were epistemic, not merely stylistic, foregrounding different aspects of design practice. This shows how situated, role-specific evaluation can be emulated by an [[llm|LLM]] when scaffolded with a semantically precise rubric, and how assessment in design-based learning emerges from the authentic task and its roles rather than being separated from them. - **Situated AI ethics.** Ethical reasoning about AI is itself best treated as *situated* — grounded in cultural-historical and ecological context rather than abstract principles.([[raffaghelli-situated-ai-ethics-2026]]) Situated learning connects closely to [[embodied-learning]] (both stress the grounding of cognition in context and action), [[distributed-cognition]] (learning distributed across people, tools, and contexts), [[experiential-learning]], and [[constructivist]] theory. In AI education it grounds the critique of decontextualized, disembodied learning: AI design that keeps learners anchored in authentic practice preserves the situatedness that durable learning requires. ## Connected Concepts - [[learner-identity]] — evolving disciplinary, professional, creative, and academic learner identities - [[learning-theories]] - [[constructivist]] - [[experiential-learning]] - [[embodied-learning]] - [[distributed-cognition]] - [[collaborative-learning]] - [[adaptive-learning]] - [[personalized-learning]] - [[teacher-role]] - [[learning-design]] - [[ai-education]] - [[virtual-and-augmented-reality]] — immersive environments as a route to authentic context ## Connected Articles - [[vargas-situated-learning-ai-review-2024]] — PRISMA systematic review of situated learning and AI in education (primary reference for this stub) - [[genai-educational-outcomes-meta-analysis]] - [[li-ai-science-situated-learning-teachers-2025]] - [[raffaghelli-situated-ai-ethics-2026]] - [[vargas-ai-catalyst-situated-learning-2026]] - [[panciroli-ai-literacy-episodes-situated-learning]] - [[fowlin-operationalizing-learning-principles-ai]] - [[videla-embodied-ai-education-choreography]] - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design --- ## [Embodied Learning](https://edtechdev.github.io/aied/concepts/embodied-learning/) > **Embodied learning** — the [[pedagogy|pedagogical]] principle that learning is grounded in bodily experience, physical interaction, and the sensory-motor context of the learner. Embodied approaches hold that cognition is not purely abstract but shaped by the body and its interaction with the environment. In [[ai-education|AI in education]], embodiment is realized through [[educational-robotics|educational robots]] and [[educational-robotics|social robots]], whose physical presence grounds abstract concepts (such as program logic or social skills) in observable, manipulable behavior. ## Questions to Consider - We often think of learning as something that happens 'in the head,' with the body just carrying the brain around. What if learning is actually grounded in bodily experience and interaction with the environment? What's one subject you learned that seemed to require your body — and could it have been learned purely abstractly? - Embodied approaches claim a physical, manipulable agent helps learners connect abstract ideas to concrete outcomes — seeing a program make a robot move, for instance. When have you noticed that doing something physical made an abstract concept finally 'click'? - Some [[research-methods-aied|researchers]] treat gesture as evidence of understanding — tracking students' hand movements alongside their speech to assess conceptual grasp. If a student's hands 'know' the concept before their words do, what might that imply about how we should assess learning? - An emerging critique challenges 'disembodied' AI that operates on abstract symbols, arguing AI should be designed around embodied intelligence to sustain learners' thinking rather than outsourcing it. Do you think an AI that has never had a body can fully support embodied learning? ## Introduction Embodied learning is closely related to [[active-learning]], [[experiential-learning]], and situated/[[constructivist]] theories. The key claim is that a physical, manipulable agent helps learners connect abstract ideas to concrete outcomes — a program that makes a robot move, or a role-play with a physical robot — in ways that pure screen-based interaction may not. Robotics is the clearest embodiment of AI in education, giving learners something to see, touch, and observe. ### How embodied learning appears in the knowledge base's research - **Grounded programming:** [[roboblockly-conversational-block-robotics-ct-2026|RoboBlockly Studio]] grounds [[cs-education|block programming]] in embodied robot execution, creating a tight loop of authoring, running, observing, and revising so learners see their code become behavior. - **Social-robotic interaction:** [[educational-robotics|Social robots]] used for [[storytelling-in-education|storytelling]] ([[motibo-digital-storytelling-robots-motivation-2026|MotiBo]], [[robobuddy-llm-social-robots-classroom-2025|RoboBuddy]]), role-play ([[remind-robot-mediated-roleplay-antibullying-2026|REMind]]), and sign language ([[pepper-robot-sign-language-lis-2025|Pepper]]) provide embodied social interaction that supports relational and [[social-emotional-learning|emotional]] learning. - **Embodiment and creative writing:** [[enhancing-creative-writing-with-robot-llm-integration-the-interplay-of-embodimen|Research on robot-LLM integration in creative writing]] examines how embodiment affects learners' interaction and outcomes. - **Human-robot interaction:** [[educational-robotics|HRI]] research ([[task-context-trust-educational-hri-2026|trust]], [[human-autonomy-agency-hri-review-2025|agency]]) examines how physical embodiment shapes trust, [[student-engagement|engagement]], and [[agency|autonomy]]. - **Gesture as evidence of understanding:** [[multimodal-embodied-cognition-oral-explanations-2026|Morphew et al.]] integrate computer-vision gesture tracking with [[llm]] analysis of speech to show that engineering students' conceptual understanding of statistics is expressed through both speech and gesture. High-confidence explanatory gestures cluster around specific concepts (especially the mean), and close gesture–speech coupling signals coherent conceptual talk while divergence marks developing ideas — positioning embodied action as evidence in [[assessment-validity|assessment]] via [[multimodal|multimodal learning analytics]], not only as a learning mechanism. ### Embodied intelligence and the critique of disembodied AI A second, more theoretical strand of the knowledge base's embodiment research concerns the role of the *body* in AI-[[sociocultural-learning|mediated learning]] — not through physical robots, but through the question of whether [[ai-technologies|AI systems]] themselves are (or can be) embodied. This work challenges the dominance of symbolic, disembodied AI models built on abstract information processing: - **Embodied AI as a design principle.** The **E3-HOT framework** argues that to sustain learners' cognitive agency and [[critical-thinking|higher-order thinking]] (rather than encouraging [[cognitive-offloading|cognitive outsourcing]]), AI should be designed around *embodied intelligence* — situational embedding, embodied participation, and cognitive creation — within a virtual–real integrated learning environment.([[zhu-e3-hot-embodied-intelligence-sustainable-learning]]) - **The limits of disembodied [[generative-ai|generative AI]].** Post-cognitivist scholarship argues that current GenAI systems lack proprioception, multimodal agency, and embodied practice, and advocates an "embodied AI" grounded in situationality, emergence, and sensorimotor coupling, proposing a perceptual–[[affective-computing|affective]] choreography for human–[[student-ai-interaction|AI interaction]].([[videla-embodied-ai-education-choreography]]) - **Embodiment, situatedness, and social construction.** In [[science-education|science learning]], AI tools function as "mediational artifacts" that enable digital communities of practice and boundary-crossing, connecting embodied, authentic inquiry to real-world and interdisciplinary contexts.([[li-ai-science-situated-learning-teachers-2025]]) This ties embodiment to [[situated-learning]] and [[distributed-cognition]]. Embodied learning connects to [[educational-robotics]], [[educational-robotics]], [[educational-robotics]], [[active-learning]], [[experiential-learning]], [[situated-learning]], [[distributed-cognition]], [[computational-thinking]], and [[social-emotional-learning]]. ## Connected Concepts - [[educational-robotics]] - [[active-learning]] - [[experiential-learning]] - [[situated-learning]] - [[distributed-cognition]] - [[computational-thinking]] - [[social-emotional-learning]] - [[learning-theories]] - [[multimodal]] - [[assessment-validity]] - [[virtual-and-augmented-reality]] — the modality that tries to exploit embodiment directly ## Connected Articles - [[multimodal-embodied-cognition-oral-explanations-2026]] — A Multimodal Framework for Embodied Cognition in Oral Explanations - [[zhu-e3-hot-embodied-intelligence-sustainable-learning]] — Fostering Sustainable Learning via Embodied Intelligence (E3-HOT) - [[roboblockly-conversational-block-robotics-ct-2026]] — RoboBlockly Studio - [[motibo-digital-storytelling-robots-motivation-2026]] — MotiBo - [[remind-robot-mediated-roleplay-antibullying-2026]] — REMind - [[pepper-robot-sign-language-lis-2025]] — Pepper and Sign Language - [[enhancing-creative-writing-with-robot-llm-integration-the-interplay-of-embodimen]] — Robot-LLM Integration in Creative Writing - [[robot-assisted-language-learning-meta-analysis-2026]] — Meta-analysis of AI-enhanced embodied robot-assisted language learning - [[white-wu-robotics-ai-education-2026]] — Robotics and AI in Education - [[ensemble-cognition-philosophy-ai-education]] - [[vargas-situated-learning-ai-review-2024]] - [[li-ai-science-situated-learning-teachers-2025]] - [[vargas-ai-catalyst-situated-learning-2026]] - [[elsayed-pedagogical-symbiosis-posthuman-learner]] - [[fowlin-operationalizing-learning-principles-ai]] - [[videla-embodied-ai-education-choreography]] - [[play-ai-pre-k-kindergarten-ai-literacy-2026]] — Play With AI (PL-AI): play-centered AI literacy curriculum for pre-K and kindergarten (Lee 2026) --- ## [Community of Inquiry](https://edtechdev.github.io/aied/concepts/community-of-inquiry/) > **Community of Inquiry (CoI)** is a framework for conceptualizing a meaningful educational experience as the dynamic interplay of **cognitive presence**, **social presence**, and **teaching presence**. Originating in computer-mediated and online learning [[research-methods-aied|research]] (Garrison, Anderson & Archer, 2000), it has become one of the most widely used models for designing, evaluating, and researching [[online-teaching-and-learning|online and blended]] inquiry-based education. In the [[generative-ai]] era the framework is being reconceptualized: machine-produced discourse can mimic authentic presence, so presences must be understood as sociotechnical accomplishments of human–GenAI assemblages rather than purely human activity. ## Questions to Consider - In an online course, what tells you that meaningful learning is actually happening — and how would you know if it weren't? - The Community of Inquiry framework names three presences: cognitive, social, and teaching. In your own online experiences, which presence most often goes missing — and what does that do to learning? - If an AI [[conversational-ai|chatbot]] writes fluent, empathetic-sounding replies in a discussion forum, is 'social presence' happening? Or is presence something only humans can genuinely create? - Generative AI can produce explanations so coherent they feel obviously right — but coherence isn't the same as correctness. Where have you seen 'fluency mistaken for warrant,' and how would you guard against it in an inquiry course? - This framework suggests presences are 'sociotechnical accomplishments' of human-and-AI working together rather than purely human activity. Does that reframe who you hold accountable when an online discussion goes shallow or goes well? ## Introduction The Community of Inquiry framework describes a worthwhile online learning experience as the product of three interacting presences: cognitive presence (learners constructing and confirming meaning through sustained reflection and discourse), social presence (participants projecting themselves socially and emotionally), and teaching presence (design, facilitation and direction of the process). Grounded in [[constructivist]] and Deweyan traditions, it treats inquiry as social and iterative, and it now serves as a design and evaluation lens for AI-mediated courses — where a [[pedagogical-agent]] may support one presence while quietly eroding another. ## The three presences - **Cognitive presence** — the extent to which learners construct and confirm meaning through sustained reflection and discourse, operationalized via the *practical inquiry* model (triggering event → exploration → integration → resolution). - **Social presence** — the ability of participants to project themselves socially and emotionally, expressed through [[affective-computing|affective]] communication, open communication, and group cohesion. - **Teaching presence** — the design, facilitation, and direction of cognitive and social processes to realize personally meaningful and educationally worthwhile outcomes, including design/organization, facilitating discourse, and direct instruction. CoI is grounded in [[constructivist]] and Deweyan pragmatic traditions: inquiry is social, iterative, and driven by a felt difficulty that motivates the search for resolution. ## CoI as a framework for online teaching CoI originated in, and remains most strongly associated with, [[online-teaching-and-learning]]. It provides a vocabulary for diagnosing *why* an online course works or fails: low [[student-engagement|engagement]] and isolation in asynchronous courses are usually failures of social and teaching presence, while surface discussion often reflects weak cognitive presence. This makes CoI a practical design lens for the very conditions the online medium creates — the removal of physical co-presence, the need for deliberate community-building, and the structuring of discussion that substitutes for face-to-face contact. In the [[generative-ai]] era, CoI is also where online instructors confront the hardest new questions: who is "present" when [[llm]] agents post, moderate, or respond, and how to keep the three presences meaningful when machine-generated discourse can mimic them. The knowledge base's online-teaching page therefore treats Community of Inquiry as the core framework for the social-presence and community-building strand of its recommended practice. ## CoI under generative-AI pressure GenAI destabilises the assumption that indicators of presence can be attributed primarily to human learners and [[teacher-role|instructor]]s (see [[reconceptualizing-community-inquiry-generative-ai|Ba, Gašević, Lim & Anderson]]). It can: - **Inflate cognitive presence** — fluent machine-generated explanations accelerate sense-making but risk premature closure when coherence is mistaken for warrant. - **Mimic social presence** — machine-produced utterances can resemble empathy and responsiveness with high linguistic credibility, complicating relational accountability. - **Redistribute teaching presence** — design, facilitation, and direct instruction become distributed accomplishments, with instructors modeling how to interrogate generated outputs and detect hallucinated citations. Rather than a tool, a dialogic partner, or a speculative "fourth presence," GenAI is best understood as an **epistemic condition** — a pervasive influence that reconfigures how presences are enacted, interpreted, evidenced, and governed through both visible outputs and invisible training-data, algorithmic, and platform logics. ## Practical implications - The relationship between GenAI involvement and inquiry quality is **conditional on human accountability**, not linear: strong presence can occur with high or low GenAI involvement when accountability is strong, and weak presence with either when accountability is weak. - [[assessment]] of inquiry should shift from polished final outputs to **process-sensitive evidence** — [[prompt-engineering|prompting]] and revision traces, disclosure and attribution practices, verification moves, and interaction logs. This aligns with the knowledge base's broader move toward [[authentic-assessment|authentic, process-revealing assessment]] and away from detection-based responses. - Pedagogically, learners often need explicit training (e.g., [[simulation|simulation-based]] practice with scripted roles and GenAI decision points) to sustain authentic inquiry under GenAI conditions — and instructors should model how to interrogate generated outputs, which depends on building [[ai-literacy]], [[critical-thinking|critical appraisal]], and calibrated [[trust-calibration|trust]]. ## Connected Concepts - [[online-teaching-and-learning]] - [[generative-ai]] - [[constructivist]] - [[pedagogy]] - [[critical-thinking]] - [[assessment]] - [[human-ai-collaboration]] - [[student-engagement]] - [[llm]] - [[ai-literacy]] - [[authentic-assessment]] - [[trust-calibration]] ## Connected Articles - [[reconceptualizing-community-inquiry-generative-ai]] — Reconceptualizing CoI in the age of generative AI - [[ai-communities-of-inquiry-2026]] — AI in communities of inquiry - [[ai-online-education-engagement-satisfaction-2026]] — AI and online education engagement - [[reflective-triangle-model-teacher-ai-2026]] — Reflective Triangle Model: AI as cognitive mediator --- ## [Self-Regulated Learning](https://edtechdev.github.io/aied/concepts/self-regulated-learning/) > Self-regulated learning (SRL) describes learners as active participants who can shape and develop their cognitive and behavioral actions in a successful way. AI tools can either [[scaffolding|scaffold]] SRL development or inadvertently short-circuit it by removing the regulatory demands that build expertise.([[scheu-mobile-chatbot-journaling-motivation-2026]])([[stanford-evidence-base-ai-k12-2026]]) ## Questions to Consider - SRL describes learners actively managing their learning through three phases: forethought (goal setting, planning, self-efficacy), performance (strategy, self-observation), and self-reflection (evaluation, adaptation). Before you read, which phase do you actually do well — and which do you skip even though you know better? - The page's core tension: AI can scaffold self-regulation or short-circuit it by removing the regulatory demands that build expertise. How can a tool that makes a task easier also make you a weaker regulator of your own learning — and can you feel the difference in your own use? - Students often show a 'production deficit': they possess self-regulation knowledge but fail to deploy it spontaneously — asking a chatbot to 'extract the main ideas' and skipping planning and monitoring entirely. Have you caught yourself doing the cognitive equivalent of this, even while knowing the better strategy? - Research found a '[[trust-calibration|miscalibration]] gap': students can *perceive* more learning with GenAI while retaining less — preferring AI over note-taking despite weaker retention. If you feel productive while using a tool, how would you ever discover that you're not actually learning more? - Whether GenAI functions as a scaffold, shortcut, or partner depends more on the learner's regulatory capacity than on the tool itself. But the page also shows self-regulation buffers — yet does not cancel — the harm of deep cognitive offloading. What does that 'buffers but doesn't cancel' caveat mean for designing better AI tools? - Set a goal before reading: pick one task you regularly use AI for, and decide in advance which of the three SRL phases (forethought, performance, reflection) you'll deliberately protect from being automated. What result will tell you it worked? ## Introduction SRL is the process whereby learners actively manage their own learning through three interrelated phases: 1. **Forethought:** Goal setting, strategic planning, [[self-efficacy]] beliefs 2. **Performance:** Strategy deployment, self-observation, [[cognitive-psychology|attention]] focusing 3. **Self-reflection:** [[self-assessment]], causal attribution, adaptation Proficient self-regulated learners employ cognitive strategies to improve success and utilize [[metacognition]] to refine their learning processes continuously.([[scheu-mobile-chatbot-journaling-motivation-2026]]) Crucially, SRL around AI is shaped by *perception* as well as behavior: [[yilmaz-genai-feedback-srl-online-higher-ed-2026|Yilmaz et al.]] demonstrate that whether students perceive feedback as coming from AI or a human significantly affects their self-regulated learning and revision behavior — a reminder that the social framing of AI, not just its content, changes how learners regulate around it. ## Digital Support for SRL ### Learning Journals Learning journals are a promising SRL intervention: by reflecting on their learning processes, students increase awareness of cognition and strengthen regulatory capacity. Key design considerations: - **Structure matters:** Open-ended journals often produce shallow entries; guided prompts and example models improve depth - **Motivation decay:** Mobile journaling apps commonly see rapid [[student-engagement|engagement]] decline after a few days - **Scaffolding trade-off:** AI assistance that writes reflections for students undermines the SRL practice; assistance that structures prompts without authoring content preserves it ### Dashboards communicating SRL profiles to teachers [[mejia-domenzain-ml-findings-teachers-blended-2026|Mejia-Domenzain et al. (2026)]] extend digital SRL support to the teacher side: their [[learning-analytics]] dashboard (DashED) communicates ML-derived self-regulated learning profiles to teachers in blended classrooms, and how teachers act on those profiles is context-dependent. In use, flipped-classroom (university) teachers followed a sequential exploration and favored course-level adaptation and showing [[visualization|dashboards]] in class, whereas vocational teachers revisited summary pages and used the tool mainly for individual coaching sessions. The actions teachers proposed were shaped by the content represented and their teaching level rather than the plot type — university teachers favored weekly tests and course adaptation, vocational teachers direct, individualized coaching. This positions the dashboard as a scaffold for teachers' regulation of instruction, with design needs that vary by context rather than a single optimal interface. ### Scheu et al.'s 2×2 Experiment (2026) In a randomized field experiment with 179 students over 22 days, two design principles were compared: | Principle | Mechanism | Effect on SRL | Effect on Motivation | Effect on Engagement | |---|---|---|---|---| | **Example-based course** | 7-day [[curriculum-design|curriculum]] [[teacher-role]] reflective journaling via modeled responses | Increased perceived competence and enjoyment | **Positive** | Constant positive | | **[[llm]] journaling assistant** | GPT-3.5 summarizes drafts, asks clarifying questions, suggests reformulations | No direct SRL skill effect measured | **No effect** | Increasing over time ([[feedback|feedback loop]]) | **Key insight:** The course improved SRL skills *and* intrinsic motivation through [[transfer-of-learning|skill transfer]], while the assistant improved engagement without affecting motivation.([[scheu-mobile-chatbot-journaling-motivation-2026]]) ## AI Tools and the SRL–Motivation Reciprocal Loop A foundational principle of SRL theory is that self-[[regulation]] skills and [[motivation]] form a **reciprocal relationship**: - Better SRL → more successful learning → higher self-efficacy → stronger motivation - Higher motivation → more effortful engagement → better SRL practice AI tools can enter this loop at different points: - **SRL-first design** (e.g., structured courses, graduated hints, reflection prompts): Strengthens the loop by building genuine skill - **Engagement-first design** (e.g., autocomplete, content generation): May boost behavioral engagement without entering the motivation loop, risking tool dependence ### Strategic Regulation of GenAI as SRL [[ai-anxiety-strategic-regulation-writing-2026|Kim (2026)]] reframes effective [[generative-ai|GenAI]] use in [[writing-education|academic writing]] as **strategic regulation** — an enacted SRL practice of verifying, revising, selectively adopting, or rejecting AI output. In a [[mixed-methods-research|mixed-methods]] study of 107 students, higher [[anxiety-and-stress|AI anxiety]] was positively associated with verification and revision (β=.24), while evaluative capacity predicted active revision and selective integration (β=.46). Students clustered into four regulatory types — Uncritical Reliance (18.7%), Selective Integration (34.6%), Evaluative Transformation (31.8%), and Strategic Rejection (14.9%) — showing that [[ai-literacy]] in [[higher-ed]] functions less as acceptance than as regulatory competence grounded in [[evaluative-judgment]] and [[ethics|ethical]] responsibility. This positions SRL as the core mechanism distinguishing critical from uncritical AI use. **The interaction itself as an object of regulation.** [[brunnstrom-ai-interaction-literacy-srl-2026|Brunnström and Palmqvist (2026)]] document the same regulation demand from the other direction: in an eight-round demonstration using a chatbot to prepare a take-home [[summative-assessment|examination]] answer, the AI's default output stayed at the *[[quantitative-research|quantitative]]*, multistructural end of the SOLO taxonomy — polished, submission-ready and pedagogically thin — and reached a usable three-step learning loop only after repeated meta-level interventions ("this is overwhelming, can you condense it?"). Their conclusion is that productive use required "the very self-regulatory skills the tool was expected to support": the learner must set incremental goals, request difficulty adjustments, and reflect on what is not yet understood, on top of the disciplinary content itself. They name this capacity [[ai-literacy|AI-interaction literacy]] and treat disengaging from the tool as a legitimate regulatory decision rather than a failure of persistence ([[metacognition]]). - **Satisfaction is not self-regulation.** [[aigc-affordance-student-self-regulation-2026|Liang et al. (2026)]] surveyed 689 undergraduates in industry-education programs and tested a serial mediation model in which the perceived affordances of AI-generated content raise AIGC [[self-efficacy]] (beta = 0.583) and, through it, learning [[motivation]] (beta = 0.565) and self-regulated learning (beta = 0.250), with motivation the heaviest single predictor of SRL (beta = 0.527). The load-bearing negative results sit alongside those paths: the quality of AI assessment feedback predicted satisfaction strongly (beta = 0.712) but not self-efficacy (beta = 0.131), and satisfaction had no significant effect on self-regulated learning (beta = 0.032). A well-liked, well-functioning assistant is therefore not evidence that regulation improved — the mechanism runs through confidence and motivation, not through the learner's experience of the tool. ## Relationship to Tutoring-Specific Designcritical from uncritical AI use. ### GenAI-aware reflection as SRL [[5p-reflection-model-genai-2026|The 5P Reflection Model (Kadel et al. 2026)]] re-centers structured reflection in the GenAI era as an enacted SRL practice. Because learners increasingly co-create meaning with AI, traditional reflection models struggle to authenticate student reflection, so the 5P model (Purpose, Process, Product, Pitfalls, Plan) fuses forethought-driven goal setting, reflection-in-action (documenting prompts and iterations), reflection-on-action (validating the probabilistic output against external sources), an explicit pitfalls stage for [[hallucination-risk|hallucination]], [[academic-integrity|plagiarism]], and [[cognitive-offloading|over-reliance]], and a forward-looking plan — embedding emotional monitoring throughout. Its "process over product" philosophy treats structured documentation of the [[human-ai-collaboration|human-AI interaction]] as the regulatory demand that preserves authenticity, [[agency]], and [[metacognition|metacognitive]] depth, positioning GenAI-aware reflection as a scaffold for self-regulation rather than a substitute for it. ## Relationship to Tutoring-Specific Design [[stanford-evidence-base-ai-k12-2026|Tutoring-specific AI]] aligns with SRL-first design: it provides graduated scaffolds that preserve [[agency]] and require strategic self-regulation. General-purpose AI often removes the regulatory demands entirely.([[stanford-evidence-base-ai-k12-2026]]) For example: - Bastani et al.'s tutoring-specific [[conversational-ai|chatbot]] preserved step-by-step reasoning (SRL demand) - The general-purpose GPT variant simply provided answers (SRL bypass) ## Evidence Across Contexts - **Mixed evidence and the miscalibration gap.** A rapid review of PreK-12 GenAI research finds metacognitive gains during supported tasks often do not persist when support is removed, and that GenAI can increase [[self-report-measures|perceived learning]] even when durable learning is absent (the miscalibration gap — students preferred GenAI over note-taking despite weaker retention). Students need explicit, stage-appropriate training to decide what to delegate and when independent effort matters.([[young-people-learning-generative-ai-rapid-review-2026]]) - **[[agentic-ai|Agentic]] initiative vs. self-regulation tension.** [[agentic-ai-pedagogical-best-practice-2026|Woollaston et al. (2026)]] note that as agents automate more of a task, the less self-regulated cognitive work the learner performs — so designs should give learners control over agent initiation (dynamic, fading scaffolding) to preserve self-regulatory capacity rather than outsourcing it. - **Self-regulation shapes AI coding-assistant use.** [[computational-thinking-aica-2026|A study of AI coding assistants]] found high-[[computational-thinking]] students showed stronger self-regulatory coherence (planning-execution-self-reflection) and used AICA for code understanding, while low-CT students used it for immediate answer retrieval. - **SRL co-occurs with lower digital distraction in [[online-teaching-and-learning|online learning]].** [[decreasing-digital-distraction-college-online-learning-2026|Shi et al. (2026)]], using unsupervised data mining on 530 college students, found that SRL strategies — goal setting, environment structuring, and time management — co-occurred most consistently with lower digital distraction in [[higher-ed|online learning]], alongside learner-instructor and learner-content engagement. The finding positions concrete SRL training as a high-leverage intervention for focused online study. ## LLM-Mediated SRL: Scaffold, Shortcut, or Partner? A cluster of Learning Letters studies (2026) converges on a central tension: [[generative-ai|GenAI]] can scaffold, short-circuit, or partner with self-regulation depending on design and how learners regulate its use. The evidence points to SRL itself — not the tool — as the decisive variable. - **[[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026|Viberg et al.]]** find that LLMs are woven into a *layered [[help-seeking]] ecosystem* rather than replacing human support: students try tasks independently first, then consult ChatGPT as a low-barrier first step, peers for conceptual negotiation, and instructors for high-stakes issues. They favor **instrumental help-seeking** (hints, step-by-step guidance) over **executive help-seeking** (direct solutions), exercising selective [[trust]] and verifying outputs against course materials — a four-stage process (deciding whether help is needed, choosing whom to ask, determining the type of help, judging the help received) that can be measured and taught. - **[[atif-dickson-deane-scaffold-shortcut-genai-srl-2026|Atif & Dickson-Deane]]** frame GenAI use as **[[cognitive-offloading]]** that can be either *scaffolded* (learners critique and adapt AI outputs, keeping [[agency]] and sense-making) or *substitutional* (learners accept outputs with minimal verification, shifting control to the tool). In a study of 267 postgraduate IT students, the same tool could scaffold or shortcut SRL depending on learner strategy — confident users showed agency in goal setting and monitoring; less confident users saw GenAI as a shortcut or misconduct. - **[[lim-bannert-student-regulation-genai-chatbot-2026|Lim & Bannert]]** show the risk concretely: students voluntarily used a genAI chatbot (73%) and scored higher on essays, but they offloaded comprehension and synthesis (asking the chatbot to "extract only the main ideas") and engaged in almost no planning or monitoring — outsourcing key regulatory decisions. This reflects a **production deficit**: students possess SRL knowledge but fail to deploy it spontaneously, so genAI tools should prompt reflection (a monitoring scaffold) when queries indicate offloading. - **[[song-genai-learning-partner-srl-over-time-2026|Song et al.]]** demonstrate that SRL is both a **stable aptitude and a dynamic state**: individual baselines are consistent, but metacognitive knowledge and [[well-being]] decline systemically over a semester, driven by assessment deadlines. They show GenAI can act as a context-aware **[[pedagogical-agent|learning partner]]** when it is given personal, temporal, and contextual data — supporting students without replacing their effort. This argues against "one-time-fits-all" [[personalized-learning|personalization]] based on baseline aptitude alone. - **[[de-barba-srl-genai-2026|de Barba]]** extends SRL theoretically, arguing the field has narrowed to task-focused regulation and to optimisable behavioral proxies in educational technology. The paper proposes a cross-scale account of **learner agency** — regulation (within tasks), integration (across time and contexts), and positioning (critically in relation to the conditions framing learning) — as a design orientation for algorithmically mediated environments. - **Self-regulation buffers but does not cancel offloading harm.** [[layer-sensitive-cognitive-offloading-writing-2026|Chen (2026)]] shows that self-regulated writing attenuates — but does not eliminate — the negative association between deep [[cognitive-offloading]] and independent no-AI outcomes in GenAI-assisted writing: the offloading-by-SRL interaction was positive (B = 0.22), flattening the harm from a slope of −0.54 (low SRL) to −0.33 (high SRL) but not canceling it. A bounded-support condition pairing delegation limits with compulsory reflection produced the strongest independent performance, evidence that metacognitive regulation partially protects learners yet cannot fully compensate for delegating the cognitive work itself. The collective lesson: **SRL is the core mechanism distinguishing critical from uncritical AI use.** Whether GenAI functions as a scaffold, shortcut, or partner depends on learners' regulatory capacity and on whether tools are designed to preserve (rather than remove) the regulatory demands that build expertise. ## Implications - **For journaling/chatbot tools:** Combine SRL instruction (course-based) with optional writing support to get both motivation and engagement gains - **For [[educational-policy-ai|AI policy]]:** Procurement criteria should ask whether a tool develops or displaces self-regulation - **For [[research-methods-aied|researchers]]:** Long-term studies measuring SRL outcomes (not just immediate performance) are essential ## Conversational Agents and SRL in Simulation Games - **Conversational agents supporting self-regulated learning in games.** Wenzel, Geiger, and Liening (2026) show that an AI conversational agent (Lara) in a business [[simulation]] game can support self-regulated learning through metric-based [[formative-assessment|formative]] feedback, on-demand guidance, and structured reflection — addressing the common limitation that simulation games provide limited formative feedback and reflection prompts. Evaluations with student teachers and BSG participants reported positive perceptions of the agent's cognitive and [[community-of-inquiry|social presence]] and its support for self-regulation. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[metacognition]] — the cognitive monitoring SRL relies on - [[self-assessment]] - [[self-efficacy]] — a forethought-phase belief driving effort - [[scaffolding]] — graduated support that preserves regulatory demand - [[feedback]] — input learners regulate around - [[feedback-literacy]] — the capacity to act on feedback - [[help-seeking]] — a strategic SRL behavior - [[motivation]] — the reciprocal partner of self-regulation - [[cognitive-offloading]] — the risk when AI removes regulatory work - [[generative-ai]] — the technology that can scaffold or short-circuit SRL - [[ai-literacy]] — regulatory competence in AI use - [[self-directed-learning]] — the broader autonomy construct - [[agency]] — the learner's capacity to act with intention, central to regulation, integration, and positioning - [[adaptive-learning]] — personalization that can support regulation - [[formative-assessment]] — continuous feedback for regulation - [[learning-by-teaching]] — a strategy building self-regulation - [[intelligent-tutoring]] — systems that scaffold SRL - [[llm]] — the underlying model of AI tools - [[retrieval-spacing-interleaving]] — scheduling, self-testing and study-strategy choices learners make - [[cognitive-surrender]] ## Connected Articles - [[genai-performance-vs-learning]] — offloading planning, monitoring and evaluating short-circuits the SRL loop (Yan et al. 2025) - [[aigc-affordance-student-self-regulation-2026]] — AIGC Affordance and Student Self-Regulation - [[brunnstrom-ai-interaction-literacy-srl-2026]] — AI-interaction literacy: steering a chatbot demanded the SRL it was meant to support (Brunnström & Palmqvist 2026) - [[5p-reflection-model-genai-2026]] — The 5P reflection model for the GenAI era (Kadel et al. 2026) - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[reclaiming-epistemic-agency-co-agency-2026]] - [[de-barba-srl-genai-2026]] — Learner agency across scales: regulation, integration, positioning - [[song-genai-learning-partner-srl-over-time-2026]] — GenAI as a context-aware learning partner over time - [[lim-bannert-student-regulation-genai-chatbot-2026]] — How students regulate learning with a genAI chatbot - [[atif-dickson-deane-scaffold-shortcut-genai-srl-2026]] — Scaffold or shortcut? GenAI dual role in SRL - [[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026]] — LLM-mediated help-seeking in STEM: layered, instrumental, and verified - [[your-brain-on-chatgpt-cognitive-debt-essay-writing]] - [[mejeh-fromm-srl-adaptive-learning-feedback-2026]] - [[banihashem-ai-srl-systematic-mapping-review-2025]] - [[yilmaz-genai-feedback-srl-online-higher-ed-2026]] — GenAI feedback and SRL: perceived source matters - [[ai-anxiety-strategic-regulation-writing-2026]] — From AI anxiety to strategic regulation - [[idea-framework-metacognitive-genai-2026]] — The IDEA framework for metacognitively regulated GenAI use - [[bilingual-llm-lecture-companion-srl-2026]] — SRL with a bilingual LLM lecture companion - [[generative-ai-reduced-study-time-math]] — Cognitive surrender as loss of self-regulated learning - [[young-people-learning-generative-ai-rapid-review-2026]] — Mixed evidence on metacognition/self-regulation with GenAI - [[agentic-ai-pedagogical-best-practice-2026]] — Agentic AI and the self-regulation tension - [[ai-cognitive-partner-co-regulation-learning]] — AI as cognitive partner in co-regulated learning - [[making-ai-annoying-constrained-writing-2026]] — Making AI annoying on purpose: constraint in AI-supported writing (Konradt, Boote & Taub 2026) - [[student-motivation-need-satisfaction-genai-sdt-2026]] — Student motivation and need satisfaction in GenAI classrooms (Schweder, Hagenauer & Raufelder 2026) - [[decreasing-digital-distraction-college-online-learning-2026]] — SRL and lower digital distraction in online learning (Shi et al. 2026) - [[conversational-agents-business-simulation-gaming-2026]] — CAIS-GBL framework for AI conversational agents in business simulation games (Wenzel et al. 2026) - [[mejia-domenzain-ml-findings-teachers-blended-2026]] — Making ML findings accessible to teachers in blended classrooms - [[scan-framework-task-assignment-generative-ai-2025]] — SCAN: a learner-facing loop of task identification, justification and post-task reflection - [[learning-analytics-genai-secondary-writing-2026]] — Using Learning Analytics to Support Secondary School Students' Writing with Generative AI --- ## [Self-Determination Theory](https://edtechdev.github.io/aied/concepts/self-determination-theory/) > **Self-Determination Theory (SDT)** — a psychological theory of human motivation positing that intrinsic motivation and [[well-being]] depend on satisfying three basic psychological needs: autonomy, competence, and relatedness. In [[ai-education|AI in education]], SDT provides a framework for designing AI tools and professional development that support rather than undermine learners' and teachers' motivation. ## Questions to Consider - The theory claims motivation isn't just how *much* you have but a *quality* shaped by the environment, built on three needs: autonomy, competence, and relatedness. Think of a learning experience that drained you. Which of those three needs was violated, and what would have restored it? - A year-long study found three motivational profiles — Disengaged, Developing, Self-Determined — that were stable over time, and students who reached the Self-Determined profile showed the greatest AI-literacy gains. Before you read, is motivation something students bring with them, or something a well-designed environment can grow? What does 'developmental, not fixed' imply? - If an AI tool makes every task effortless and 'easy to use,' which psychological need might it be satisfying — and which might it be quietly undermining? How could a tool that boosts short-term engagement still erode long-term motivation? - Need-supportive professional development for teachers enhanced their AI literacy and sustained engagement. Does this suggest that how we *train* educators about AI matters as much as what the AI itself does? What would 'autonomy-supportive' AI training for you personally look like? - A study found ChatGPT could support autonomy, relatedness, and competence in language learning. But could the same tool undermine those needs for a different learner? What would need to be true about *how* it's used for the theory to hold? - Before reading further, name one way you've felt your own competence, autonomy, or sense of connection affected by using an AI tool — and reflect on whether you'd have noticed that change without being prompted to look for it. ## Introduction SDT is increasingly used in AI in education [[research-methods-aied|research]] as a theoretical lens for both learner-facing and teacher-facing AI systems. The theory's central claim — that motivation is not simply a quantity learners have but a quality shaped by the social and technological environment — makes it directly relevant to questions about how AI tools affect [[student-engagement|engagement]], persistence, and [[learning-gains|learning outcomes]]. The articles in this knowledge base apply SDT across three main contexts: teacher professional development, AI-mediated learning engagement, and affective computing. ### Key research themes **SDT-based teacher professional development** applies the theory's need-supportive principles to prepare educators for AI. **[[teacher-education-ai-literacy-sdt-2026|Chiu et al.]]** studied 382 [[k-12|secondary school]] teachers, finding that need-supportive professional development grounded in SDT enhances teachers' [[ai-literacy]] and fosters sustained behavioral engagement in online professional learning communities. [[qualitative-research|Qualitative]] analysis identified nine design strategies supporting autonomy, competence, and relatedness — bridging the gap between isolated professional development and professional learning communities. **SDT in AI-mediated learning engagement** examines how [[generative-ai|generative AI]] tools shape student motivation. **[[students-engagement-with-generative-ai-in-academic-learning-a-self-determination|Isaeva et al.]]** combined SDT with [[network-analysis|epistemic network analysis]] to study students' engagement with generative AI in academic learning. **[[ai-availability-student-motivation]]** explores how AI availability affects student motivation and persistence, connecting to [[cognitive-offloading|Over-Reliance]] concerns about motivation erosion. **[[liang-ai-learning-motivation-sdt-2026|Liang et al. (2026)]]** extend SDT to AI learning with a latent transition analysis of **2,086 secondary students** in a year-long AI [[curriculum-design|curriculum]], identifying **three motivational profiles (Disengaged, Developing, Self-Determined)** that were stable across time and showing that most students maintained or advanced toward higher profiles. Crucially, students who reached or remained in the Self-Determined profile showed the **greatest [[ai-literacy]] gains** — direct longitudinal evidence that satisfying autonomy, competence, and relatedness predicts better AI-learning outcomes, and that motivation is a developmental (not fixed) learner property. **Dual pathways from learning climate to AI use.** [[dual-ai-learning-pathways-sdt-2026|Shen and Arunrugstichai (2026)]] integrate SDT with the Hook model of behavioral engagement to explain why GenAI use ranges from constructive support to compulsive dependence. In cross-sectional survey data from **508 university students** with different high-school backgrounds and current contexts (China and Thailand), retrospective reports of **high-school pressure vs. autonomy support** differentially predicted which of two pathways students followed into university — one toward constructive, autonomous GenAI use and another toward compulsive dependence — with results consistent across cross-contextual multi-group analyses. The model links SDT motivational processes to perceived quality of AI-supported learning, framing autonomy support as a lever that steers students toward productive rather than dependent AI use. - **ChatGPT and SDT needs in language learning:** [[chatgpt-english-language-learning-malaysia|Annamalai et al. (2026)]] used an SDT lens with 25 Malaysian university students, finding that ChatGPT supports autonomy, relatedness, and competence in [[language-learning|English language learning]] — enhancing grammar, writing, and conversational tasks while letting [[teacher-role|educators]] focus on higher-order training. **SDT applied to instructors' own AI-mediated practice.** [[claassen-learning-analytics-genai-learning-design-2026|Claassen et al. (2026)]] used SDT as the interpretive lens on how instructors integrate [[learning-analytics|learning analytics]] and generative AI into [[learning-design|learning design]] — finding that supporting instructors' basic needs (autonomy, competence, relatedness) fosters the creative [[problem-solving]] their design work requires. In their ENA analysis, GenAI use was associated with designing for student self-determination (e.g., co-creating assessment rubrics with students), extending SDT from learners to the educators who build need-supportive AI-mediated environments. **Autonomy support as the frame for children's GenAI use.** [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] relocate the question of responsible use from restriction to need support, arguing that the distinction that matters is whether adults around a child support autonomy rather than control it, and distinguishing dependent from autonomous [[cognitive-offloading]] within SDT terms: dependent offloading transfers [[agency]] and lowers intrinsic motivation, autonomous offloading scaffolds while the learner retains epistemic control. Two features of the review are directly relevant to SDT application: it insists that autonomy support is not permissiveness, and it treats the family-school coordination that current guidance assumes as an untested hypothesis, formalizing additive, synergistic and compensatory versions that only a factorial trial contrasting family-only, school-only, coordinated and usual-practice guidance could discriminate. ## Connections to related concepts SDT connects directly to [[motivation]] as its parent construct, to [[affective-computing]] and [[affective-tutoring]] for emotion-aware AI design, and to [[student-experience]] for how learners experience AI-mediated environments. The theory's emphasis on autonomy connects to [[self-regulated-learning]], while its competence dimension connects to [[self-efficacy-tutoring-learning]] and [[teacher-ai-competency]]. SDT is particularly relevant to [[professional-training]] and [[educational-development]] because need-supportive design is a transferable principle for preparing educators to use AI. ## Connected Concepts - [[motivation]] - [[student-experience]] - [[affective-computing]] - [[affective-tutoring]] - [[self-regulated-learning]] - [[teacher-ai-competency]] - [[educational-development]] - [[professional-training]] - [[cognitive-offloading]] - [[student-engagement]] - [[ai-education]] - [[learning-theories]] ## Connected Articles - [[family-school-autonomy-support-genai-2026]] — Family-School Autonomy Support for Children's Responsible Use of Generative AI - [[dual-ai-learning-pathways-sdt-2026]] — High-school pressure/autonomy support and dual AI learning pathways (Shen & Arunrugstichai 2026) - [[reclaiming-epistemic-agency-co-agency-2026]] - [[claassen-learning-analytics-genai-learning-design-2026]] — LA and GenAI in learning design decision-making - [[teacher-education-ai-literacy-sdt-2026]] - [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination]] - [[ai-availability-student-motivation]] - [[chatgpt-english-language-learning-malaysia]] — Students' ChatGPT experiences in English language learning - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions - [[liang-ai-learning-motivation-sdt-2026]] — SDT latent transition analysis of students' AI learning motivation (2,086 secondary students) - [[student-motivation-need-satisfaction-genai-sdt-2026]] — Student motivation and need satisfaction in GenAI classrooms (Schweder, Hagenauer & Raufelder 2026) - [[ai-literacy-course-satisfaction-pbl-scale-2026]] — AI-PBLS scale; empowerment and ethical awareness mediating PBL-to-satisfaction in AI literacy courses (Zhu & Kong 2026) --- ## [Motivation](https://edtechdev.github.io/aied/concepts/motivation/) > **Motivation** — the psychological processes that initiate, direct, and sustain goal-directed behavior. In [[ai-education|AI in education]], motivation [[research-methods-aied|research]] examines how AI tools affect learners' and teachers' motivation — whether AI [[scaffolding|scaffolds]] or undermines persistence, curiosity, and intrinsic engagement — and how motivational states shape the effectiveness of AI-mediated learning. ## Questions to Consider - Motivation is often treated as a trait some students 'have' and others lack. The page describes it instead as psychological processes that initiate, direct, and sustain behavior — and as something developmental and socially scaffolded. How does that reframe who is responsible for student motivation? - AI tools can remove friction and make learning more accessible, but they can also reduce the cognitive effort and struggle that support intrinsic motivation. When has making something 'easier' actually made it less motivating or less satisfying for you? - Self-determination theory says intrinsic motivation grows from autonomy, competence, and relatedness. If an [[intelligent-tutoring|AI tutor]] does most of the work, which of those three might it threaten — and which might it enhance? - Research finds distinct motivational profiles among students in an AI [[curriculum-design|curriculum]], and that reaching a self-determined profile predicted the largest AI-literacy gains. What might it take for a learner to move from passively disengaged to genuinely self-determined in using AI? - The page reports that teacher support drives AI-assisted engagement largely through mastery-approach and performance-approach goals — the 'approach' rather than 'avoidance' orientations. How does the way a teacher frames AI use ('to get it right' vs. 'to avoid looking wrong') shape whether students engage deeply? - If you're designing for motivation, is the goal to make learning easier, more engaging, or more meaningfully effortful? Where do [[accessibility]] and intrinsic motivation pull in opposite directions? ## Introduction Motivation is a foundational construct in education research, and the rise of AI in education has made it more consequential: AI tools can remove friction and make learning more accessible, but they can also reduce the cognitive effort and struggle that support intrinsic motivation and deep learning. The articles in this knowledge base explore motivation across learner-facing AI tools, teacher-facing AI systems, and the psychological mechanisms — [[self-determination-theory|self-determination]], [[self-efficacy-tutoring-learning|self-efficacy]], emotions — through which AI shapes motivated behavior. - **[[lee-wu-gender-motivation-genai-achievement-2026|Lee & Wu]]** show gender and motivation drive differential engagement with GenAI, with distinct [[learning-gains|achievement]] trajectories. ## Key research themes **AI effects on student motivation** is the most direct line of research. **[[ai-availability-student-motivation]]** examines how the availability of AI assistance affects student motivation and persistence, connecting to [[cognitive-offloading|Over-Reliance]] research on motivation erosion when AI does the work. **[[scheu-mobile-chatbot-journaling-motivation-2026]]** explores mobile [[conversational-ai|chatbot]] journaling as a motivational intervention. **[[ai-learning-tools-engineering-education-needs]]** examines what motivates students to adopt AI learning tools in [[engineering-education|engineering education]]. **Motivation in AI-mediated engagement** examines how motivational quality (not just quantity) changes with AI. **[[students-engagement-with-generative-ai-in-academic-learning-a-self-determination|Isaeva et al.]]** combined self-determination theory with [[network-analysis|epistemic network analysis]] to study engagement with [[generative-ai|generative AI]]. **[[liang-ai-learning-motivation-sdt-2026|Liang et al. (2026)]]** traced motivation developmentally via latent transition analysis of **2,086 [[k-12|secondary]] students** in a year-long AI curriculum, finding three stable profiles (Disengaged, Developing, Self-Determined) and that reaching the Self-Determined profile predicted the largest [[ai-literacy]] gains. **[[wang-goal-setting-ai-engagement-2026|Wang & Wang (2026)]]** used goal-setting theory with **758 [[higher-ed|university]] English learners**, showing that **teacher support** drives AI-assisted engagement primarily through mastery-approach and performance-approach goals (the approach, not avoidance, goal orientations). Together these studies show that motivation in AI contexts is both developmental and socially scaffolded — it shifts over time and responds to teacher support and goal framing, not just tool design. **Motivation gains are construct-specific, not general.** [[genai-writing-program-primary-l2-motivation-engagement|Lu et al. (2026)]] found that a nine-week GenAI-supported [[writing-education|writing]] program for 301 Grade 5 and 6 learners raised their ideal L2 writing self and academic buoyancy — the aspirational and the resilience components of motivation — while leaving growth mindset unchanged; the only growth-mindset gain appeared in the control group and did not survive correction for multiple comparisons. Students attributed the shift to seeing fluent text built from vocabulary they already knew, which made successful writing feel attainable. So motivation is not a single dial that AI turns up: what improved was the belief that one *can* write well, not the belief that ability grows with effort. The same construct-specificity appears when GenAI is itself the relevance intervention. [[genai-math-relevance-intervention-2026|Guo, Fryer and Shum (2026)]] had **218 high-school students** spend one hour in semi-structured dialogue with a [[conversational-ai|chatbot]] aimed at personal relevance to [[math-education|math]]; the collective, class-level version raised relevance as identification (F = 4.35, p = .014, η² = 0.04; against the control, F = 11.11, p = .001, η² = 0.073 — a medium effect) and, in the SEM, predicted relevance to a specific lesson one week later (β = 0.18, p < .01), while interest in the math class did not move (F = 0.29, p = .75, η² = 0.003). The authors attribute the decoupling to dose: a single one-hour session is too brief for relevance gains to consolidate into interest in the class. **Teacher motivation and persistence** examines motivation among educators. **[[framing-5-percent-problem-teachers-persistence|Framing the 5 Percent Problem]]** studies teacher persistence with AI tools, and **[[teacher-education-ai-literacy-sdt-2026|Chiu et al.]]** found need-supportive [[educational-development|professional development]] fosters sustained behavioral engagement in professional learning communities. - **Cross-cultural motivation of future teachers:** [[motivation-shape-future-education-ai-switzerland-china|Martínez-Moreno et al. (2026)]] validated the (D)FIT-Choice scale with 416 student teachers in Switzerland and China, finding Swiss teachers report stronger social utility and intrinsic motivation while Chinese teachers show higher perceived digital competence and enthusiasm for integrating AI — highlighting how cultural and systemic factors shape motivation to shape the future of education with AI. **The effort paradox and the vicious cycle of assistance.** [[zohar-bloom-inzlicht-against-frictionless-ai-2026|Zohar, Bloom and Inzlicht (2026)]] argue that motivation is not simply helped or hindered by AI but redistributed: humans generally take the path of least resistance, yet they also seek effort out — the effort paradox — because effort signals that actions matter and because reward attached to process rather than product increases the tendency to strive and persevere. Two claims follow. First, the effort–meaning relationship is an inverted U, so the motivational target is moderate friction, and the risk of frictionless AI is overshooting into too little. Second, a vicious cycle: as AI replaces effort in a domain, the motivational benefits of effort there erode, which makes users more dependent on AI, which erodes motivation further. They also separate supplement from substitute by developmental stage — learners in earlier stages risk bypassing the experiences that build perseverance, while those with established skills can use AI to save time ([[desirable-difficulties]], [[self-efficacy]]). ## Connections to related concepts Motivation is the parent construct of [[self-determination-theory]], which specifies the psychological needs (autonomy, competence, relatedness) that sustain intrinsic motivation. It connects to [[student-experience]] as the experiential layer of motivated engagement, to [[student-engagement]] as its measurable dimension, and to [[affective-computing]] for the emotional mechanisms that shape motivation. Motivation also connects to [[cognitive-offloading|Over-Reliance]] (AI reducing [[desirable-difficulties|productive struggle]]), [[self-regulated-learning]] (motivated learners self-regulate), and [[teacher-role]] (motivation applies to educators as well as students). ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[self-directed-learning]] - [[self-determination-theory]] - [[student-experience]] - [[student-engagement]] - [[affective-computing]] - [[affective-tutoring]] - [[cognitive-offloading]] - [[self-regulated-learning]] - [[teacher-role]] - [[ai-education]] - [[framing-ai-use-for-students]] - [[social-emotional-learning]] — Social-Emotional Learning ## Connected Articles - [[zohar-bloom-inzlicht-against-frictionless-ai-2026]] — The effort paradox and the vicious cycle of frictionless assistance - [[cui-motivation-roles-metacognitive-genai-2026]] — Motivation and roles in metacognitive GenAI engagement - [[lee-wu-gender-motivation-genai-achievement-2026]] — Gender, motivation, and GenAI achievement trajectories - [[oby-chatgpt-use-learning-framework-2026]] - [[genai-thoughtless-use-self-directed-learning-2026]] - [[ai-student-engagement-online-learning-review-2025]] - [[ai-online-education-engagement-satisfaction-2026]] - [[chatgpt-perception-online-learning-engagement-2026]] - [[ethical-ai-higher-ed-game-theory]] — Coordination game framework for ethical AI use in higher education (Ogbo et al. 2026) - [[genai-student-experiences-uk-he-survey-2026]] - [[ai-availability-student-motivation]] - [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination]] - [[teacher-education-ai-literacy-sdt-2026]] - [[scheu-mobile-chatbot-journaling-motivation-2026]] - [[framing-5-percent-problem-teachers-persistence]] - [[self-efficacy-tutoring-learning]] - [[instructor-designed-ai-tutors-foreign-language-sdt-2026]] — Instructor-Designed AI Tutors in University Foreign Language Education: A Mixed-Methods Study of Learner Motivation and Reflective Learning Experience Based on Self-Determination Theory - [[context-based-ai-secondary-chemistry-2026]] — Context-based 7E + AI instruction in secondary chemistry - [[guillen-curriculum-genai-teacher-competence-2026]] — Assessing Teacher Digital Competence for GenAI Curriculum Design (Guillén-Gámez 2026) - [[motivation-shape-future-education-ai-switzerland-china]] — Motivation to shape the future of education with AI - [[chatgpt-english-language-learning-malaysia]] — Students' ChatGPT experiences in English language learning - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions - [[student-perceptions-ai-study-productivity-2026]] — Students' Perceptions of Artificial Intelligence Tools for Study Productivity and Learning: An Exploratory Survey Study - [[liang-ai-learning-motivation-sdt-2026]] — SDT latent transition analysis of students' AI learning motivation (2,086 secondary students) - [[wang-goal-setting-ai-engagement-2026]] — Goal-setting theory: teacher support, achievement goals, and engagement in AI-assisted English learning (758 Chinese students) - [[utility-value-intervention-teach-responsibly-genai-2026]] — Utility-value intervention effects in learning to teach responsibly with GenAI (Boos, Eder & Lachner 2026) - [[student-motivation-need-satisfaction-genai-sdt-2026]] — Student motivation and need satisfaction in GenAI classrooms (Schweder, Hagenauer & Raufelder 2026) - [[ai-literacy-course-satisfaction-pbl-scale-2026]] — AI-PBLS scale; empowerment and ethical awareness mediating PBL-to-satisfaction in AI literacy courses (Zhu & Kong 2026) - [[predicting-attrition-competitive-programming]] — Predicting Student Attrition in Competitive Programming - [[genai-writing-program-primary-l2-motivation-engagement]] — Construct-specific motivation gains in a primary L2 GenAI writing program (Lu et al. 2026) - [[air-scale-motivations-ai-reading-2026]] — The AIR Scale: four motive families for reaching for AI while reading - [[genai-math-relevance-intervention-2026]] — GenAI relevance dialogue raised relevance as identification but left class interest flat --- ## [Self-Efficacy](https://edtechdev.github.io/aied/concepts/self-efficacy/) > **Self-efficacy** — a learner's belief in their capability to successfully perform a task or achieve a goal. Drawing on social cognitive theory (Bandura), self-efficacy shapes motivation, effort, persistence, and learning [[student-engagement|engagement]]. In [[ai-education|AI in education]], self-efficacy matters in two ways: AI tools can build learners' confidence and autonomy (e.g., by providing [[feedback]] and [[scaffolding]]), and learners' AI self-efficacy — their confidence in using AI [[ai-technologies|technologies]] — influences how effectively they engage with AI, including how AI-related knowledge translates into career-relevant readiness. ## Questions to Consider - Self-efficacy is a belief about your capability — distinct from actual competence. Have you ever been highly capable at something yet doubted yourself, or confidently wrong about something you couldn't do? What does that gap between belief and ability tell you about why self-efficacy matters? - AI self-efficacy (confidence working with AI) is a separate construct from AI literacy, and research finds literacy translates into readiness only when learners also have confidence. If someone knows *about* AI but doesn't believe they can use it, what happens to that knowledge — and what does that imply for training? - Research found that using AI to support understanding was fully mediated by academic self-efficacy in its link to performance, while shortcut use predicted worse outcomes partly independent of self-efficacy. Why would the *same* tool build confidence when used one way and fail to when used another? - The page distinguishes self-efficacy from the everyday word 'confidence.' Before you read, how are they different, and why would a [[research-methods-aied|researcher]] insist on the distinction rather than treating them as the same thing? - Robotics and embodied learning build confidence by grounding tasks in observable outcomes, and feedback can build learner self-efficacy. Think of a task where you gained real confidence only after seeing a concrete result. What does that say about what kinds of AI learning experiences are most likely to build — rather than merely report — self-efficacy? - [[teacher-role|Teacher]] self-efficacy affects adoption and integration of AI. If a teacher doesn't believe they can use AI effectively, does any amount of AI literacy fix it? What would build a teacher's confidence, and how is that different from giving them more information? ## Introduction Self-efficacy is distinct from actual competence: it is a belief about capability that drives behavior. It is closely related to — and often used interchangeably with — the everyday notion of *confidence* in one's abilities. Self-efficacy connects closely to [[motivation]], [[self-regulated-learning]], and [[student-experience]]. In the AI context, AI self-efficacy (confidence in working with AI) is a distinct construct from AI literacy, and research shows it plays a crucial role in whether learners actually activate and apply AI-related knowledge. ### How self-efficacy appears in the knowledge base's research - **Declines under GenAI-plus-XR studio work:** in a 27-student architectural [[arts-design-and-media-education|design studio]], teams using a [[generative-ai|GenAI]] and multi-user XR pipeline showed larger relative pre–post declines in design self-efficacy confidence (β = −1.675, p = 0.021) and outcome expectancy (β = −2.088, p = 0.002) than teams continuing the normal workflow, with no significant difference in blinded panel ratings of their presentations ([[genai-xr-architectural-design-education-2026|Xiao et al., 2026]]). Tool-rich environments can depress efficacy beliefs even when the work itself is judged equivalent. - **AI self-efficacy and career readiness:** [[ai-literacy-career-adaptability-business-2026|Research on AI readiness]] shows that AI self-efficacy moderates the relationship between AI literacy and AI readiness: literacy translates into readiness only when learners have confidence in using AI, and self-efficacy directly predicts [[career-development-and-readiness|career adaptability]]. - **Trust as the pivot between literacy and confidence:** [[hu-psychological-predictors-continued-chatgpt-use-2026|Hu (2026)]] surveyed 450 university students who already use ChatGPT and found an ordered chain rather than two parallel correlates: [[ai-literacy|AI literacy]] related to [[trust]] in the tool (beta = 0.50), trust to academic self-efficacy (0.48), and self-efficacy to continued use, with the serial indirect effect significant (0.07, 95% CI [0.04, 0.10]). [[anxiety-and-stress|AI anxiety]] weakened the literacy to trust link (interaction beta = -0.25, simple slopes falling from 0.76 to 0.25 across the anxiety range), so the same knowledge bought less confidence, and reached less use, among more anxious students. Confidence is thus a downstream link in the chain rather than a starting point. - **Profiles of academic self-efficacy and who reaches for AI:** [[suria-martinez-academic-self-efficacy-motor-disabilities-2026|Suriá-Martínez et al. (2026)]] ran a latent profile analysis of academic self-efficacy among 102 university students with motor disabilities in Spain and found three profiles (low 29.4%, moderate 41.2%, high 29.4%) whose reported AI use rose stepwise with profile level (means 2.41, 3.56, 4.68; F(2, 99) = 27.84, p < .001, eta squared = .36), with Excellence, the planning and goal-setting dimension, most strongly associated with AI use (beta = .47). The authors read this as an [[equity-in-ai-education|equity]] problem: if confidence tracks with uptake, students with lower self-efficacy may be the least likely to reach for AI support that could reduce [[accessibility|access]] barriers, so support for academic self-efficacy belongs inside [[inclusive-learning|inclusion]] frameworks. - **AI use patterns and self-efficacy:** [[stamatoulis-genai-use-patterns-2026|Stamatoulis et al. (2026)]] found that using [[generative-ai|GenAI]] to *support understanding* (evaluative integration) was fully mediated by academic self-efficacy in its association with performance — understanding-oriented AI use builds confidence — whereas shortcut use (low-verification uptake) predicted worse outcomes partly independently of self-efficacy. Self-efficacy is thus both a pathway through which productive AI use helps and a factor that shortcut use may fail to build. - **Robotics and [[experiential-learning|hands-on learning]]:** [[remind-robot-mediated-roleplay-antibullying-2026|REMind]]'s robot-mediated role-play built children's self-efficacy in anti-bullying intervention; robotics and [[embodied-learning|embodied learning]] generally build confidence by grounding tasks in observable outcomes. - **Creative self-efficacy in children:** [[niu-genai-children-creative-thinking-cognitive-development-review-2026|Niu et al. (2026)]] report, in a systematic scoping review of 22 primary studies of generative AI with children aged 6 to 15, that gains in creative self-efficacy cluster with divergent thinking and narrative creativity, mostly through text-to-image tools that lower the barrier between idea and artifact. They stop short of claiming efficacy: the designs are heterogeneous, the review performed no critical appraisal, and the corpus is nearly silent on disability, low-connectivity and underserved learners, so creative confidence is an outcome reported to rise rather than one the field has confirmed. - **Teacher self-efficacy:** [[teacher-ai-competency|Teacher AI competency]] research examines how [[educational-development|professional development]] builds teachers' confidence in using AI, which affects adoption and integration. [[ai-supported-ementoring-efl-preservice-2026|Ismael, Luo & Li (2026)]] add quasi-experimental evidence that an AI-supported e-mentoring model raises EFL pre-service teachers' self-efficacy and emotional intelligence during the practicum: the experimental group gained substantially more than controls on both measures, with a large between-group effect and gains across all self-efficacy subdomains — and the authors attribute them to mentoring AI-mediated within structured reflective cycles rather than to the AI alone. - **Feedback and confidence:** [[ai-feedback-quality|AI feedback]] and [[intelligent-tutoring|tutoring]] can build learner self-efficacy by providing actionable, supportive feedback. - **Empowerment in AI [[problem-solving]]:** Zhu and Kong (2026) find that students' empowerment in using AI for problem solving mediates the relationship between perceived [[project-based-learning|project-based learning]] and satisfaction with an AI literacy course. In their SEM analysis of 1,027 students, PBL fostered conditions that empowered students to use AI for problem solving, which in turn drove course satisfaction — evidence that building students' confidence and capability with AI is a key mechanism of effective AI literacy education. Self-efficacy connects to [[motivation]], [[self-regulated-learning]], [[student-experience]], [[ai-literacy]], [[agency]], and [[educational-robotics]]. Building self-efficacy is a key mechanism through which AI supports engagement and learning. Self-efficacy is measured almost entirely by [[self-report-measures|self-report]], so its associations with observed behavior deserve the usual caution. - **AIGC self-efficacy as the pivot between tool and learning.** [[aigc-affordance-student-self-regulation-2026|Liang et al. (2026)]] find that perceived affordances of AI-generated content raise AIGC self-efficacy (beta = 0.583), which then mediates the paths to learning motivation (indirect effect 0.329) and to self-regulated learning (0.145), with the serial path affordance to self-efficacy to motivation to self-[[regulation]] also significant (0.173). The contrast that makes the finding useful is that the quality of AI [[assessment]] feedback did *not* predict self-efficacy (beta = 0.131, n.s.) even though it strongly predicted satisfaction — confidence with the tool is built by directing it, not by receiving good output from it. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[self-directed-learning]] - [[motivation]] - [[self-regulated-learning]] - [[student-experience]] - [[ai-literacy]] - [[agency]] - [[educational-robotics]] - [[self-report-measures]] - [[social-emotional-learning]] — Social-Emotional Learning ## Connected Articles - [[genai-performance-vs-learning]] — confidence rising while technological dependence grows: self-efficacy as a misleading AI-era outcome (Yan et al. 2025) - [[ai-supported-ementoring-efl-preservice-2026]] — AI-supported e-mentoring raises EFL pre-service teachers' self-efficacy and emotional intelligence (quasi-experimental) - [[aigc-affordance-student-self-regulation-2026]] — AIGC Affordance and Student Self-Regulation - [[oby-chatgpt-use-learning-framework-2026]] - [[genai-thoughtless-use-self-directed-learning-2026]] - [[ai-literacy-career-adaptability-business-2026]] — AI Literacy, AI Readiness, and Career Adaptability - [[remind-robot-mediated-roleplay-antibullying-2026]] — REMind - [[teacher-education-ai-literacy-sdt-2026]] — Teacher Education for AI Literacy (SDT) - [[social-robot-study-companions]] — Social Robots as Study Companions - [[hcap-human-centric-ai-pedagogy-framework-2026]] — HCAP Framework - [[learnai-just-in-time-ai-cocreation-university-2026]] — LearnAI: Just-in-Time AI Co-Creation Across Disciplines - [[student-dependency-on-ai-literacy-self-efficacy-2026]] - [[ai-advice-suppresses-ikt-suspension-2026]] - [[luo-ibl-patterns-llm-bloom-2026]] — IBL patterns in LLM-driven environments (Bloom's perspective) - [[guillen-curriculum-genai-teacher-competence-2026]] — Assessing Teacher Digital Competence for GenAI Curriculum Design (Guillén-Gámez 2026) - [[stamatoulis-genai-use-patterns-2026]] — Patterns of GenAI use and academic self-efficacy - [[ai-literacy-course-satisfaction-pbl-scale-2026]] — AI-PBLS scale; empowerment and ethical awareness mediating PBL-to-satisfaction in AI literacy courses (Zhu & Kong 2026) - [[predicting-attrition-competitive-programming]] — Predicting Student Attrition in Competitive Programming - [[genai-xr-architectural-design-education-2026]] — Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study - [[hu-psychological-predictors-continued-chatgpt-use-2026]] — AI literacy, trust and academic self-efficacy in a serial chain to continued ChatGPT use - [[suria-martinez-academic-self-efficacy-motor-disabilities-2026]] — Academic self-efficacy profiles and reported AI use among university students with motor disabilities - [[niu-genai-children-creative-thinking-cognitive-development-review-2026]] — Reported creative self-efficacy gains from generative AI in children (scoping review) --- ## [Self-Directed Learning](https://edtechdev.github.io/aied/concepts/self-directed-learning/) > **Self-directed learning (SDL)** — the process by which learners take initiative and responsibility for diagnosing their own learning needs, setting goals, identifying resources, choosing and implementing strategies, and evaluating outcomes, often with limited external structure. In the AI era, SDL is both a key outcome (does AI use support or erode learners' capacity to direct their own learning?) and a vulnerability (the convenience of generative AI can undermine the very autonomy and [[self-efficacy]] SDL requires). ## Questions to Consider - Self-directed learning means diagnosing your own needs, setting goals, and evaluating outcomes with limited external structure. Before you read, how comfortable are you actually directing your own learning — and has an AI tool ever made you *less* able to, without you noticing? - The page frames SDL as both a hoped-for outcome and a vulnerability: generative AI's convenience can erode the very autonomy and self-efficacy SDL requires. Why would a tool that gives you instant answers make it *harder* to direct your own learning later? - Thoughtless use of GenAI — adopting outputs without [[critical-thinking|critical evaluation]] — was found to harm SDL both directly and by eroding self-efficacy and motivation. Can you recall a time you accepted an AI answer without evaluating it? What, if anything, did that cost you? - SDL and self-regulated learning (SRL) are closely related but distinct: SRL concerns in-the-moment [[regulation]] of learning, while SDL concerns overarching responsibility across time. Where do you see the boundary between 'managing this task' and 'directing my own learning' in your own practice? - The harm from thoughtless AI use hit motivation harder for some students and self-efficacy harder for others. If the erosion of these psychological resources is uneven across learners, what [[equity-in-ai-education|equity]] concern does that raise about who loses the most from AI convenience? - Set a goal before you read: name one learning goal you're currently pursuing mostly on your own, and one way an AI tool helps you toward it and one way it might be quietly taking the directing away from you. ## Introduction Self-directed learning is closely related to — but distinct from — [[self-regulated-learning|self-regulated learning (SRL)]]. While SRL emphasizes the in-the-moment cognitive, motivational, and behavioral regulation of learning (planning, monitoring, controlling, reflecting), SDL emphasizes the learner's overarching responsibility for the direction and management of their own learning across time, often in informal or self-chosen contexts. SDL is foundational to [[adult-learning|adult learning]] and [[lifelong-learning|lifelong learning]], and is a prominent theory in distance and online education, where learners must sustain autonomy without scheduled class time. SDL is also increasingly *tractable* to empirical study: analyzing the clickstreams of 315 online learners who built 822 models in VERA, [[an-goel-self-directed-modeling-2026|An, Hammock & Goel (2025)]] identified three behavioral signatures of self-direction — Observation, Construction, and Exploration — and found learners progressing from hands-on construction toward fuller, hypothesis-driven Exploration while Observation persists across all phases, showing that the degree and kind of autonomy learners exercise in an unstructured online task can be distinguished from their trace data alone. ## How generative AI reshapes self-directed learning The knowledge base's [[research-methods-aied|research]] documents both sides of the GenAI–SDL relationship. - **AI can support SDL.** [[ai-lifelong-learning-policy|AI and lifelong learning]] and [[self-directed-growth-generative-ai-learning-analytics|self-directed growth with GenAI + learning analytics]] show that AI tools can [[scaffolding|scaffold]] independent inquiry, provide on-demand resources, and personalize learning paths in ways that strengthen learner autonomy. [[genai-educational-outcomes-meta-analysis|Meta-analytic evidence]] on generative AI educational outcomes and [[conversational-ai-informal-learning|conversational AI in informal learning]] suggest positive potential when AI is used as a resource the learner directs. - **Thoughtless use undermines SDL.** [[genai-thoughtless-use-self-directed-learning-2026|Zhao & Gu (2026)]] show that the **thoughtless use of GenAI** — adopting AI outputs without critical evaluation — significantly harms undergraduates' SDL both directly and through erosion of [[self-efficacy]] and [[motivation]] (the model explained 75.3% of SDL variance; TUGA β = −0.42). The negative effect on motivation was stronger for male students and on self-efficacy stronger for female students. This connects to the broader [[cognitive-offloading|over-reliance]] risk documented in the knowledge base. - **Cognitive offloading and delegation.** [[andragogy-cognitive-delegation-genai-2026|Andragogy and cognitive delegation]] and [[learning-by-chatting-genai-impact|learning by chatting with GenAI]] examine how learners may delegate cognitive work to AI in ways that bypass the effortful processing SDL requires — a failure mode of otherwise autonomy-supportive tools. [[critical-thinking-genai-scaffolding|Scaffolding critical thinking with GenAI]] and [[test-driven-ai-assisted-learning|test-driven AI-assisted learning]] model more productive designs. ## The SDL–SRL distinction in practice Because SDL emphasizes learner-initiated direction, interventions to protect it focus on preserving [[agency]] and self-efficacy rather than merely regulating moment-to-moment behavior. The evidence that thoughtless AI use erodes motivation and self-efficacy — the psychological resources SDL depends on — suggests that promoting [[ai-literacy|responsible AI use]] is not just an integrity issue but a developmental one: protecting students' capacity to direct their own learning. ## AI Extraction Scaffolding Research-Based Learning - **AI extraction as a scaffold for research-based learning.** An and colleagues (2026) design an AI-powered information extraction system that converts research publications into structured, traceable datasets to support undergraduate thesis completion in [[stem-education|STEM]], positioned as an **epistemic scaffold** that enables inspection of evidence-claim relationships while reducing low-level data-handling demands. In a 20-student [[mixed-methods-research|mixed-methods]] pilot across 80 documents, students extracted over 90% of targeted parameters, self-reported literature-review time dropped ~65%, and their ability to identify influential variables rose 50% — supporting the idea that [[generative-ai|AI]] can rebalance cognitive load toward higher-order, self-directed research reasoning rather than routine summarization. ## Connected Concepts - [[self-regulated-learning]] - [[agency]] - [[self-efficacy]] - [[motivation]] - [[metacognition]] - [[cognitive-offloading]] - [[ai-misuse-learning-harm]] - [[ai-literacy]] - [[adult-learning]] - [[lifelong-learning]] - [[higher-ed]] ## Connected Articles - [[genai-thoughtless-use-self-directed-learning-2026]] — Thoughtless GenAI use and self-directed learning (SEM, gender differences) - [[ai-lifelong-learning-policy]] — AI and lifelong learning policy - [[self-directed-growth-generative-ai-learning-analytics]] — Self-directed growth with GenAI and learning analytics - [[genai-educational-outcomes-meta-analysis]] — Meta-analysis of generative AI educational outcomes - [[andragogy-cognitive-delegation-genai-2026]] — Andragogy and cognitive delegation with GenAI - [[kim-ai-andragogy-2026]] — AI Applications in Supporting Andragogy (Kim et al. 2026) - [[ai-information-extraction-undergraduate-thesis-2026]] — AI-powered information extraction supporting undergraduate thesis and research-based learning (An et al. 2026) - [[an-goel-self-directed-modeling-2026]] --- ## [Metacognition](https://edtechdev.github.io/aied/concepts/metacognition/) > Metacognition — thinking about one's own thinking — is both a target of [[ai-education|AI education]] [[research-methods-aied|research]] (can AI tools develop students' metacognitive skills?) and a risk factor (AI completing tasks may suppress metacognitive practice).([[stanford-evidence-base-ai-k12-2026]])([[scheu-mobile-chatbot-journaling-motivation-2026]]) ## Questions to Consider - 'Metacognition' is thinking about your own thinking — knowing what you know, monitoring yourself, and adjusting your strategies. When you study or solve a problem, how aware are you in the moment of whether you actually understand versus just recognizing the material? - A striking finding: students who used AI essay assistance were often unable to recall quotes from their own essays, because they hadn't engaged with the content during production. When a tool produces the output, what practice is the learner losing — and is that practice important? - The page frames metacognition as both a target (can AI build it?) and a risk (can AI suppress it?). Could the same AI tool either strengthen or weaken a learner's metacognition depending on how it's designed or used? What determines which way it goes? - Structured prompts that ask students to self-explain, evaluate strategies, or identify gaps preserve metacognitive demand, while AI that simply completes tasks displaces it. If you were designing an AI study tool, what would you build so that it invites reflection instead of replacing it? - The page finds that whether AI use is metacognitively rich depends on the learner's motivation and stance as much as on the technology. Have you ever used a tool in a shallow way and then realized you learned nothing — and what was different about times you used it deeply? ## Introduction Metacognition in education refers to learners' awareness, monitoring, and [[regulation]] of their own cognitive processes: - **Metacognitive knowledge:** Understanding what one knows, what strategies are available, and when to deploy them - **Metacognitive regulation:** Planning, monitoring, and evaluating one's own learning in real time Within [[self-regulated-learning]] frameworks, metacognition is the central mechanism that enables learners to adapt strategies, recognize confusion, and seek help appropriately.([[scheu-mobile-chatbot-journaling-motivation-2026]]) How learners actually deploy metacognition around AI is shaped by more than the tool itself: [[cui-motivation-roles-metacognitive-genai-2026|Cui et al.]] find that student motivation and the interaction role they adopt shape their metacognitive [[student-engagement|engagement]] with [[generative-ai|GenAI]] — meaning whether AI use is metacognitively rich depends on the learner's stance as much as on the technology. [[miles-prompt-literacy-human-centered-genai-framework-2026|Miles, Haber-Curran and Arar (2026)]] add the [[student-ai-interaction|AI interaction]] itself as an object of that reflection: the closing step of their [[prompt-engineering|Prompt Literacy]] Cycle asks learners to examine what the process revealed about how prompts function and what assumptions shaped the response, and they make reflection and revision the phase in which authorship and critical judgment develop. ## How AI Tools Affect Metacognition ### The Suppression Risk (Stanford SCALE, 2026) When AI completes reasoning tasks for students — solving math problems, writing essays, generating code — the student loses practice in monitoring their own understanding and selecting strategies.([[stanford-evidence-base-ai-k12-2026]]) Key findings: - **Kosmyna et al. (2025):** Students who used AI essay assistance were **83% unable to recall quotes** from their own essays, vs. 11% for non-AI users — indicating they did not engage with the content during production. - **Stadler et al. (2024):** General-purpose AI reduced cognitive load but produced **lower-quality reasoning** vs. traditional search, suggesting metacognitive engagement was displaced. - **Lehmann et al. (2025):** General AI for [[cs-education|programming]] harmed understanding for low-[[prior-knowledge]] students — the students most in need of metacognitive scaffolding received answers instead. ### The Augmentation Opportunity (Scheu et al., 2026) When AI is designed to support reflection rather than replace it, metacognition can be strengthened: - **Learning journals** are a classic metacognitive practice: by reflecting on learning processes, students increase awareness of their cognition - **Structured prompts** that ask students to self-explain, evaluate strategies, or identify knowledge gaps preserve metacognitive demand. CoMeT (Hou et al. 2026) gives that phrase a definition and an empirical warrant: it treats metacognitive demand as a quantity distinct from [[cognitive-offloading|cognitive load]] — what the learner must decide, state, or judge before help arrives, not simply what remains when help is withheld — and held it statistically equivalent to a tutor that withheld answers by design (p_TOST = .004) while its own support escalated and faded one rung at a time. Fading held when the learner's turn was aimed at the decision under support: turns aimed elsewhere drew a later concession 40.3% of the time against 28.8% for aimed turns, an 11.5-point difference, so what a tutor must read for is where the learner's attention sits rather than how much effort the turn displays. - The **example-based course** in Scheu et al.'s [[conversational-ai|chatbot]] increased **perceived competence** (a metacognitive [[self-assessment]]) even when the [[llm]] assistant alone did not - **Surfacing interaction patterns that learners cannot see.** [[student-ai-interaction-consecutive-interpreting-2026|Kuang, Li and Weng (2026)]] tracked eye movements, note-taking and speech while 22 interpreting trainees worked with a speech-recognition and machine-translation system, and found that the way students divided [[cognitive-psychology|attention]] between AI output and their own notes was invisible to them: 58.3% changed profile between task stages, and the heaviest readers of AI output scored lowest on delivery fluency and target language quality. The pedagogical consequence is that reflection has to be scaffolded by external evidence, because a learner's strategy is not introspectable — the authors argue for guiding students to describe and evaluate why they worked a given way at each stage. ## The Engagement–Motivation Distinction Scheu et al. (2026) found a critical split: | Dimension | LLM Assistant Effect | Course Effect | |---|---|---| | **Intrinsic motivation** (willingness to engage) | **No effect** | **Positive** | | **Behavioral engagement** (amount written) | **Increasing over time** ([[feedback|feedback loop]]) | **Constant positive** | This suggests that **metacognitive support and [[motivation]] are not identical**. The LLM assistant's [[scaffolding]] of journal entries increased how much students wrote (behavioral engagement) but did not make them *want* to write more (intrinsic motivation).([[scheu-mobile-chatbot-journaling-motivation-2026]]) ## The Beliefs-vs-Experiences Distinction [[cognitive-offloading-metacognitive-review-2026|Guo & Ye (2026)]] offer a theoretically sharper account of how metacognition governs strategy selection, distinguishing two components that operate in different phases: - **Metacognitive beliefs** — stable, self-referential self-conceptions stored in long-term memory (e.g., beliefs about one's memory capability, or the reliability of a tool). These anchor strategy choices *before* task initiation. - **Metacognitive experiences** — dynamic, task-specific feelings (perceived difficulty, confidence, mental workload) that drive belief *updating* during task execution. This distinction yields the principle of **timing-component matching**: feedback that targets beliefs (e.g., comparative rankings) is most effective in the pre-task preparation phase, whereas feedback that targets experiences (e.g., immediate correctness indicators) is most effective during task execution. Abstract ranking feedback can become separated from — or overridden by — the task-specific experiences that dominate immediate decision-making, explaining why some feedback interventions fail to change behavior. This gives [[teacher-role|educators]] a phase-contingent rationale for designing metacognitive scaffolds around AI tools: calibrate beliefs before use, provide immediate task-specific feedback during use. ### Calibration is trainable: prediction + feedback [[metacognitive-training-optimal-cognitive-offloading-2026|Ngai & Gilbert (2026)]] provide direct causal evidence that metacognitive calibration is a *trainable* skill. In two preregistered experiments (N=164, N=416), **just five practice trials pairing a performance prediction with veridical feedback** improved calibration and reduced bias. A four-group additive design isolated the causal component: **making predictions alone was ineffective; adding performance feedback drove the improvement; explicitly labeling over-/underconfidence added nothing further**. Critically, the improvement acted on *absolute* calibration — raising confidence in the underconfident and lowering it in the overconfident — so it corrected [[trust-calibration|miscalibration]] in both directions rather than shifting everyone one way (which is why signed/directional effects were null). This strengthens the "experiences not beliefs" account above and shows the *minimum viable metacognitive training*: prediction + immediate, task-specific feedback. - **A brief reflection prompt sharpens monitoring during AI-supported decisions.** [[ren-metacognitive-awareness-genai-reliance-2026|Ren (2026)]] added three reflection prompts before finalizing answers in a three-condition experiment with 342 undergraduates: acceptance of incorrect ChatGPT advice fell from 62.4% to 39.7% (*OR* = 0.40) and awareness calibration rose (0.59 vs. 0.41), while recommendation accuracy and alignment with correct advice stayed high. Reflection made reliance more discriminative rather than uniformly defensive, which supports treating reliance as a monitoring problem rather than a question of how much AI is used. ## Implications for Tool Design 1. **Preserve the "friction" of thinking:** If AI writes the reflection, the student does not build metacognitive skill. Journaling assistants should scaffold, not author. 2. **Model metacognitive language:** The example-based course worked partly because it exposed students to proficient models' metacognitive self-talk. 3. **Separate support for motivation vs. skill:** Metacognitive skill development (course-structured) and productivity enhancement (AI-assisted) may require different design strategies. AI may alter the **metacognitive threshold** for deciding one knows enough to answer: [[ai-advice-suppresses-ikt-suspension-2026|Marcoccia et al. (2026)]] found that mere access to AI advice suppressed people's willingness to suspend judgment under uncertainty, even with wrong advice and accuracy incentives — an effect that survived unsolicited AI output and monetary stakes. Proactive [[agentic-ai|agentic AI]] can displace the learner's own metacognitive loop: [[agentic-ai-pedagogical-best-practice-2026|Woollaston et al. (2026)]] argue that when agents pre-fetch, initiate, and self-correct, the agent's planning, monitoring, and evaluation replace the learner's, removing the [[retrieval-spacing-interleaving|retrieval practice]] and self-monitoring that [[desirable-difficulties|desirable difficulties]] and metacognitive training depend on. - **Mistake-based [[pedagogy]] as metacognitive training:** [[pedagogy-ai-mistakes|Hosseini (2026)]] shows that deliberately exposing students to AI-generated errors in a database design course activates metacognitive monitoring — students inspected outputs, identified errors, and revised designs rather than accepting AI output at face value. [[self-report-measures|Self-reported]] [[ai-literacy|AI literacy]] correlated weakly and negatively with objective competency (*r*=−0.39), a calibration gap the critique-refinement cycle is designed to narrow. - **[[productive-failure|Productive failure]] engages metacognitive monitoring.** [[kim-ai-productive-failure-adult-2026|Kim et al. (2026)]] show productive-failure-based learning activates reflection on one's own attempts; [[lukesova-clue-before-correction-2026|clue-before-correction]] tasks require learners to diagnose and correct their own errors — a metacognitive activity where AI gives clues rather than answers. - **Self-regulation buffers offloading harm but cannot cancel it.** [[layer-sensitive-cognitive-offloading-writing-2026|Chen (2026)]] shows that metacognitive regulation (self-regulated writing) attenuates the negative association between deep [[cognitive-offloading|cognitive offloading]] and independent no-AI outcomes in GenAI-assisted writing (interaction B = 0.22), but does not eliminate it — a bounded-support condition pairing delegation limits with compulsory reflection about how AI suggestions were accepted/rejected produced the strongest independent performance. - **Explanation-seeking depth predicts task quality, not recall.** [[llm-interaction-depth-task-quality-recall-2026|Tsiligkiris (2026)]] shows explanation-seeking prompts (depth) in LLM interaction predict task quality but not immediate recall, interpreting the dissociation via elaboration (comprehension) vs. retrieval practice (consolidation) — and suggesting explanation-seeking correlates with metacognitive monitoring, though retrieval demands must be added for durable retention. - **Self-reported metacognition is a weak proxy for regulation *with* an LLM.** [[clerc-ai-literacy-workshop-llm-regulation-2026|Clerc et al. (2026)]] gave 116 [[k-12|middle-school]] students a two-hour AI literacy workshop and then measured their LLM interaction during science problems: trained students accepted underspecified prompts less often (51.5% vs. 66.7%), asked follow-up questions after a weak response far more often (59.2% vs. 27.9%, *d* = 0.80) and judged answer correctness more sensitively to prompt quality (interaction OR = 2.52). Neither a general metacognitive-awareness scale (Jr. MAI) nor GenAI self-reports predicted those behaviors or final performance (*r* = .04 and *r* = .01) — monitoring and control during [[generative-ai|generative AI]] use is task-specific, and observable behavior carries more information than the self-report instruments built to capture it. - **Automated scoring reached the performance phase, not the planning that precedes it.** [[chen-automated-scoring-interpreting-self-regulated-learning-2026|Chen and Liu (2026)]] gave 46 interpreting students 14 weeks of weekly automated scoring with a returned score, marked errors, and a reference rendition: the automated group gained more overall (*d* = 1.03), but only monitoring during practice correlated with score gains (*r* = 0.42) while pre-learning planning sat near the scale midpoint (M = 3.01). Evaluation and reflection carried the second-highest mean (3.87 of 5) yet showed a near-zero link to gains (*r* = 0.10), so a high reflection score should not be read as productive reflection. - **Verification literacy pays off only through metacognitive self-regulation.** [[davor-ai-supported-learning-higher-order-outcomes-2026|Davor, Larbi and Boateng (2026)]] surveyed 533 university students and found that AI verification literacy had no direct association with critical thinking or technical problem-solving; it mattered only indirectly, through metacognitive self-regulation (a full mediation pattern). [[cognitive-offloading|Cognitive offloading]] tendency ran the other way, predicting lower self-regulation (-.294) along with lower critical thinking (-.240) and problem-solving (-.312). - **The evaluator's own monitoring is metacognitive work too.** [[hoppe-teachers-diagnostic-skills-ai-formative-assessment-2026|Hoppe, Loibl and Leuders (2026)]] argue that an AI-generated diagnostic inference is not raw evidence but an interpretation already made, so teachers must integrate it with their own observations in a process they call *meta-diagnosis*, deciding deliberately whether to accept, reject, or modify it. That places a second metacognitive loop beside the learner's: not only how the student regulates thinking with AI, but how the teacher evaluates what the system claims about that thinking. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[self-regulated-learning]] - [[self-assessment]] - [[cognitive-offloading]] - [[scaffolding]] - [[agentic-ai]] - [[formative-assessment]] - [[ai-literacy]] - [[retrieval-spacing-interleaving]] — judgments of learning and the fluency illusion that retrieval practice corrects - [[cognitive-surrender]] ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Explainability and productive struggle as metacognitive design - [[genai-performance-vs-learning]] — the performance/learning distinction, and metacognitive laziness as offloaded evaluation (Yan et al. 2025) - [[clerc-ai-literacy-workshop-llm-regulation-2026]] — a two-hour AI literacy workshop shifted middle-school students' LLM-interaction regulation, unlike their self-reported metacognition (Clerc et al. 2026) - [[student-ai-interaction-consecutive-interpreting-2026]] — Student-AI Interaction in Computer-Assisted Consecutive Interpreting - [[du-yuan-epistemic-dependence-2026]] — Epistemic dependence in AI-mediated learning (Du & Yuan 2026) - [[pearls-epistemic-verification-2026]] — PEARLS framework for epistemic agency and verifying AI output (Wang 2026) - [[llm-interaction-depth-task-quality-recall-2026]] — What students ask matters: LLM interaction depth, task quality, and immediate recall (Tsiligkiris 2026) - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[lim-bannert-student-regulation-genai-chatbot-2026]] — How students regulate learning with a genAI chatbot - [[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026]] — LLM-mediated help-seeking in STEM: layered, instrumental, and verified - [[cui-motivation-roles-metacognitive-genai-2026]] — Motivation and roles in metacognitive GenAI engagement - [[metacognitive-training-optimal-cognitive-offloading-2026]] — Metacognitive training facilitates optimal cognitive offloading (Ngai & Gilbert 2026) - [[cognitive-offloading-metacognitive-review-2026]] — Meta-cognitive insights into cognitive offloading: mechanisms, interventions, and educational implications (Guo & Ye 2026) - [[idea-framework-metacognitive-genai-2026]] — The IDEA framework for metacognitively regulated GenAI use - [[haiml-human-centered-ai-metacognitive-model-2026]] — HAIML: a human-centered AI metacognitive learning model (agency & reflective learning) - [[metacognitively-discordant-completion-genai-2026]] — Metacognitively discordant completion and aware pass-through of non-understanding - [[ai-metacognition-stem-review]] — AI tools scaffolding metacognition in STEM - [[ai-making-us-stupid]] — Is AI making us stupid? critique of cognitive offloading - [[stanford-evidence-base-ai-k12-2026]] — General-purpose AI suppresses metacognition by completing reasoning - [[young-people-learning-generative-ai-rapid-review-2026]] — Miscalibration gap and metacognitive inequity with GenAI - [[ai-advice-suppresses-ikt-suspension-2026]] — AI advice suppresses willingness to say "I don't know", even with wrong advice and accuracy incentives - [[agentic-ai-pedagogical-best-practice-2026]] — Agentic AI and pedagogical best practice: the tension between automation and learning - [[cognitive-offloading-speedup-illusion]] — Cognitive offloading and the speedup illusion in human-AI interaction - [[lodge-loble-cognitive-offloading-2026]] — AI, cognitive offloading and implications for education (Lodge & Loble 2026) - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and cognitive surrender: how AI reshapes human reasoning (Shaw & Nave 2026) - [[pedagogy-ai-mistakes]] — The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking (Hosseini 2026) - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[lukesova-clue-before-correction-2026]] — Clue Before Correction: ChatGPT for Autonomous Language Learning - [[miles-prompt-literacy-human-centered-genai-framework-2026]] — Reflection on the prompting process and authorship development in the Prompt Literacy Cycle (Miles, Haber-Curran & Arar 2026) - [[learning-analytics-genai-secondary-writing-2026]] — Using Learning Analytics to Support Secondary School Students' Writing with Generative AI - [[adaptive-scaffolding-contingency-comet-tutor-2026]] — Adaptive Scaffolding Needs Contingency: An AI Tutor That Escalates and Fades on What the Learner Does - [[ren-metacognitive-awareness-genai-reliance-2026]] — A reflection prompt cut acceptance of incorrect AI advice and improved awareness calibration (Ren 2026) - [[chen-automated-scoring-interpreting-self-regulated-learning-2026]] — Automated scoring reinforced monitoring but not planning, and self-reported reflection stayed unproductive (Chen & Liu 2026) - [[davor-ai-supported-learning-higher-order-outcomes-2026]] — Verification literacy acting only through metacognitive self-regulation (Davor, Larbi & Boateng 2026) - [[hoppe-teachers-diagnostic-skills-ai-formative-assessment-2026]] — From diagnosis to meta-diagnosis: teachers judging AI-generated inferences (Hoppe, Loibl & Leuders 2026) --- ## [Desirable Difficulties](https://edtechdev.github.io/aied/concepts/desirable-difficulties/) > **Desirable difficulties** — the finding (Bjork) that harder, effortful retrieval conditions — spacing, [[retrieval-spacing-interleaving|retrieval practice]], interleaving, and generation — improve long-term learning more than easier, massed conditions — is the theoretical counterweight to AI that smooths away [[cognitive-offloading|cognitive work]]. In the AI era the principle warns that tools which eliminate productive struggle may raise immediate performance while undercutting durable learning. **Desirable difficulties, cognitive friction, and productive friction are used as overlapping synonyms** for this intentional effort: the knowledge base treats them as the same core idea viewed from different fields, with the nuances between the labels spelled out in the section below. Closely allied concepts — **confusion**, and **productive struggle** — mark the zone where this effortful processing is expected (and desirable) to occur. ## Questions to Consider - Have you ever felt you understood something because it felt easy and fluent in the moment — only to fail when you had to recall it later? That's the illusion of competence. What created it for you? - Desirable difficulties say that effortful conditions — spacing, retrieval practice, interleaving — build durable learning better than easy, massed ones. Where in your own learning have you resisted a 'harder' strategy that probably would have worked better? - Generative AI is, by default, a friction-removing machine: it answers instantly and produces polished output on demand. If removing struggle raises immediate performance but undercuts durable learning, how would you know whether an AI is helping or harming a student? - This page distinguishes desirable difficulties (memory optimization from [[cognitive-psychology|cognitive psychology]]) from productive/cognitive friction ([[student-engagement|engagement]] [[guardrails]] from UX design). Can you see why the same educational goal needs both — and where they'd diverge? - Some [[intelligent-tutoring|AI tutors]] are found to 'over-scaffold' — removing the very effortful processing desirable difficulties require. If you were evaluating an AI tutor, what concrete behavior would tell you it's preserving productive struggle rather than collapsing to answer-giving? - Confusion is framed here as a resource, not a bug — when resolved productively it drives deep processing, but unaddressed it decays into frustration. Where's the line between productive struggle worth preserving and frustration that's just harmful? ## Introduction Desirable difficulties are the conditions of practice that make learning feel harder in the moment — effortful retrieval, generation and explanation, spacing, interleaving — yet produce stronger retention and [[transfer-of-learning]] than conditions that feel easy. The same phenomenon appears in the literature as **cognitive friction** and **productive friction**, labels borrowed from human–computer interaction and UX design that emphasize deliberately placing resistance between a learner and an easy answer; the families overlap but are not identical, and the differences are set out below. Because generative AI is optimized to be frictionless, the concept has become a first-order design concern rather than a niche finding: systems that answer instantly remove difficulty that may have been doing the learning ([[cognitive-offloading]], [[scaffolding]]). ## The Effort–Learning Trade-Off Desirable difficulties rest on the insight that conditions that make learning feel harder in the moment — requiring effortful retrieval, generation, or explanation — frequently produce stronger retention and [[transfer-of-learning|transfer]] than conditions that feel easy. Conversely, conditions that feel easy (fluent presentation, immediate answers) can produce an illusion of competence: [[learners]] feel they know the material because recognition was smooth, while later free recall fails. This is the theoretical core of the **performance–learning gap**: what looks like good performance during practice is not the same as durable learning. ## Confusion, Cognitive Friction, and Productive Struggle Three related constructs describe the zone in which desirable difficulties operate: - **Confusion** — a learning *epistemic emotion* (see [[affective-computing]] and [[epistemic-emotions-collaborative-problem-solving]]) that signals a gap between a learner's mental model and incoming information. Confusion is not uniformly bad: when resolved through productive inquiry it can drive deep processing, but when unaddressed it can decay into frustration or disengagement. AI systems increasingly detect confusion (e.g. capture buttons, affective sensing) to anchor [[personalized-learning|personalized support]] — as in [[knowloop-confusion-to-consolidation-2026]], where marked confusion points become review anchors and [[learning-by-teaching|teach-back]] prompts surface conceptual gaps. - **Cognitive friction** — the deliberate resistance a learning environment places between a learner and an easy answer, forcing them to think before receiving help. AI tools that answer instantly remove this friction; designs that withhold, hint, or scaffold preserve it. [[generative-refusal-ai-tools-for-thought]], [[sequenced-ai-feedback-learning]], and [[critical-thinking-genai-scaffolding]] each examine how intentionally preserved friction supports reasoning. - **Productive struggle** — the effortful phase of [[problem-solving|problem solving]] in which a learner wrestles with a challenge before (or while) receiving support. The knowledge base's evidence base documents both its value and its cost: [[generative-ai-reduced-study-time-math]] shows removing struggle reduced study time but impaired learning, while [[curiobot-llm-tutoring-exploratory-learning]] and [[rethinking-scaffolding-llm-tutors]] explore how tutors can keep learners in the productive-struggle zone rather than collapsing to answer-giving. Productive failure is the structured, theory-driven version of this idea: [[productive-failure|Kapur's productive failure]] (PF) formalizes productive struggle as a two-phase design (generation & exploration *before* instruction, then consolidation & knowledge assembly). The AI-era PF literature gives the knowledge base a concrete design vocabulary for preserving desirable difficulty — [[kim-ai-productive-failure-adult-2026|Kim et al. (2026)]] derive AI design principles ([[human-ai-collaboration|human-AI collaboration]], reflective design, non-directive support) that keep AI from erasing the struggle; [[puech-pedagogical-steering-llm-productive-failure-2025|Puech et al. (2025)]] show [[llm]] tutors can be steered to withhold solutions and elicit multiple attempts; [[wang-safety-gap-productive-struggle-2026|Wang & Shan (2026)]] formalize the "Safety Gap" — the divergence between AI-assisted performance and unassisted capability — as the cost of removing struggle; and [[rhaimi-productivemath-2025|ProductiveMath]] uses AI to lower the burden of designing PF problems. These show that desirable-difficulty principles translate into concrete AI design choices. ## Desirable Difficulties vs. Cognitive Friction vs. Productive Friction Because AI is designed to be frictionless — instantly generating summaries, solving equations, and writing essays — it can inadvertently bypass the very struggle required for a student to learn. To combat this, educators and technologists rely on two overlapping but distinct frameworks: **desirable difficulties** and **productive (or cognitive) friction**. Both advocate making things harder for the learner, but they originate from different fields and target different parts of the learning process. In this knowledge base they are treated as synonyms for the same intentional-effort idea; the table below details the nuance between the labels. | Feature | Desirable Difficulties | Productive / Cognitive Friction | |---|---|---| | Primary goal | Maximizing long-term memory and knowledge transfer | Preventing [[cognitive-offloading]] and maintaining [[active-learning|active engagement]] | | Scientific root | Cognitive science & psychology (Bjork, 1994) | Human–Computer Interaction (HCI) & UX design | | The "threat" | The illusion of competence (thinking you know it because it feels easy now) | Automation bias (letting the machine do the thinking for you) | | AI implementation | Algorithms that time and structure practice (spacing, interleaving, retrieval) | [[conversational-ai|Chatbot]] guardrails and UI roadblocks that force the learner to do the work | **Desirable difficulties: the memory optimizer.** Coined by Robert and Elizabeth Bjork (1994), this framework comes from [[cognitive-offloading|cognitive psychology]]. Its core idea is that learning strategies which feel harder and slow initial performance actually produce better long-term retention and [[transfer-of-learning|transfer]]. Desirable difficulties are about *how the brain encodes and retrieves information*: if learning feels too easy or fluent in the moment (like re-reading a highlighted textbook), the brain likely isn't doing the deep processing required to make the memory stick. In AI, a tool using this framework changes the *[[pedagogy]]* of the session — for example, asking the student to retrieve from memory before offering a summary (retrieval practice), scheduling review just before forgetting (spacing), or mixing problem types (interleaving) rather than grouping them by category. Notably, the benefit of these effortful strategies is itself content-dependent: [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit, Koedinger & Carvalho (2025)]] show that retrieval practice chiefly strengthens verbatim memory (by delaying forgetting), whereas acquiring a generalizable skill requires integrating worked examples with practice — so the "difficulty" that helps must be matched to the type of knowledge being learned rather than applied uniformly. **Productive (cognitive) friction: the engagement guardrail.** This framework comes from UX and interaction design, where "friction" is normally the enemy (one-click checkout, instant search). In educational technology, zero friction means zero thinking: productive friction introduces intentional "speed bumps" into the software to prevent the user from offloading cognition to the machine. It is about the *interaction between human and machine*, keeping the user actively engaged and preventing automation bias — blindly trusting the AI's output without evaluating it. In AI, a tool using this framework changes its *behavior and design* to prevent shortcuts — for example, a [[socratic-method|Socratic]] guardrail that withholds the direct answer and asks what symbols the student noticed, effort checkpoints that refuse to generate a draft until a thesis and outline are entered, or delayed [[feedback]] that requires committing to an answer and explaining reasoning before the solution is revealed. **In short:** you use productive friction to ensure the student actually interacts with the material instead of letting the AI do the heavy lifting; you use desirable difficulties to structure *how* they interact with that material so they remember it a month from now. ## Desirable Difficulties in the AI Era The central tension for AI-supported learning is that [[generative-ai|generative AI]] is, by default, a friction-removing technology: it answers, generates, and produces polished artifacts on demand. Across the knowledge base, this plays out in two directions: - **The cost of removing struggle.** When AI erases spacing, retrieval, and generation, learners may show immediate performance gains but forfeit durable learning and transfer. This connects directly to the [[cognitive-offloading|Over-Reliance]] and [[ai-misuse-learning-harm]] findings: an AI that removes desirable difficulty produces the performance–learning gap documented across the knowledge base's evidence base. [[agentic-ai-pedagogical-best-practice-2026]] calls explicitly for intentional friction. - **Designing struggle back in.** Instructional designs can deliberately preserve productive processing: draft-first routines, hint-not-answer tutoring, delayed feedback, and teach-back/explanation protocols. These are the concrete scaffolds explored under [[reducing-ai-misuse]] and [[structured-llm-feedback-programming]]. **The inverted U and the effort paradox.** [[zohar-bloom-inzlicht-against-frictionless-ai-2026|Zohar, Bloom and Inzlicht (2026)]] supply the sharpest recent statement of why AI's friction-removal is not automatically good. They distinguish AI from earlier labor-saving [[ai-technologies|technologies]] on two grounds: it targets intellectual and creative work rather than physical or clerical work, and its friction removal is *extreme* — prior technologies eliminated excess friction, "tedious or insurmountable obstacles that offer little benefit for learning or meaning", whereas a chatbot lets a learner move from ideation to evaluation "without exerting meaningful effort, without questioning the output, and without engaging the cognitive processes that foster ownership, retention, or critical thought". Their organizing claim is that the effort–meaning relationship is curvilinear: moderate friction enhances meaning and motivation while excessive friction overwhelms, so AI's risk is overshooting into too little friction rather than excess. Two consequences matter pedagogically — effort is itself a trainable skill (rewarding process rather than product increases the tendency to strive and persevere), and the motivational benefits of effort erode in exactly the domains where AI substitutes for it, producing a cycle of increasing dependence ([[cognitive-offloading]], [[motivation]]). ## Design Implications 1. **Do not optimize for effort-free fluency.** An AI tutor that always answers immediately may raise satisfaction while lowering durable learning; favor interventions that require retrieval and generation first. 2. **Treat confusion as a resource, not a bug.** Detect and target confusion points as personalized review anchors rather than smoothing them away — the KnowLoop Recognize→Resolve→Consolidate model is a concrete pattern. 3. **Preserve cognitive friction deliberately.** Use hint-not-answer [[scaffolding]], sequential feedback, and refusal-to-answer where the goal is reasoning, not production. 4. **Match friction to learner readiness.** Desirable difficulties benefit learners who can engage in effortful processing; over-challenge without support risks frustration. [[scaffolding]] must keep learners in the productive-struggle zone, not past it. TutorMoments operationalizes desirable-difficulty principles as evaluation criteria: [[zhang-tutormoments-2026|Zhang et al. (2026)]] test whether AI tutors preserve productive struggle by scaffolding for access (when needed) and pushing for rigor (when ready), and find that LM tutors default to over-scaffolding — removing the effortful processing that desirable difficulties require. ## Connected Concepts - [[learning-by-teaching]] - [[self-regulated-learning]] - [[metacognition]] - [[transfer-of-learning]] - [[scaffolding]] - [[learning-gains]] - [[cognitive-offloading]] - [[ai-misuse-learning-harm]] - [[reducing-ai-misuse]] - [[affective-computing]] - [[active-learning]] - [[constructivist]] - [[motivation]] - [[learning-theories]] - [[productive-failure]] — Productive Failure - [[retrieval-spacing-interleaving]] — the operational techniques that instantiate this principle: testing effect, spacing, interleaving ## Connected Articles - [[zohar-bloom-inzlicht-against-frictionless-ai-2026]] — Against frictionless AI: the inverted-U argument for preserving beneficial friction - [[evaluation-age-ai-output-evidence-2026]] — Evaluation in the Age of AI - [[critical-thinking-paradox-genai-learning-2026]] — The critical-thinking paradox in GenAI-integrated learning - [[brcic-effortless-trap-productive-struggle-2026]] — Six-move model of learning and AI placement (Brcic & Frljic 2026) - [[agentic-ai-pedagogical-best-practice-2026]] - [[finkelstein-principled-ai-education-2025]] - [[structured-llm-feedback-programming]] - [[generative-ai-reduced-study-time-math]] - [[curiobot-llm-tutoring-exploratory-learning]] - [[rethinking-scaffolding-llm-tutors]] - [[knowloop-confusion-to-consolidation-2026]] - [[generative-refusal-ai-tools-for-thought]] - [[sequenced-ai-feedback-learning]] - [[critical-thinking-genai-scaffolding]] - [[epistemic-emotions-collaborative-problem-solving]] - [[stanford-evidence-base-ai-k12-2026]] — Tutoring-specific AI preserves productive struggle vs. general-purpose chatbots - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[young-people-learning-generative-ai-rapid-review-2026]] — Productive friction built into GenAI tools supports learning - [[zhang-tutormoments-2026]] — When Help is Unhelpful: evaluating AI tutors for productive struggle - [[lodge-loble-cognitive-offloading-2026]] — AI, cognitive offloading and implications for education (Lodge & Loble 2026) - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[wang-safety-gap-productive-struggle-2026]] — The Safety Gap: Restoring Productive Struggle - [[rhaimi-productivemath-2025]] — ProductiveMath: AI to Support PF Problem Design - [[making-ai-annoying-constrained-writing-2026]] — Making AI annoying on purpose: constraint in AI-supported writing (Konradt, Boote & Taub 2026) - [[rachatasumrit-example-problem-ratio-2026]] --- ## [Transfer of Learning](https://edtechdev.github.io/aied/concepts/transfer-of-learning/) > **Transfer of Learning** — the extent to which knowledge or skills acquired in one context (e.g., practice with an AI tool) persist and apply in a different context (e.g., independent performance without the tool). In [[ai-education|AI in education]], transfer is the central open question: whether performance gains students show *with* AI tools translate into durable learning they can demonstrate *without* them. ## Questions to Consider - Here's a striking pattern the page documents: students often show immediate gains on AI-assisted tasks, yet those gains can vanish — or even reverse — when the AI is removed. Before reading the explanations, why do you think a tool that clearly helps in the moment could end up leaving students worse off without it? - Recall something you learned to do with a tutor, calculator, or assistant and then had to do alone. Did the skill carry over, or did you feel dependent on the aid? What was different about the experiences that transferred well versus those that didn't? - A common intuition is that 'practice is practice' — that doing a task with help builds the same skill as doing it alone. Where might that intuition mislead, especially when the help is an AI that completes the reasoning for you rather than guiding you through it? - The page draws a distinction between 'effects with' a technology and 'effects of' it — performing better while using the tool versus becoming more capable without it. If you're an instructor, designer, or student, which of these is your real goal, and how would you know you'd achieved it? - The evidence suggests that how much cognitive work you delegate matters: offloading surface tasks like grammar hurt transfer less than offloading deep reasoning and structure. Think about the last time you used AI on an assignment. Which 'layer' did you delegate, and what does your choice predict about what you'd retain? - The page proposes conditions that might support positive transfer — [[pedagogy|pedagogical]] [[guardrails]], fading support, calibration to the learner's readiness. If you were designing (or were the user of) an AI learning tool, what would you insist on so that gains while using it become durable ability without it? ## Introduction Transfer of learning is a foundational concern in education [[research-methods-aied|research]], and AI tools have made it urgent. The defining empirical pattern documented across AI in education studies is a **transfer paradox**: students using AI typically show immediate, measurable gains on tasks where AI is available, but those gains often fail to persist — or even reverse — when AI is removed and students must demonstrate understanding independently. This pattern implicates [[cognitive-offloading|Over-Reliance]], Cognitive Load Theory, and [[metacognition]] as the mechanisms at work, and connects directly to debates about [[intelligent-tutoring|AI Tutoring]] design. ### The transfer paradox Students using AI typically show **immediate, measurable gains** on the tasks where AI is available. Yet when AI is removed: - Effects become **mixed or negative** - Gains often **fail to transfer** to unassessed settings - Students may become **dependent on the tool** at the expense of independent reasoning The evidence base, synthesized in the [[stanford-evidence-base-ai-k12-2026|Stanford Evidence Base on AI in K-12]] review, is consistent across domains: | Study | Context | Immediate Effect | Transfer Effect | Mechanism | |---|---|---|---|---| | Bastani et al. (2025) | High school math | Higher practice grades | **~17% worse** on closed-book finals | General-purpose chatbot did the work | | Chen et al. (2025) | Programming homework | Higher homework scores | **No improvement** on unassisted exams | [[llm]]-Tutor solved problems for students | | Lehmann et al. (2025) | Programming | More topics covered | **Harmed understanding**; widened gaps | General AI for low-prior learners | | Stadler et al. (2024) | Academic research | Faster task completion | **Lower-quality reasoning** vs. search | Reduced cognitive [[student-engagement|engagement]] | | Kosmyna et al. (2025) | Essay writing | Higher essay quality | **83% failed to recall** their own quotes | Outsourced authorship | All five studies show a **negative or null transfer** pattern when general-purpose AI is the intervention. ### Mechanisms undermining transfer **Metacognitive displacement.** AI completing reasoning reduces opportunities for students to monitor their own understanding and select strategies. Students who used AI were less able to explain their answers when queried. This connects to [[metacognition]] research on self-monitoring and the [[vibe-compiler-metacognition-genai-agency-2026|evidence that structured courses increase metacognitive competence while raw LLM assistants do not]]. **Germane load suppression.** General-purpose AI reduces not just extraneous (distracting) cognitive load but also *germane* load — the productive mental effort that encodes durable knowledge. Easier practice feels better but stores weaker traces. See Cognitive Load Theory and the distinction between [[stanford-evidence-base-ai-k12-2026|tutoring-specific vs general AI]]. **Over-reliance / expertise reversal.** Novices given answers do not build schemas. General AI provides answers; effective tutoring provides structured guidance. When novices are given expert-level shortcuts, learning is disrupted — the [[desirable-difficulties]] principle in reverse. **Tool-dependent performance.** Students may optimize for the specific affordances of the AI tool ([[prompt-engineering|prompt engineering]], reliance on generated code structure) rather than building domain generalization — a form of [[cognitive-offloading-speedup-illusion|cognitive offloading]] that feels productive but displaces durable learning. **Layer-sensitive offloading and transfer.** [[layer-sensitive-cognitive-offloading-writing-2026|Chen (2026)]] directly tests Salomon, Perkins & Globerson's "effects with vs. effects of technology" distinction in [[generative-ai|GenAI]]-assisted writing: an eight-week quasi-experiment found open AI collaboration maximized supported-writing performance but produced the *lowest* independent no-AI near-transfer outcomes, while bounded support with reflection preserved independent competence. Deeper offloading layers (reasoning, structure) predicted worse transfer than surface layers (grammar). This is direct classroom evidence that AI's *with*-support performance gains do not transfer to *of*-support independent performance — and that the depth of delegation, not just whether AI is used, shapes transfer. A complementary, if confounded, instance comes from [[physics-education|physics]]: the Ruhr University Bochum redesign of an introductory nuclear and particle physics course ([[ai-particle-physics-education-redesign-2026|Mikhasenko et al., 2026]]) had students successfully complete collaborative, resource-rich research problems with AI assistance, yet those same students averaged 20.6/80 on a conventional unaided written exam, with several serious attempts unable to complete standard calculations. The authors read this as evidence that assisted performance does not automatically transfer to unprompted performance, and their remedy is deliberate design: making the written exam the sole grade determinant, releasing tutorial problems in advance so class time becomes prepared discussion, and adding prerequisite preparation, worked examples and consolidation around the exploratory AI-permitted work. ## Conditions supporting positive transfer The limited evidence suggests transfer is possible when: - **Pedagogical guardrails are present** — step-by-step hints, [[misconceptions|misconception]] targeting, [[socratic-method|Socratic questioning]] (Bastani et al., 2025 tutoring variant) - **Traditional strategies are preserved** — note-taking paired with AI use improved retention (Kreijkes et al., 2026) - **AI is used for [[formative-assessment|formative]], not [[summative-assessment|summative]], practice** — scaffolding during learning, not during assessment - **The practice format is matched to the knowledge being transferred.** [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit, Koedinger & Carvalho (2025)]] find that retrieval-practice gains frequently fail to transfer to unfamiliar problems — they strengthen memory for a procedure without enabling its use in new contexts — and that durable generalization to novel applications requires pairing practice with worked examples that support skill induction; the optimal example–problem ratio therefore depends on whether the content is a verbatim fact or a generalizable skill. - **Learner expertise is calibrated** — the tool adapts support to readiness rather than defaulting to full assistance - **Transfer as the criterion that separates learning from assistance.** [[yan-agentivism-learning-theory-ai-2026|Yan and Gašević (2026)]] build their theory of human-AI learning around transfer under reduced support: assisted performance counts as learning only if the capability persists once the support is withdrawn, which makes transfer the test rather than one outcome among several. Their proposition is directional, that requiring source checking or justification during AI-supported work should improve delayed performance, while repeated low-friction delegation without reconstruction should weaken learners' calibration of their own competence. This aligns with [[intelligent-tutoring|AI Tutoring]] research showing that tutoring-specific tools with pedagogical guardrails outperform general-purpose [[conversational-ai|chatbots]], and with [[scaffolding]] principles about fading support as competence grows. ### Unanswered questions 1. **Time scale:** Does transfer improve over weeks/months of use, or does dependence deepen? 2. **Domain differences:** Is transfer better in well-structured domains (math) vs. ill-structured domains (writing)? 3. **Individual differences:** Do high-[[prior-knowledge]] students suffer less transfer loss than novices? 4. **Skill remediation:** Can explicit "AI-off" practice sessions reverse tool dependence? ### Connections to related concepts Transfer of learning connects to [[metacognition]] (self-monitoring of understanding), Cognitive Load Theory (germane vs extraneous load), [[desirable-difficulties]] (productive struggle), [[scaffolding]] (fading support), [[cognitive-offloading|Over-Reliance]] (tool dependence), and [[sociocultural-learning]] (general-purpose AI operates outside the ZPD by completing work for students). It is the bridge between assisted performance and genuine learning — the distinction between [[stanford-evidence-base-ai-k12-2026]] and the central question for [[intelligent-tutoring|AI Tutoring]] effectiveness. ## Connected Concepts - [[metacognition]] - [[desirable-difficulties]] - [[cognitive-offloading]] - [[scaffolding]] - [[sociocultural-learning]] - [[intelligent-tutoring]] - [[k-12]] - [[self-regulated-learning]] - [[learning-theories]] - [[productive-failure]] — Productive Failure ## Connected Articles - [[yan-agentivism-learning-theory-ai-2026]] — A mid-range learning theory for human-AI interaction, with four mechanisms and six testable propositions (Yan and Gašević 2026) - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[critical-thinking-paradox-genai-learning-2026]] — The critical-thinking paradox in GenAI-integrated learning - [[stanford-evidence-base-ai-k12-2026]] - [[educational-llm-alignment]] - [[cognitive-offloading-speedup-illusion]] - [[vibe-compiler-metacognition-genai-agency-2026]] - [[hazra-safetutors-pedagogical-safety-2026]] - [[learnity-graphs-lifelong-learning-framework-2026]] - [[genai-assisted-problem-posing-physics-2026]] - [[young-people-learning-generative-ai-rapid-review-2026]] — Performance-learning distinction and durable transfer - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[rachatasumrit-example-problem-ratio-2026]] - [[ai-particle-physics-education-redesign-2026]] — AI in Particle Physics Education: Research Problems and Foundational Skills - [[shi-genai-experiential-learning-management-education-2026]] — argues that protected classroom simulations can form decision habits that fail outside them --- ## [Prior Knowledge](https://edtechdev.github.io/aied/concepts/prior-knowledge/) > **Prior knowledge** — the existing knowledge, skills, beliefs, and mental models a learner brings to a new learning task. It is the single most powerful predictor of subsequent learning: new information is interpreted through — and integrated with — what the learner already knows, so instruction that activates and builds on prior knowledge produces stronger, more durable learning than instruction that treats every learner as a blank slate. In [[ai-education|AI in education]], prior knowledge is central to [[student-modeling]] (adapting [[personalized-learning|instruction]] to the learner's current state), to the [[constructivist]] principle that knowledge is actively constructed atop existing mental models, and to the risk that AI tools which pre-fetch and surface content bypass the [[retrieval-spacing-interleaving|retrieval practice]] that activates prior knowledge. ## Questions to Consider - What's something you learned deeply and something you struggled to learn? How much of the difference came down to what you already knew when you started? - Having prior knowledge isn't enough—it must be actively retrieved and connected. When has recalling what you already knew (or failing to) changed how well you learned something new? - The page says prior knowledge can *interfere* when it's wrong (a misconception). Can you think of a belief you held that made new, correct information harder to learn? - Generative AI that pre-fetches answers can bypass the retrieval practice that activates prior knowledge. How might a tool designed to help you learn actually prevent you from recalling what you know? - If an AI must estimate your prior-knowledge state to personalize, what happens when that estimate is wrong? How confident are you that a system could accurately know what you already know? - How is 'activating prior knowledge' different from simply asking students a question before [[teacher-role|teaching]]? What would make that activation genuinely deepen the learning that follows? ## Introduction Prior knowledge activation is one of the most robust findings in [[learning-sciences|the learning sciences]]: learners do not absorb new material in a vacuum but map it onto existing schemas, and the quality of that mapping determines retention and [[transfer-of-learning|transfer]]. The concept underpins Ausubel's advance organizers, activation of prior knowledge before new instruction, retrieval practice as a form of activating and strengthening what is known, and diagnostic [[assessment]] of what learners already know. In the AI era, prior knowledge has taken on new urgency because [[generative-ai|generative AI]] can either *support* activation ([[prompt-engineering|prompting]] learners to recall and connect what they know) or *bypass* it entirely (instantly supplying an answer or pre-fetched content that the learner never had to retrieve or integrate). ## The role of prior knowledge - **It is the strongest predictor of learning.** Decades of [[research-methods-aied|research]] show that what a learner already knows correlates with [[learning-gains|learning outcomes]] more strongly than almost any other factor, because new information is encoded relative to existing mental models. AI systems that adapt to each learner's prior-knowledge state therefore hold particular promise for efficiency and [[transfer-of-learning|transfer]]. - **Activation matters, not just possession.** Having prior knowledge is not enough — it must be actively retrieved and connected to the new material. This is why "activating prior knowledge" is a standard [[pedagogy|instructional]] move, and why retrieval practice (recalling what you know before adding to it) improves learning beyond simple re-exposure. - **It shapes interpretation.** Learners interpret new information through what they already believe. When those beliefs are wrong ([[misconceptions]]), prior knowledge can *interfere* with learning, which is why instruction must surface and address misconceptions rather than assume a neutral starting point. - **It drives student modeling.** To personalize, an AI system must estimate the learner's prior-knowledge state — the basis of [[knowledge-tracing]], student modeling, and adaptive [[scaffolding]]. The quality of these estimates determines whether adaptation is genuinely helpful or misleading. - **Knowledge content determines which process practice recruits.** Whether learning hinges on memory or on induction is set by the prior-knowledge structure of the target: [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit, Koedinger & Carvalho (2025)]] follow the [[learning-theories|KLI]] framework in distinguishing knowledge components with constant conditions and responses (facts, acquired through memory and retrieval practice) from those with variable conditions and responses (skills, acquired through induction and generalization to novel inputs) — which is why the optimal mix of worked examples and practice differs for fact content versus skill content. ## Prior knowledge in the AI era Generative AI has made prior knowledge a central design consideration rather than a background variable: - **The bypass risk.** [[agentic-ai-pedagogical-best-practice-2026|Proactive agentic AI]] that pre-fetches and surfaces content can bypass the retrieval practice that activates prior knowledge — the learner never has to recall or integrate what they know before receiving an answer. This is one of the six pedagogical risks identified in the [[agentic-ai|agentic]]-education best-practice framework, and it connects directly to [[cognitive-offloading|Over-Reliance]] and the [[desirable-difficulties]] principle that effortful processing supports durable learning. - **Priming and activation as design.** [[genai-mindtool-generative-learning|GenAI mindtool approaches]] deliberately "prime the learning task" by activating prior knowledge and curiosity through prompting questions, AI-generated visuals, and analogies (e.g., "What do you already know about ecosystems?") before introducing new content — modeling the retrieval-and-integration path rather than the answer-supply path. - **Student modeling and memory.** AI systems increasingly model learners' prior-knowledge state and longitudinal memory (e.g., incorporating prior-knowledge state and forgetting curves into tutoring memory), enabling spaced repetition and adaptive review that build on what each learner already knows.([[nie-personavlm-long-term-personalization-2026]]) - **A personalized-adaptation lever.** Because learners differ widely in prior knowledge, adaptation must be tuned to the individual — a core argument for [[personalized-learning]] and adaptive [[scaffolding]] that meet learners at their actual current state rather than a class-average assumption. ## Implications for designing AI in education 1. **Activate before you supply.** Design AI interactions to prompt learners to retrieve and articulate what they already know before providing new content or answers — preserving retrieval practice rather than bypassing it. 2. **Model the learner's prior-knowledge state.** Build student modeling and adaptation on estimated prior knowledge (and its misconceptions), not on assumed uniformity, to make personalization genuinely responsive. 3. **Surface and address misconceptions.** When prior knowledge is incorrect, it will interfere; instruction should elicit and correct misconceptions rather than add new content on top of faulty foundations. 4. **Weigh the friction trade-off.** Activating prior knowledge adds desirable difficulty (retrieval, integration) that AI's friction-removing defaults tend to erase — a tension to manage deliberately rather than let automation resolve by default. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[constructivist]] - [[personalized-learning]] - [[student-modeling]] - [[misconceptions]] - [[icap-framework]] - [[knowledge-tracing]] - [[scaffolding]] - [[self-regulated-learning]] - [[transfer-of-learning]] - [[metacognition]] - [[cognitive-offloading]] - [[desirable-difficulties]] - [[learning-theories]] - [[productive-failure]] — Productive Failure - [[retrieval-spacing-interleaving]] — how what a learner already knows determines what retrieval practice can do ## Connected Articles - [[agentic-ai-pedagogical-best-practice-2026]] — The tension between automation and learning (prior knowledge activation risk) - [[genai-mindtool-generative-learning]] — GenAI as a mindtool: priming and activating prior knowledge - [[nie-personavlm-long-term-personalization-2026]] — LLM student modeling and memory - [[critical-thinking-paradox-genai-learning-2026]] — The critical-thinking paradox in GenAI learning - [[lodge-loble-cognitive-offloading-2026]] — Lodge & Loble on cognitive offloading - [[cognitive-offloading-llm-synthesis-writing]] — Cognitive offloading in LLM synthesis writing - [[bridging-instructional-design-framework-math]] — An instructional-design framework for math - [[chudziak-ai-math-tutoring-platform]] — AI math tutoring platform - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[rachatasumrit-example-problem-ratio-2026]] --- ## [ICAP Framework](https://edtechdev.github.io/aied/concepts/icap-framework/) > **The ICAP Framework** (Interactive–Constructive–Active–Passive) — a taxonomy of cognitive engagement developed by Michelene Chi that classifies learner behavior into four modes of knowledge change, ordered from least to most cognitively engaged: *passive*, *active*, *constructive*, and *interactive*. In AI in education, ICAP provides both a design target (build tools that elicit constructive and interactive engagement rather than passive consumption) and an evaluation lens (measure whether learners and AI systems are actually engaged at the higher modes).([[hingle-collaborative-ai-literacy-2025]])([[icap-cognitive-engagement-llm-agents]]) ## Questions to Consider - Think of the last time you 'learned' something by watching a video or reading. The ICAP framework would call that passive. What do you actually retain from passive exposure versus from explaining it to someone else? - ICAP orders engagement from passive to active to constructive to interactive. Where do AI tools you've used tend to keep learners — and does clicking through adaptive practice count as real engagement or just activity? - The page argues the most consequential shift is from Active to Constructive — generating explanations or new artifacts rather than just applying knowledge. Why might producing something new be the step that actually changes understanding? - One study found human experts far outperform AI models at labeling engagement levels. If automated systems systematically underestimate engagement, how should we treat 'engagement' metrics generated by AI? - ICAP shows that a tool that answers for you keeps you passive, while one that prompts and questions pushes you toward constructive and interactive engagement. Which design choice would you make for your learners? - The framework is used both as a design target and an evaluation lens. How could you use ICAP in your own teaching or design to tell whether learners are genuinely engaged rather than merely active? ## Introduction ICAP is grounded in the assumption that *what learners do* determines how much and what they learn. Chi's framework posits that as engagement moves from passive to active to constructive to interactive, the nature of knowledge change deepens — from storing, to attending, to integrating new knowledge with prior knowledge, to co-creating knowledge through dialogue. This makes ICAP a powerful analytic tool for AI in education, where the central design question is whether AI assistance supports or displaces learners' cognitive engagement. ## The four modes | Mode | Learner behavior | Nature of knowledge change | |------|------------------|---------------------------| | **Interactive** | Dialogue with another learner or agent, co-constructing meaning; e.g. defending a position, [[collaborative-learning|collaborative problem-solving]] | Co-creating new knowledge through joint, reciprocal activity | | **Constructive** | Generating new output beyond the given; e.g. self-explaining, comparing, reflecting, drawing | Integrating new information with prior knowledge to produce novel understanding | | **Active** | Manipulating or acting on the material; e.g. taking notes, underlining, pausing to think | Attending to and storing information, sometimes without deep integration | | **Passive** | Receiving information without overt action; e.g. listening to a lecture, reading | Storing information, with limited further processing | ## ICAP in AI in education ### A design target for AI tools ICAP reframes the central design question for AI in education: an AI tool that *answers for* the learner keeps them in passive/active modes, while a tool that *prompts, questions, and [[scaffolding|scaffolds]]* can push learners toward constructive and interactive engagement. This aligns ICAP with [[constructivist]] pedagogy and with [[active-learning]] research.([[multimodal-learning-genai]])([[hingle-collaborative-ai-literacy-2025]]) ### An evaluation lens for AI agents ICAP also serves as a measurement framework. In one study, researchers extended ICAP to a 7-point scale to characterize cognitive engagement in collaborative dialogue, then compared trained human annotators with LLM-based labeling (in-context learning, zero-shot prompting, and reflective agents). Human interrater reliability (kappa = 0.906–0.998) far exceeded LLM annotation (kappa = 0.541–0.609), highlighting ICAP's role — and current limits — in automated engagement measurement for [[learning-analytics]] pipelines.([[icap-cognitive-engagement-llm-agents]]) ### Guiding collaborative-dialogue facilitation Because interactive engagement is the highest ICAP mode, the framework helps locate the value of AI facilitation in [[collaborative-learning|online collaborative discussion]]. [[llm-facilitation-timing-online-discussions|Research on LLM facilitation timing]] shows that *when* an AI intervenes in a discussion shapes whether it supports or interrupts interactive knowledge co-construction — an ICAP-informed caution that autonomous moderation agents need calibration toward human-like restraint rather than over-eager facilitation. ### ICAP and learning analytics design ICAP underlies critiques of shallow "engagement" metrics: interacting with a dashboard by clicking filters is *active*, not *interactive*, engagement. Effective learning-analytics designs elicit self-assessment and two-way dialogue rather than merely displaying data — an implication drawn directly from Chi's framework.([[interactive-learning-dashboards-engagement]]) ### The Active→Constructive transition as the pivotal step Although ICAP describes a hierarchy, the most consequential shift for learning is the jump from *Active* to *Constructive* modes (Chi & Boucher, 2023). Active engagement (applying knowledge to similar-but-non-identical scenarios) prepares learners, but it is Constructive engagement — generating explanations, summaries, or new artifacts — that equips them to create new knowledge. This is the crux for AI in education: a tool that keeps learners in the Active mode (e.g., clicking through adaptive practice) may look productive but never pushes them into the constructive generation that yields durable understanding. Collaborative and literacy-focused interventions that deliberately scaffold the Active→Constructive leap tend to show the strongest gains.([[hingle-collaborative-ai-literacy-2025]]) ### ICAP as an adaptive-scaffolding signal in ITS ICAP's modes can be operationalized as *target states* that an adaptive tutor selects among to scaffold cognitive engagement based on an evolving student model. In a logic ITS, [[adaptive-scaffolding-cognitive-engagement-its|Dey Tithi et al.]] dynamically chose between an *Active* "Guided" worked-example mode and a *Constructive* "Buggy" example mode. Comparing Bayesian Knowledge Tracing (BKT) against Deep Reinforcement Learning (DRL) and a non-adaptive baseline over 113 students, both adaptive policies improved posttest performance — but in a differentiated way: BKT gave the largest gains to low-prior-knowledge students (helping them catch up), while DRL produced the highest posttest scores among high-prior-knowledge students. This is a concrete demonstration that effectively *personalizing* the ICAP mode of an intelligent tutor depends on modeling the learner's current knowledge — and that no single mode or adaptive method suits every learner. It connects the ICAP hierarchy directly to [[adaptive-learning]] and [[knowledge-tracing]] design. ### ICAP as a model of cognitive state for generating human-like agents Beyond selecting task modes, ICAP has been embedded directly into the *cognitive model* of a generative educational agent. [[cogevolution-student-cognitive-evolution-agent-2026|CogEvolution]] builds an ICAP-based "cognitive depth perceptron" that maps inputs to a probability distribution across the four ICAP levels, fusing this with evolutionary-inspired state updates and item-response-theory memory retrieval to simulate a student's cognitive evolution (including transitions such as confusion → insight). Ablations show that removing the ICAP perception module collapses the agent's ability to distinguish shallow from deep learning — evidence that the ICAP taxonomy can serve as a fine-grained, internal measure of cognitive engagement for [[simulating-students|student simulation]], not merely an external evaluation lens. ### ICAP anchors assessment of reflective GenAI interaction ICAP's emphasis on generative, process-level engagement has been adopted by assessment frameworks that evaluate *how* students learn with generative AI. [[assessing-student-drive-framework-2025|The DRIVE framework]] explicitly aligns its core construct — deep reflective interaction with GenAI output — with the kind of generative engagement ICAP identifies as leading to deeper learning, and uses it to distinguish surface consumption from effortful, reflective reworking of AI-generated content. This positions ICAP as a theoretical anchor for designing and measuring meaningful [[generative-ai|GenAI]] learning interactions rather than merely tracking usage. ## Implications for design and research 1. **Design for the higher modes.** AI tools should prompt learners to generate, explain, and dialogue — constructive and interactive activity — rather than deliver passive content or act as answer machines.([[multimodal-learning-genai]]) 2. **Engage learners across modes.** Effective [[ai-literacy|AI literacy]] instruction engages learners at multiple ICAP levels — passive exposure, active manipulation, constructive generation, and interactive dialogue — selecting the mode that fits the learning goal.([[hingle-collaborative-ai-literacy-2025]]) 3. **Measure engagement honestly.** ICAP gives researchers and designers a common vocabulary for distinguishing genuine cognitive engagement from mere activity — a corrective to shallow [[student-engagement]].([[icap-cognitive-engagement-llm-agents]]) 4. **Watch the human–LLM annotation gap.** If automated systems are used to code engagement, their systematic shortfall relative to trained humans must be accounted for.([[icap-cognitive-engagement-llm-agents]]) ## Connected Concepts - [[active-learning]] - [[collaborative-learning]] - [[student-engagement]] - [[learning-analytics]] - [[constructivist]] - [[learning-design]] - [[metacognition]] - [[ai-literacy]] - [[human-in-the-loop-ai]] - [[limitations-in-aied-research]] ## Connected Articles - [[icap-cognitive-engagement-llm-agents]] — Extended ICAP framework for measuring engagement with human vs. LLM annotation - [[hingle-collaborative-ai-literacy-2025]] — Collaborative AI literacy across the four ICAP modes - [[interactive-learning-dashboards-engagement]] — ICAP as a critique of shallow learning-analytics engagement - [[multimodal-learning-genai]] — ICAP and cognitive engagement in multimodal learning design - [[llm-facilitation-timing-online-discussions]] — LLM facilitation timing in online collaborative discussions - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[cogevolution-student-cognitive-evolution-agent-2026]] — ICAP cognitive-depth model in a generative student-simulation agent - [[assessing-student-drive-framework-2025]] — ICAP-anchored assessment of reflective GenAI interaction - [[code-to-learn-genai-artifact-construction-2026]] — CtL-GenAI: constructionism framework for artifact construction --- ## [Misconceptions about AI](https://edtechdev.github.io/aied/concepts/misconceptions/) > **Misconceptions about AI** — the inaccurate beliefs people hold about what AI systems are, what they do, and what using them means for learning and work. Misconceptions are not a single falsehood but a family of calibration errors that cluster around two core mistakes: misjudging what the model is (authority vs. tool, neutral vs. biased, understanding vs. generating) and misjudging what learning requires (output vs. process). They are held not only by students but also by teachers, administrators, policymakers, and the broader public — and correcting them is a core aim of [[ai-literacy]] and [[trust-calibration]] education. ## Questions to Consider - Many people believe an AI chatbot's answer is a verified fact because it sounds confident and fluent. The page calls this the 'authority fallacy.' When you read an AI-generated explanation, how do you decide whether to accept it — and how often do you actually verify it? - The 'learning-equals-output' misconception is the belief that producing work with AI is the same as having learned it. Have you ever felt you 'learned' something by letting a tool do the drafting? What was actually missing afterward? - People often assume AI is objective and unbiased. The page calls this the 'neutrality illusion' — models encode biases from training data, which in writing contexts can homogenize ideas across an entire class. Where might bias hide in a tool that feels neutral? - Misconceptions about academic integrity cluster at two extremes: some students treat AI output as 'not copying a person' and so permissible, while others think any use is cheating. Where do you think the line should fall, and who should decide it? - A common belief is that one query is enough and that AI output is deterministic — the same question always yields the same answer. The page describes this as the determinism error. How might that misconception lead someone to over-trust a single output? - Misconceptions are described as stable, plausible, and resistant to correction — much like misconceptions in any domain. If simply telling people the truth rarely changes their minds, how should AI literacy actually be taught? ## Introduction Misconceptions about AI matter because they are the cognitive precursor to the harmful behaviors the knowledge base documents under [[cognitive-offloading|Over-Reliance]] and [[academic-integrity]] concerns. People rarely set out to [[ai-misuse-learning-harm|misuse]] AI; they do so because inaccurate mental models lead them to misplace trust, skip verification, and treat output as understanding. In education these errors shape everything from how students study to how teachers and institutions design curricula, [[assessment]], and policy. ### What AI misconceptions are A misconception here is not mere ignorance of how a model works — it is an actively held, often self-reinforcing belief that produces systematic errors in how students interact with AI. They are directly analogous to the domain misconceptions studied in [[learning-theories|learning science]]: stable, plausible, and resistant to correction until confronted. Correcting them is a core aim of [[ai-literacy]] and [[trust-calibration]] education. ### Common misconceptions in academic contexts - **The authority fallacy** — treating [[llm]] output as verified fact rather than a probabilistic completion. Drives uncritical acceptance and the answer-seeking-over-understanding pattern documented in [[intelligent-tutoring|AI-tutoring]] research, where learners accept a model's answer without checking it against [[hallucination-risk]]. - **Learning-equals-output** — believing that producing work *with* AI is the same as having learned it. This is the exact error behind [[cognitive-offloading|Over-Reliance]]: the drafting, recall, and revision processes that build durable knowledge get outsourced. - **The neutrality illusion** — assuming AI is objective and unbiased. Students often miss that models encode training-data biases and that in [[writing-education]] contexts this produces idea homogenization across a cohort. - **The integrity gray zone** — misjudging whether [[academic-integrity|AI use is acceptable]]. Some students see AI output as "not copying a person" and therefore permissible; others over-correct and think *any* use is cheating. [[governance|Institutional]] inconsistency feeds both errors. - **Anthropomorphism** — believing the model has intent, memory, and understanding of *their* context. This over-trust is especially risky academically, because students may rely on plausible-sounding explanations the model cannot actually ground. - **The determinism error** — expecting one query to be enough and not realizing output is non-deterministic and prompt-sensitive. Underestimating this produces the "[[prompt-engineering|prompting]] gap," where students mistake shallow results for the tool's ceiling. - **The [[ai-detection|detection]] miscalibration** — underestimating both institutional detection and, more importantly, the self-harm of submitting work they cannot later explain or defend. - **The efficiency illusion** — treating time saved as pure gain, missing that unexercised foundational skills decay and that novices cannot yet tell good output from bad. - **The friendliness-safety illusion (children)** — believing a friendly-seeming AI is inherently safe and less likely to spy. [[children-ai-safety-misconceptions-2026|Leisten et al. (2026)]] surveyed 71 children aged 10–16 (*M*age = 12.90) and ran six focus groups (n = 36) around the [[open-source]] social robot Blossom, finding foundational AI knowledge rose reliably with age (*M* = 4.41 of 6; β = 0.17) while safety attitudes stayed ambivalent (importance *M* = 2.51, current safety *M* = 2.77 on a 1–4 scale), captured in a 12–13-year-old's belief that "maybe if they are good friends he doesn't spy so much" — alongside the belief that AI's knowledge can be deleted at the press of a button, that hacking leads to kidnapping, and that "the wifi radiates into the brain." ### Institutional and public AI myths Misconceptions are not confined to students — they saturate the institutional and public discourse about AI that students inherit. [[rudolph-ai-myths-critical-higher-ed|Rudolph et al. (2025)]] dismantle eight entrenched "myths" that shape higher-education policy and teaching: that AI is genuinely "artificial" (rather than built from exploited human labor), that it is truly "intelligent" and [[agentic-ai|agentic]], that it will unproblematically "make the world a better place," that it is "objective and unbiased," that the US holds a sole superpower monopoly, that it will not disrupt the job market, that it "revolutionises [[higher-ed|higher education]]," and that teachers can reliably detect AI-generated work. These institutional myths are the upstream source of many student misconceptions documented above — most directly the [[trust|neutrality illusion]] ("AI is objective") and the [[trust-calibration|authority fallacy]] ("AI is intelligent"), and the detection miscalibration that leads students to assume undetectable, unverifiable use is safe. Where students absorb and act on institutionally-repeated myths, correcting them requires confronting not only the learner's belief but the discourse that feeds it. ### Why misconceptions matter for learning Misconceptions translate directly into the behaviors that cause learning harm. The belief that "AI is always right" suppresses verification; the belief that "using AI is learning" suppresses effortful processing; the belief that "it's not cheating" bypasses the metacognitive review that consolidates understanding. In this sense misconceptions are upstream of the [[ai-misuse-learning-harm]] documented across the knowledge base's evidence base. ### Misconceptions beyond students: teachers, institutions, and the public AI misconceptions are not confined to learners — they are pervasive among the adults who shape education: - **Teachers and faculty** may overestimate AI's ability to reliably grade or detect misuse, or underestimate its bias, leading to either uncritical adoption or reflexive banning. That assumption of reliability is partly testable and partly false: [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] red-teamed an everyday [[automated-assessment|AI grading]] workflow and found that instructions hidden inside a submitted file raised a failing essay's grade with no visible warning, in 9 of 9 iterations for one strategy and 17 of 18 for another. When teachers hold the [[trust|authority fallacy]] about AI outputs, they model the same uncritical posture they should be correcting in students. Preparing [[teacher-role|educators]] with accurate mental models of AI is a prerequisite for [[teacher-ai-competency|responsible AI integration]] and [[pedagogical-safety|safe pedagogy]]. - **Administrators and policymakers** inherit and propagate institutional myths — that AI is "objective," that it will "revolutionise" education, or that detection tools are trustworthy — which then shape [[educational-policy-ai|policy]], procurement, and assessment rules. The [[trust-calibration|trust]] students develop is partly a product of the institutional framing they inherit. - **The general public** absorbs media and vendor narratives about AI's capabilities and risks. Because students learn within this discourse, public myths become the substrate from which student misconceptions grow. Correcting AI misconceptions is therefore an [[ai-literacy]] task aimed at the whole educational ecosystem, not only at learners. This breadth is why the knowledge base treats misconceptions as a cross-cutting foundational theme rather than a purely student-facing one: the same calibration errors recur across learners, teachers, institutions, and the public, and correcting them requires confronting both individual beliefs and the discourse that feeds them. ### Correcting misconceptions Correction is not a one-time disclosure but an ongoing [[ai-literacy]] process that develops [[metacognition]] and [[self-regulated-learning]]: helping students (and the adults around them) monitor their reliance, calibrate when to trust and when to question a model, and see the cost of bypassing their own [[cognitive-offloading|cognitive work]]. Because misconceptions are resistant, they are best addressed through direct confrontation with evidence — including the finding that students often *do not perceive* the learning harm of AI misuse. **Refutation text is a core correction technique.** Because misconceptions are actively held and resistant, the most direct evidence-based strategy is the [[refutation-text|refutation text]] — an instructional text that states the misconception, explicitly refutes it, and presents the correct conception. This is the same family of technique used to correct the domain misconceptions studied in learning science, applied here to students' beliefs about AI itself. The knowledge base's [[refutation-text]] concept page synthesizes how this plays out in [[ai-education|AI in education]] in three complementary ways: - **AI as the corrector.** [[conversational-ai|Conversational AI]] tutors can deliver *personalized* refutation, adapting the refutation to a learner's specific misconception on the fly. [[ai-tutors-vs-tenacious-myths-personalized-dialogue-2026|Corbett & Tangen (2026)]] found personalized AI dialogue produced larger and faster belief reductions than static textbook-style refutation, with higher [[student-engagement|engagement]] and confidence — though the advantage faded by two months without reinforcement. - **AI as the generator of refutation content.** [[akdogan-heat-temperature-conceptual-change-thesis-2025|Akdoğan (2025)]] found AI-generated conceptual-change/refutation text matched expert-written quality (and both outperformed a prompted interactive dialogue in that science context), showing AI can produce effective correction materials at scale. - **AI-generated misconceptions as a learning resource.** Rather than treating AI-generated misconceptions as merely harmful, [[llms-misconception-collaborative-learning-healthcare-2026|Cheah et al. (2026)]] propose generating misconceptions and addressing them through structured peer discussion — a [[collaborative-learning|collaborative]] form of refutation that promotes conceptual change and [[critical-thinking|critical thinking]]. For misconceptions about AI, this means correction should combine **direct confrontation** (refutation-style materials that name and rebut specific myths) with **scaffolded practice** — using [[ai-literacy]] instruction and [[metacognition]] to help people see both the false belief and the correct model. The evidence cautions that the *format* matters: personalized, interactive correction is more engaging and initially more effective, but needs reinforcement to persist; and the outcome measured (knowledge vs. attitudes vs. skills) shapes how large a correction effect appears. Because misconceptions span learners and the adults who shape learning, effective correction must reach [[teacher-role|teachers]], [[administrator|administrators]], and [[educational-policy-ai|policymakers]] as much as students. Bernstein and Sibia (2026) show that [[generative-ai|GenAI]]-generated analogies introduce structural misconceptions that only source-domain knowledge can catch ([[student-reception-genai-analogies-computing-2026]]): a circular-route analogy for a linked list implies a loop back to the start, and a badminton-rally analogy for recursion carries no guaranteed shrinking input. Students who knew the source domain identified these flaws and proposed repairs, while participants noted that a flawed analogy may still be memorable — indicating that a familiar analogy source can help learners detect, rather than absorb, an AI-produced misconception, and that framing flaws as deliberate artifacts for critique turns the risk into an assessment opportunity. **Model-generated misconceptions are a measurable capability, not an accident.** [[milicevic-socratic-trap-strategic-misconceptions-2026|Miličević et al. (2026)]] built SocraticTrap-CS, which prompted seven open-weight models to write a "[[socratic-method|Socratic]] trap" for 35 core CS concepts — an explanation that is fluent and authoritative while resting on a subtle, domain-specific error. Of 241 prompted segments, 221 (91.7%) were confirmed as strategic misconceptions by expert majority vote (Fleiss' κ = 0.9487), with no significant differences between CS domains; 66.5% of the confirmed errors were conceptual rather than factual and none were purely logical. Fluency is the mechanism rather than a defense: persuasiveness averaged 3.71 on a five-point scale and was strongly model-dependent, and frequency and severity dissociated, with the two models that produced traps most often also rated most convincing. Because a student's poorly framed question can itself act as an adversarial prompt, the authors treat the rate as a capability under adversarial prompting rather than a base rate for ordinary study sessions — and argue the [[ai-literacy]] task shifts from fact-checking individual statements to conceptual verification and mental-model validation. This is the darker face of the generative use above: the same capability that can seed productive peer discussion can also entrench an error a learner was already forming. ### Refutation-style corrections for common AI misconceptions Because misconceptions are actively held and resistant, the most direct way to address them — including on this page — is the [[refutation-text]] structure: **name the misconception, explicitly refute it, and state the correct conception.** The entries below apply that structure to the most consequential misconceptions about AI and about learning, teaching, and education: **"AI is always right."** *That's a misconception.* AI output is a probabilistic completion, not a verified fact. *The correction:* LLMs generate plausible-sounding text based on statistical patterns; they can [[hallucination-risk|hallucinate]], be biased, and be confidently wrong. Treat output as a draft to be checked against sources, not an authority to be accepted. This is the core of [[trust-calibration]] and why "always verify" beats "always trust." **"Using AI is learning."** *That's a misconception.* Producing work *with* AI is not the same as acquiring the knowledge or skill the work is supposed to demonstrate. *The correction:* durable learning happens through the effortful processes of drafting, recalling, revising, and metacognitively reviewing — exactly the processes that [[cognitive-offloading|offloading]] to AI short-circuits. Use AI as a tool alongside that effort, not a replacement for it. **"AI is neutral and objective."** *That's a misconception.* Models inherit the biases, gaps, and perspectives of their training data. *The correction:* AI can reproduce and amplify [[bias-mitigation|bias]]; treat its outputs with the same source-critical scrutiny you would apply to any other text. Awareness of this is part of [[ai-literacy]] and helps counteract the [[equity-in-ai-education|equity]] harms of uncritical adoption. **"AI will replace teachers."** *That's a misconception.* AI augments but does not displace the [[pedagogy|pedagogical]] work of [[teacher-role|teachers]] — judgment, contextualization, and the relational and [[ethics|ethical]] dimensions of teaching. *The correction:* AI increases the need for pedagogical mediation and critical judgment; teachers who understand AI become more effective, not obsolete. This reframing matters because it shapes whether institutions invest in [[teacher-ai-competency|teacher AI competency]] or reflexively resist or over-adopt. **"AI understands like a person."** *That's a misconception.* Models have no intent, memory of you, or genuine understanding of your context. *The correction:* anthropomorphizing AI leads to over-trust and reliance on explanations the model cannot actually ground. Keep the boundary clear: AI is a powerful tool, not a mind. **"AI will transform education automatically."** *That's a misconception.* Technology alone does not change learning; it is the pedagogy around it that does. *The correction:* AI's benefits depend on intentional [[learning-design|instructional design]], teacher preparation, and institutional support — not on simply deploying the tool. This is why [[ai-ed-evaluation|evidence]] and [[research-methods-aied|rigorous evaluation]] matter, and why the knowledge base frames responsible AI use as a [[governance]] and [[educational-policy-ai|policy]] question rather than a purely technical one. **"One prompt should give me the answer."** *That's a misconception.* Output is non-deterministic and prompt-sensitive. *The correction:* expect to iterate, refine, and cross-check; the "prompting gap" — mistaking shallow first results for the tool's ceiling — is a skill problem, not a tool limit. Developing this is part of [[prompt-engineering]]. **"It's not cheating if a person didn't write it."** *That's a misconception.* Academic integrity is about the honest, attributable production of work, not just about not copying a person. *The correction:* undisclosed AI-generated submission can violate [[academic-integrity]] even when no human was copied; the question is whether the work is genuinely the learner's. When in doubt, disclose and check your institution's policy. These refutations are deliberately written in the [[refutation-text]] form so they can themselves be used (or adapted into interactive [[conversational-ai|AI dialogue]]) to confront and correct misconceptions about AI — and about learning, teaching, and education more broadly. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[ai-literacy]] - [[trust-calibration]] - [[cognitive-offloading]] - [[metacognition]] - [[self-regulated-learning]] - [[academic-integrity]] - [[hallucination-risk]] - [[generative-ai]] - [[student-experience]] - [[framing-ai-use-for-students]] - [[refutation-text]] - [[teacher-role]] - [[educational-policy-ai]] - [[trust]] ## Connected Articles - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[rudolph-ai-myths-critical-higher-ed]] — Don't believe the hype: eight AI myths and the need for a critical approach in higher education - [[drawedumath-vlm-struggling-students-2026]] — VLMs misdiagnose student math errors (DrawEduMath, Lucy et al. 2026) - [[student-rationalization-ai-writing]] — Student Rationalization of AI Writing - [[genai-skill-bypass-literacy]] — GenAI Skill Bypass and Literacy - [[trust-reliance-ai-education-2026]] — Trust and Reliance in AI Education - [[contextual-sycophancy-ai-literacy]] — Contextual Sycophancy and AI Literacy - [[sycophantic-ai-social-interaction-2026]] — Sycophantic AI in Social Interaction - [[llm-fallacy-misattribution]] — LLM Fallacy Misattribution (Kim et al.) - [[generative-ai-guardrails-harm-learning]] — GenAI Without Guardrails Can Harm Learning - [[student-reception-genai-analogies-computing-2026]] — Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education - [[milicevic-socratic-trap-strategic-misconceptions-2026]] — SocraticTrap-CS: fluent, authoritative explanations that are wrong conceptually rather than factually (Miličević et al. 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Hidden instructions in a submitted file can raise an AI-graded mark with no visible warning (Humble 2026) - [[children-ai-safety-misconceptions-2026]] — Children's AI-safety misconceptions: friendship with a robot misread as a privacy guarantee (Leisten et al. 2026) - [[mental-health-literacy-students-llms-2026]] — Mental Health Literacy Across Psychology Students and Large Language Models --- ## [Refutation Text](https://edtechdev.github.io/aied/concepts/refutation-text/) > **Refutation text** — a misconception-correction technique in which a text explicitly states a common misconception, directly refutes it, and then presents the scientifically correct conception. Originating in the [[misconceptions|conceptual-change]] literature of science education, refutation texts are a proven, low-tech intervention for dislodging stable, intuition-aligned misconceptions that resist ordinary instruction. In AI in education, refutation texts are increasingly used in two ways: as a **comparison condition** for AI-based interventions (personalized dialogue, LLM-generated content), and as **AI-generated content** — conceptual-change texts and misconception texts produced by [[generative-ai|generative AI]] to correct beliefs or to seed collaborative discussion. ## Questions to Consider - Have you ever 'corrected' a student's wrong idea by simply presenting the right answer, only to have the misconception resurface later? The page argues misconceptions aren't gaps but actively held beliefs that resist ordinary instruction. What does that reframe about why your correction failed? - A refutation text states the misconception explicitly, refutes it, and offers the correct conception — unlike a standard expository text that just presents the truth. Why would naming the wrong idea out loud help change it, when teaching only the right idea apparently doesn't? - The research is mixed on whether personalized AI dialogue beats static refutation text: in one study interactive dialogue produced larger, faster belief change; in another, well-crafted texts outperformed a prompted AI chat. What might explain these contradictory results, and what does it tell you about 'interactivity is always better'? - AI can now generate effective refutation texts that match expert-written quality, and even generate misconceptions to seed structured peer discussion. Does the idea of deliberately teaching from AI-generated wrong ideas feel risky or productive to you — and under what conditions would you try it? - Refutation effects appear concentrated among high-achieving students and moderated by epistemology and metacognition. If the technique helps the strong most, what obligations does that create for an instructor using it with a mixed classroom? - Before you read further, name one misconception you currently hold about a subject you teach, and imagine writing the explicit 'wrong' claim and its refutation yourself. What did that exercise reveal about how hard good refutation is to write? ## Introduction ### The concept Refutation texts rest on the idea that misconceptions are not mere gaps in knowledge but actively held, plausible, self-reinforcing beliefs that resist correction — a claim central to conceptual-change research. A refutation text works by making the misconception explicit, naming it as wrong and explaining why, and then offering the correct conception in a way the learner can integrate. This differs from a standard expository text, which simply presents correct information and assumes the misconception will be displaced. In AI in education, the core finding is that the *format* and *interactivity* of the correction matter. Converging evidence ([[ai-tutors-vs-tenacious-myths-personalized-dialogue-2026|Corbett & Tangen 2026]]) shows that static textbook-style refutation reliably corrects beliefs but that **personalized, interactive AI dialogue** can produce larger and faster belief reduction by targeting the learner's specific misconception and engaging them motivationally. However, this advantage may be context- and design-dependent: in [[akdogan-heat-temperature-conceptual-change-thesis-2025|science education (Akdoğan 2025)]], well-structured conceptual-change texts (expert *or* AI-generated) outperformed a prompted interactive ChatGPT dialogue — suggesting that the dialogue's design (personalized vs. generic) and the domain shape which format wins. ### Why refutation text matters for AI in education - **AI as the corrector.** Conversational AI tutors can deliver *personalized* refutation — adapting the refutation to the learner's specific misconception on the fly, which pre-written texts cannot do. This produces stronger immediate belief change and higher engagement/confidence than static refutation ([[ai-tutors-vs-tenacious-myths-personalized-dialogue-2026|Corbett & Tangen 2026]]), though effects may need spaced reinforcement to persist. - **AI as the generator of refutation content.** [[generative-ai|Generative AI]] can produce effective conceptual-change texts that match expert-written quality ([[akdogan-heat-temperature-conceptual-change-thesis-2025|Akdoğan 2025]]), and can generate large numbers of context-specific misconception texts cheaply — scaling misconception-based learning that would otherwise depend on educator experience ([[llms-misconception-collaborative-learning-healthcare-2026|Cheah et al. 2026]]). - **AI-generated misconceptions as a learning resource.** Rather than viewing AI-generated misconceptions as harmful, structured peer discussion of them — a form of collaborative refutation — can promote conceptual change and critical thinking ([[llms-misconception-collaborative-learning-healthcare-2026|Cheah et al. 2026]]). - **Complementing misconception education.** Refutation texts are a recommended strategy for correcting the conceptual misconceptions that underpin students' mistaken beliefs about AI itself (see [[misconceptions]] and [[critical-genai-use-predictors]]). ### Refutation text vs. related techniques Refutation texts are one member of the conceptual-change toolkit, alongside analogies, discrepant events, and interactive dialogue. Their advantage is that they are **scalable, low-cost, and demonstrably effective**; their limitation is that static texts cannot adapt to the learner. AI dialogue addresses the adaptation gap but introduces design-dependence (personalization, prompt quality) and, in some studies, no advantage over well-crafted text. The relationship between refutation text and AI dialogue is thus complementary: text offers reliable baseline correction at scale; personalized AI dialogue offers stronger, faster, more motivating correction when well designed. ### Key research themes - Whether personalized AI dialogue outperforms static refutation text, and under what conditions. - Whether AI-generated refutation/conceptual-change text matches expert-written quality. - Using AI to generate misconceptions for collaborative, misconception-based learning. - The role of learner characteristics (achievement, epistemology, metacognition) in moderating refutation effectiveness. - Spaced reinforcement to sustain the initial advantages of interactive refutation. ### Practical implications For educators, refutation texts remain a reliable, low-barrier way to correct stubborn misconceptions. For those integrating AI, the evidence suggests: (1) use AI to *generate* effective refutation/conceptual-change content at scale; (2) where feasible, deliver refutation through personalized AI dialogue for stronger immediate engagement and belief change; (3) expect AI-generated misconceptions to be pedagogically useful when structured discussion is used to confront them; and (4) design for the learner — refutation effects can be concentrated in high-achieving students and moderated by epistemology and metacognition, so scaffolding and follow-up matter. ## Connected Concepts - [[misconceptions]] - [[scaffolding]] - [[metacognition]] - [[generative-ai]] - [[stem-education]] - [[physics-education]] - [[medical-education]] - [[collaborative-learning]] - [[intelligent-tutoring]] ## Connected Articles - [[ai-tutors-vs-tenacious-myths-personalized-dialogue-2026]] — Personalized AI dialogue vs. textbook refutation for belief correction - [[akdogan-heat-temperature-conceptual-change-thesis-2025]] — Expert/AI conceptual change text vs. interactive AI dialogue - [[llms-misconception-collaborative-learning-healthcare-2026]] — LLM-generated misconceptions for collaborative learning - [[chatgpt-inoculation-training-verification-2026]] — Inoculation training as an adjacent refutation-style intervention - [[critical-genai-use-predictors]] — Recommends refutation texts to target conceptual misconceptions - [[ai-learning-companions-framework]] — AI companions and misconception correction --- ## [Activity Theory](https://edtechdev.github.io/aied/concepts/activity-theory-aied/) > **Activity theory** (Cultural-Historical Activity Theory, CHAT) — a Vygotskian framework that analyzes learning and work as *tool-mediated, object-oriented, collective activity systems* composed of subject, object, tools/mediating artifacts, community, rules, and division of labor. In [[ai-education|AI in education]], activity theory is used both as an **analytic lens** (to understand how AI reshapes the activity systems of teaching, learning, and research) and as a **design tool** (to diagnose systemic contradictions and redesign interventions). It frames AI systems — [[generative-ai|generative AI]] included — as *mediating artifacts* that reconfigure the division of labor, rules, and community of educational activity, for better and worse. ## Questions to Consider - Activity theory sees learning not as individual cognition but as a collective system of subject, tools, rules, community, and division of labor. If you bring a new AI tool into a classroom, which of these do you expect to change — and which to stay stubbornly the same? - A key idea is that contradictions within an activity system are the driving force of development, not bugs to be eliminated. Can you think of a tension in your own work or study that actually pushed you to change how you did things? - Activity theory reframes a teacher's adoption of AI as a property of the whole system — norms, rules, workload distribution — rather than just individual attitudes. What systemic reasons might explain why a capable teacher resists a genuinely good AI tool? - When AI takes over tasks, it reconfigures who does what, shifting the division of labor between students, teachers, and tools. What cognitive work in your own context has quietly moved from humans to machines — and who noticed? - Because each discipline functions as its own activity system, the same AI tool can produce different outcomes across subjects. Why might a tool that transforms writing classes barely change a math class, or vice versa? - Activity theory is used both to analyze how AI reshapes teaching and to design interventions that fix the tensions it exposes. How might seeing your own teaching or study practice as an activity system change how you diagnose a problem? ## Introduction ### The concept Activity theory descends from Vygotsky's and Leontiev's cultural-historical psychology (with later development by Engeström) and is a central strand of the broader [[sociocultural-learning|sociocultural]] tradition. Its core claim is that human activity is not reducible to individual cognition or isolated tool use; it is a **collective, object-oriented, tool-mediated system**. Engeström's canonical model of an activity system comprises six interacting elements: - **Subject** — the individual or group whose agency is the point of view of the analysis (e.g., a student, a teacher, a department). - **Object** — the motive or goal toward which activity is directed; the "raw material" that the activity transforms. - **Tools / mediating artifacts** — the instruments, technologies, signs, and language through which the subject acts on the object. - **Community** — the group that shares the object and constitutes the social context of the activity. - **Rules** — the explicit and implicit norms, conventions, and regulations that govern activity. - **Division of labor** — how tasks, power, and responsibility are distributed across the community. A defining feature is **contradiction**: activity systems contain historically accumulating structural tensions (within an element, between elements, or between an old and a newly introduced element/tool). Contradictions are not bugs to be eliminated but the *driving force of development* — when aggravated, they prompt participants to question and deviate from established norms, opening the way for expansive transformation (Engeström, 2001). ### Why activity theory matters for AI in education AI introduces new **tools/mediating artifacts** into existing educational activity systems, which produces both opportunities and contradictions. Activity theory gives AIED research a vocabulary and method for analyzing these changes: - **AI as a mediating artifact that reconfigures the division of labor.** AI systems do not simply transmit information; they take over tasks, change who does what, and redistribute cognitive work across humans and machines. Studies use activity theory to examine how AI shifts the division of labor between students, teachers, and tools — e.g., in collaborative writing or tutoring — and what this means for [[agency|human agency]]. - **Analyzing teacher adoption and [[educational-development|professional development]].** Activity theory reframes teacher uptake of AI as a property of the whole activity system (community norms, rules, division of labor, [[governance|institutional]] expectations) rather than of individual attitudes alone. [[quantitative-research|Quantitative]] work [[activity-theory-teachers-adoption-ai-sem-2026|models the six AT components as measurable constructs]] predicting teachers' intention to adopt AI; [[qualitative-research|qualitative]] work [[lee-anson-k12-teachers-ai-activity-theory|documents teachers' sixfold sentiments]] (unsuitable, impersonal, imperfect, uncertain, assisting, inevitable); and intervention work [[activity-theory-teacher-pd-ai-agent-design-2026|uses CHAT to diagnose and redesign]] teacher professional development, treating disengagement as a rational response to need-thwarting systems. - **Norm change and systemic disruption.** Because AI introduces a new tool into the activity system, it generates contradictions with existing rules and norms. Studies of students' GenAI use [[ai-disruption-engineering-education-chat-2026|show how new implicit rules emerge]] as students adapt — transforming norms around self-direction, learning objectives, the teacher's role, and [[ethics]]. - **Anchoring [[learning-analytics|learning analytics]] and measurement.** Activity theory can ground the *design* of analytics by mapping measurement facets onto activity-system elements. A CHAT-anchored analytics pipeline [[chat-anchored-learning-analytics-ai-literacy-2026|maps temporal participation, discourse quality, and concept sophistication to CHAT elements]] to detect early at-risk participation in small discussion-based classes. - **Disciplinarity and cross-context variation.** Because activity systems are historically and culturally situated, activity theory explains why the same AI tool produces different outcomes across disciplines and contexts — each discipline functioning as an activity system with its own rules, community, and division of labor [[jiang-genai-activity-theory-disciplines-2026|(e.g., differences in undergraduates' GenAI use and disclosure across academic domains)]]. ### Activity theory and related frameworks Activity theory sits within the [[sociocultural-learning|sociocultural]] family and is often paired with, or contrasted against, other frameworks: **[[distributed-cognition|distributed cognition]]** and **[[situated-learning|situated learning]]** share its emphasis on context and mediation; **[[community-of-inquiry|community of inquiry]]** and **communities of practice** share its collective orientation; and **technology-acceptance** models (e.g., [[technology-acceptance-model|TAM]]) offer a contrasting, more individual-belief account of adoption that activity theory critiques for flattening the collective and structural dimensions. In AIED research, activity theory is frequently combined with other theories — e.g., paired with [[self-determination-theory|Self-Determination Theory]] in teacher-PD intervention design, or with ecological systems theory in situated-AI-ethics frameworks [[raffaghelli-situated-ai-ethics-2026|(Raffaghelli et al., 2026)]]. ### Key research themes - **How AI redistributes the division of labor** in teaching, learning, and research activity systems. - **Systemic contradictions** as drivers of norm change and expansive transformation in AI-augmented education. - **Teacher adoption and professional development** as activity-system (not merely individual) phenomena. - **Disciplinary and cross-cultural variation** in AI use, explained by differences in activity systems. - **Theory-anchored learning analytics** that ground measurement in activity-system elements. - **The object of AI-mediated activity** — whether the object of learning shifts from mastery to output, and what that means for [[learning-gains|learning outcomes]]. ### Practical implications For designers and educators, activity theory counsels looking beyond the AI tool itself to the whole activity system it enters: the rules that govern acceptable use, the community and its norms, the division of labor between human and machine, and the object/motive of the activity. Sustainable AI integration requires *redesigning the system* — addressing contradictions, updating norms, and rebalancing roles — not merely supplying better tools or individual training. For researchers, activity theory offers a rigorous unit of analysis (the activity system, not the individual) and a method ([[formative-assessment|formative]] intervention, contradiction analysis) for studying and shaping AI's role in education. ## Connected Concepts - [[sociocultural-learning]] — the theoretical family activity theory belongs to (Vygotskian, cultural-historical) - [[learning-theories]] — the hub that situates activity theory among learning theories - [[distributed-cognition]] — shared emphasis on tool-mediated, contextually distributed cognition - [[situated-learning]] — shared emphasis on context and participation in activity - [[agency]] — how AI redistributes human agency across the activity system - [[teacher-role]] — activity theory frequently analyzes teachers' adoption and practice - [[learning-analytics]] — CHAT anchors the design of analytics - [[generative-ai]] — the AI artifact that enters and disrupts educational activity systems - [[technology-acceptance-model]] — the contrasting individual-belief account of adoption ## Connected Articles - [[jiang-genai-activity-theory-disciplines-2026]] — Disciplinary differences in GenAI use and disclosure through an activity theory lens - [[genai-runaway-object-math-higher-ed]] — CHAT analysis of GenAI reshaping teaching and research activity systems - [[raffaghelli-situated-ai-ethics-2026]] — Situated AI ethics fusing ecological systems theory with CHAT - [[zhang-ai-students-disabilities-meta-analysis-2024]] — CHAT framing of AI interventions for students with disabilities - [[activity-theory-teachers-adoption-ai-sem-2026]] — Activity theory as a lens on teachers' adoption of AI (SEM) - [[lee-anson-k12-teachers-ai-activity-theory]] — K-12 teachers' perspectives on AI use through activity theory - [[ai-disruption-engineering-education-chat-2026]] — Changing student norms in engineering education via CHAT - [[activity-theory-teacher-pd-ai-agent-design-2026]] — CHAT-SDT redesign of teacher professional development - [[chat-anchored-learning-analytics-ai-literacy-2026]] — CHAT-anchored learning analytics pipeline for AI literacy - [[chatgpt-critical-creative-thinking-review]] — CHAT as one theoretical lens on ChatGPT pedagogy --- ## [Retrieval, Spacing and Interleaving](https://edtechdev.github.io/aied/concepts/retrieval-spacing-interleaving/) > **Retrieval, spacing and interleaving** — the three concrete study techniques that make practice effortful and, as a result, make what is learned durable. Retrieval practice asks the learner to produce an answer from memory rather than recognize one; spacing distributes that practice across time instead of massing it into one session; interleaving mixes problem types instead of blocking them. What unites them is a shared signature: each lowers performance *during* practice while raising retention *after* it — the [[transfer-of-learning|performance–learning gap]] in operational form. They are the workhorses of [[cognitive-psychology]], and in [[ai-education|AI-supported learning]] they are the specific behaviors that [[generative-ai|generative AI]] most easily erases, because an [[llm|LLM]] that answers on demand removes the retrieval, flattens the schedule, and smooths away the mixing. This page is the *techniques* page. The principle behind them — Bjork's effort–learning trade-off, why harder practice produces more durable learning — is treated on [[desirable-difficulties]]; this page covers what each technique is, what the corpus's studies measured, and how AI implements or undermines them. ## Questions to Consider - Have you ever re-read a chapter until it felt familiar, then been unable to explain it a week later? What did the re-reading actually give you? - Retrieval practice means answering from memory *before* checking. When an AI assistant is one keystroke away, what would have to be true about your study routine for that retrieval to still happen? - Spaced repetition systems schedule review at intervals a learner would not choose. Is a schedule that feels wrong a feature or a bug? - Interleaving mixes problem types and feels messier than blocking. If the corpus offers almost no direct AI evidence on interleaving, how confident should you be in adopting it — and on what basis would you decide? - [[adaptive-pretesting-retention|Akgun and Toker (2026)]] found that students who chatted freely with an AI scored worst, even after the best pretesting. What did the free-chatting condition lack that the structured one had? - [[llm-interaction-depth-task-quality-recall-2026|Tsiligkiris (2026)]] found deeper LLM questioning improved task quality but not recall. What does that tell you about using fluency as evidence of learning? ## Introduction Retrieval practice, spacing and interleaving are usually taught and cited together because they were discovered separately but behave alike: each sacrifices how good practice *looks* in exchange for how much survives. In the knowledge base they appear as instances of the [[desirable-difficulties]] principle, but they are also independently actionable — a reviewer can schedule intervals, a tutor can withhold the answer, a [[curriculum-design|curriculum]] can mix item types, without invoking the underlying theory. Their operational character is why they are the natural bridge between [[cognitive-psychology]] and AI design, and why the studies below tend to report two outcome measures rather than one: performance during practice, and retention after a delay. ## The Testing Effect: Retrieval Practice Beats Re-Reading Retrieval practice — attempting to recall material rather than re-reading it — strengthens memory more than additional study does, and the corpus treats the finding as settled enough to build on rather than re-litigate. The learning-to-learn [[meta-analysis-systematic-review|scoping review]] places retrieval practice in the **Tools** layer of its three-layered framework (Dimensions of cognitive and [[metacognition|metacognitive]] skill, Processes of [[self-regulated-learning|self-regulation]], Tools such as retrieval practice), positioning it as a concrete, teachable component of learning-to-learn that is meant to counterbalance [[cognitive-offloading|GenAI overreliance]] and protect [[agency|learner agency]]. The most informative test in the corpus is a dissociation. [[llm-interaction-depth-task-quality-recall-2026|Tsiligkiris (2026)]] logged 22 postgraduate students' prompt-by-prompt [[llm]] interactions during a neuroeconomics case task and separated *depth* (the proportion of explanation-seeking "why/how/explain" prompts) from volume and pacing. Depth predicted independently marked task quality (β = 6.27, p = .006) but had a null association with immediate post-test recall (β = −0.014, p = .728), with recall gains driven by baseline knowledge instead. The reading the author gives is that elaboration drives comprehension while retrieval drives consolidation — and that in LLM-supported study without explicit retrieval demands, learners can experience high fluency with little need to retrieve anything unaided. That is a retrieval-practice finding stated as an absence: the mechanism was not engaged, so the retention outcome did not move. Retrieval practice is not a universal lever, and the corpus says so. [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit, Koedinger and Carvalho (2025)]] ran a 2×2 experiment with 95 participants crossing knowledge content (verbatim facts vs. generalizable skills, geometry-area materials) with training schedule. They found a content–treatment interaction (β = 0.41, p = .038, d = 0.38): pure practice testing produced higher learning gains for **facts**, while example-integrated practice produced higher gains for **skills**. The authors also note that retrieval-practice gains frequently fail to extend to unfamiliar problems. So retrieval practice is strongest where the target is memory for specific content, and it needs to be paired with worked examples where the target is a generalizable skill — which is also why their AI design implication is that [[intelligent-tutoring]] should adapt the example-to-practice ratio to the knowledge component being learned. The same logic runs through the cheating-sheet study: [[student-cheat-sheets-make-or-take|Chen, Sakhnini and Istead]] treat *constructing* a cheat sheet as an active, generative study strategy (selection, condensation, organization) rather than logistics, and frame the AI-era risk as offloading the artifact's construction and losing the metacognitive rehearsal that building it provided. ## Spacing and Distributed Practice The spacing effect — that the same total practice produces more durable learning when distributed across sessions than massed into one — is the most algorithmically exploited of the three techniques, and the corpus's applied systems are built around it. [[memdora-ai-spaced-repetition|Memdora]] is the clearest example of spacing treated as an engineering problem. It grounds scheduling in the Ebbinghaus forgetting curve, cites the figure that roughly **70% of newly learned material is forgotten within 24 hours** without review, and integrates **FSRS-6**, described as the current state-of-the-art spaced-repetition algorithm. Its argument is that scheduling alone is not enough: existing tools reduce flashcard interaction to a single binary gesture, "flip and self-rate", which the authors call an impoverished model that fails to exploit cognitive-science evidence on retrieval practice. Memdora's contribution is therefore a **taxonomy of 17 cognitively-grounded interaction types** across Language, By Heart and Exam categories, each mapped to peer-reviewed evidence displayed on the card, plus an effort-based reward system that compensates actual cognitive engagement rather than app presence, a unified generation pipeline that creates cards at the point of reading, and a classroom layer that reports [[learning-gains|learning outcomes]] at the individual card level. [[adaptive-pretesting-retention|Akgun and Toker (2026)]] supply the corpus's strongest experimental evidence on how spacing is *structured*. In a three-arm randomized study of 89 undergraduates in an applied statistics course, all three groups practiced in spaced sessions with identical number and timing over seven weeks; only the interaction structure differed. Adaptive spaced retrieval (G1) reached the highest posttest score (M = 78.19) and the highest observed practice effort (M = 0.85); fixed spaced retrieval (G2) followed (M = 74.55, effort 0.74); learner-directed AI study with no enforced retrieval came last on both (M = 67.28, effort 0.49). The multivariate effect was significant (Wilks' Λ = .664, partial η² = .185), with G1 ahead of G3 on retention at d = 0.92 (p = .003). The detail that matters pedagogically is that G1's agent was configured to *withhold*: response-contingent probes for [[misconceptions]], requests for elaboration after superficial answers, advancement only on adequate conceptual engagement, and direct solutions explicitly excluded from its allowed outputs. The authors' own reading is that pretesting is a front-loaded catalyst rather than a standalone intervention whose benefits survive open-ended AI access. Two further corpus findings sharpen what spacing does and does not yet get right. [[schuetze-knowledge-tracing-forgetting-2026|Schuetze, Yan and Carvalho (2025)]] fit [[knowledge-tracing]] models to a six-session successive-relearning dataset and found that BKT, BKT-with-Forgetting and the Additive Factors Model reproduce learning trends when fit retrospectively (AUC 0.74–0.79) but, under time-based cross-validation of the kind a deployed tutor would need, overestimate future performance by roughly 58%, 51% and 47% respectively and **fail to capture the spacing effect** — sometimes predicting the opposite ordinal ordering across spacing conditions. In other words, the scheduling decisions made by [[adaptive-learning|adaptive systems]] may be based on models that do not represent the very effect spacing depends on. And [[nie-personavlm-long-term-personalization-2026|LLM student modeling and long-term memory]] names the design gap directly: most tutoring systems lack longitudinal memory, and how such memory should interact with spaced repetition and forgetting curves remains an open question. ### What the corpus does not yet show Neither applied system has produced retention evidence of its own: the Memdora article reports improved retention relative to traditional spaced-repetition tools, but describes no delayed-interval outcome study behind that claim, and the bilingual lecture companion — which generates flashcards automatically to remove the authoring cost that blocks evidence-based retrieval and spaced practice, while simplifying scheduling to a binary "know it / still learning" rating rather than the graded SM-2 and FSRS algorithms — states plainly that **no learning-outcomes study has been run**, only a pre-registered protocol. The scheduling machinery is better evidenced than the learning it is meant to produce. ## Interleaving Interleaving — mixing problem or item types within a practice session rather than blocking them by type — is the weakest-evidenced of the three in this corpus, and that should be stated rather than padded. Interleaving appears on [[desirable-difficulties]] only as one instance in a list of effortful conditions, and no study in the knowledge base tests an interleaving manipulation against a blocked control. What the corpus does contain is adjacent and worth distinguishing carefully. [[simulating-learner-task-selection|Simulating learner task-selection]] models **Interleaving** and **Blocking** as two of eight candidate learner strategies in a [[mastery-learning]] [[simulation]], alongside strength targeting, weakness targeting and outcome-informed rules. The finding there is about choice architecture rather than memory: some learner strategies systematically delay progression through over-practice, and task-selection constraints repaired the maladaptive strategies while leaving Interleaving, Blocking, Strength Targeting and Weakness Targeting at steady over-practice levels. That is evidence that interleaving is a *selectable behavior* a system can accommodate or constrain — not evidence that interleaving improves retention. The remaining mentions in the corpus are unrelated senses of the word (interleaved agent stages in [[knowloop-confusion-to-consolidation-2026]]; interleaved visual-textual solution trajectories in a geometry [[benchmark]]), and they should not be counted as interleaving evidence. Interleaving's inclusion here therefore rests on the same theoretical family as the other two techniques and on its status in the [[cognitive-psychology]] literature, not on a study in this knowledge base. Treat it as the open slot. ## How AI Implements and Undermines These Techniques **Implementation.** AI's most defensible contribution is scheduling and generation, the parts of these techniques that are tedious rather than pedagogically interesting. Adaptive scheduling is implemented in Memdora via FSRS-6 and in knowledge-tracing architectures that estimate mastery over time; adaptive *pretesting* is implemented by Akgun and Toker's response-contingent agent, which reads prior-session performance signals to raise conceptual depth in weaker areas while maintaining challenge in stronger ones; automatic flashcard and pretest generation is implemented in the bilingual lecture companion, where card authoring is the cost that blocks spaced practice in the first place. [[nie-personavlm-long-term-personalization-2026|Longitudinal student memory]] is the architecture that would let any of this persist across semesters, and [[agentic-ai-pedagogical-best-practice-2026|Woollaston and colleagues (2026)]] supply the design vocabulary for doing it deliberately: intentional friction, dynamic [[scaffolding]], [[human-in-the-loop-ai|human-in-the-loop]] oversight, and considered rather than maximal AI utilization. **Undermining.** The failure mode is specific and documented repeatedly. [[agentic-ai-pedagogical-best-practice-2026|Woollaston et al. (2026)]] name the first of their six [[pedagogy|pedagogical]] risks as prior-knowledge activation: [[agentic-ai|agents]] that pre-fetch and surface content **bypass the retrieval practice that activates prior knowledge**, so the learner never recalls or integrates what they know before receiving an answer. [[prior-knowledge]] states the same bypass risk as a general property of [[generative-ai|generative AI]], and adds the design countermeasure — activate before you supply. The fluency illusion is the mechanism on the learner's side: an LLM supplies complete, coherent explanations on demand, so the learner allocates less effort to internal retrieval and reconstruction while the session *feels* productive, which is precisely the dissociation [[llm-interaction-depth-task-quality-recall-2026|Tsiligkiris (2026)]] measured. The three-arm pretesting study makes the same point experimentally: the arm with free AI access and no enforced retrieval performed worst on end-of-semester retention, and interaction volume did not differ significantly across arms, so the gap was produced by structure rather than by how much students engaged. **Preserving retrieval demand.** The corpus converges on a small set of design features. The pretesting agent's policy is the most concrete: withhold direct solutions by configuration, probe misconceptions, request elaboration after superficial attempts, and advance only on adequate conceptual engagement. Tsiligkiris recommends embedding retrieval demands *after* LLM use through closed-tool outputs — short-answer questions, concept maps drawn from memory, [[learning-by-teaching|teach-back]] explanations without the model — and separating scaffolding from checking, so the model is available for clarification and [[feedback]] but distinct checkpoints require independent recall. [[knowloop-confusion-to-consolidation-2026|KnowLoop]] shows the teach-back form in a deployed system: a Peer agent scaffolds reflective teach-back that surfaced gaps learners could not articulate, and its participants asked to move fluidly between clarification and consolidation rather than treating them as strict phases. [[student-cheat-sheets-make-or-take|Chen, Sakhnini and Istead]] add the assessment-side version — value the construction process rather than the artifact — while [[learning-to-learn-in-the-age-of-generative-ai-a-scoping-review-and-conceptual-fr|Schorr and colleagues (2026)]] position these tools as learnable skills that reduce overreliance and support [[agency|learner agency]]. Underneath all of them is the one non-negotiable: something must require the learner to produce an answer that the model has not already given. ## Limits of the Evidence - **Small samples and single contexts.** The pretesting study's analytic sample is 89 undergraduates in one applied statistics course, with 27–34 students per arm; the LLM interaction study has n = 22 with a single-group design that supports no causal claim; the teach-back study has 22 participants. The pretesting authors state that replication across domains and task types is needed for generalizability. - **Short or absent retention intervals.** The interaction-depth study tested immediate recall only, with no delayed measure, so its null is a null for short-term retention specifically. Memdora's retention advantage and the [[llm-interaction-depth-task-quality-recall-2026|LLM-depth dissociation]] alike are not tracked over the weeks that the spacing literature would require; the pretesting study's seven-week window is the longest in the corpus. - **[[self-report-measures|Self-report]] and proxy measures.** Practice effort in the pretesting study is a behavioral indicator scored by two raters against a rubric from conversation logs, which the authors explicitly note is not a measure of internal [[motivation|motivational]] state; the LLM-depth measure is a keyword-based proxy capturing the surface form of explanation-seeking rather than its quality. - **Instrumentation and modelling limits.** [[knowledge-tracing]] models in the corpus fail exactly where spacing decisions need them, and the authors attribute past successes to retroactive full-dataset fitting rather than deployment-realistic validation. The bilingual companion's flashcard quality is unverified against ground truth, with no automated fact-checking layer, and it has run no outcomes study at all. - **Interleaving has essentially no direct evidence here.** As stated above, the corpus contains no blocked-versus-interleaved comparison, only a simulation in which Interleaving is one selectable strategy among several. - **Subject populations and content.** The evidence base concentrates on [[higher-ed|higher education]], on STEM and statistics material, and on fact-versus-skill distinctions from geometry and multiple-regression content; generalization to [[k-12|K-12]] settings, to [[humanities-education|humanities]] and writing, and to non-Western classrooms is not established by these studies. ## Connected Concepts - [[desirable-difficulties]] — the principle these three techniques operationalize - [[cognitive-psychology]] — memory, encoding and retrieval as the theoretical home of all three - [[prior-knowledge]] — retrieval practice as activation of what the learner already holds - [[metacognition]] — judging whether fluency reflects learning - [[self-regulated-learning]] — studying consistently what an AI can schedule or answer for you - [[mastery-learning]] — retention after the mastery threshold, and the spacing that sustains it - [[knowledge-tracing]] — the models that make or miss scheduling decisions - [[cognitive-offloading]] — the failure mode when retrieval is delegated to a tool - [[transfer-of-learning]] — the performance–learning gap the techniques address - [[formative-assessment]] — closed-tool checkpoints as retrieval events - [[productive-failure]] — attempting before receiving help - [[learners]] — the stakes of durable rather than fluent learning ## Connected Articles - [[adaptive-pretesting-retention]] — Akgun & Toker: adaptive spaced retrieval vs fixed retrieval vs learner-directed AI over seven weeks - [[memdora-ai-spaced-repetition]] — Memdora: FSRS-6 scheduling and a taxonomy of grounded flashcard interactions - [[llm-interaction-depth-task-quality-recall-2026]] — Tsiligkiris: explanation depth predicts task quality but not immediate recall - [[rachatasumrit-example-problem-ratio-2026]] — Retrieval practice for facts, worked examples for skills; the content–treatment interaction - [[schuetze-knowledge-tracing-forgetting-2026]] — Knowledge-tracing models fail to capture the spacing effect under time-based validation - [[bilingual-llm-lecture-companion-srl-2026]] — Automatic flashcard generation removing the authoring cost of spaced practice - [[student-cheat-sheets-make-or-take]] — Constructing a cheat sheet as generative study; offloading the construction with AI - [[knowloop-confusion-to-consolidation-2026]] — Teach-back consolidation surfacing gaps clarification did not - [[simulating-learner-task-selection]] — Interleaving and Blocking as modeled learner strategies under task-selection constraints - [[nie-personavlm-long-term-personalization-2026]] — Longitudinal student memory and its open relation to spaced repetition and forgetting curves - [[learning-to-learn-in-the-age-of-generative-ai-a-scoping-review-and-conceptual-fr]] — Retrieval practice in the Tools layer of a learning-to-learn framework - [[agentic-ai-pedagogical-best-practice-2026]] — Pre-fetching agents bypass retrieval practice; the case for intentional friction - [[ai-tutor-modality-randomized-field-experiment-2026]] — When AI Tutors Speak: Evidence from a Randomized Field Experiment --- ## [Student Engagement](https://edtechdev.github.io/aied/concepts/student-engagement/) > **Student engagement** — the degree and quality of a learner's active involvement in the learning process, most often decomposed into behavioral, cognitive, and [[affective-computing|affective]] dimensions. In [[ai-education]] research, student engagement is both a key outcome (does an AI tool keep students engaged?) and a mechanism (does engagement mediate between AI design and learning?). It is conceptually distinct from learning itself — engagement is participation in learning, not proof of cognitive gain — and from the specific metrics used to measure it. ## Questions to Consider - The page insists engagement is not the same as learning — a student can be behaviorally active (clicking, spending time) while cognitively shallow. Where have you seen high 'engagement' that produced little learning, and how did you notice? - Think of a moment you were deeply cognitively engaged in something — truly wrestling with an idea. What was different about it compared to times you were merely busy or entertained, and could an AI tool reliably create that state? - Engagement is broken into behavioral, cognitive, and affective dimensions that can diverge. Why do you think researchers insist on treating these separately rather than as one thing, and what would you measure to tell them apart? - The research suggests deep cognitive engagement with AI predicts learning, while shallow engagement predicts over-reliance. If a tool is 'engaging' but shallow, who is at fault — the design, the task, or the learner? - How might an AI tool satisfy [[agency|autonomy]], competence, and relatedness (the needs behind engagement) without those features turning into shallow entertainment that displaces real learning? ## Introduction Engagement is a multidimensional construct rooted in educational psychology. **Behavioral engagement** refers to participation, effort, persistence, and on-task activity. **Cognitive engagement** refers to the depth of mental processing — elaboration, [[critical-thinking|critical analysis]], self-[[regulation]], and the investment of mental effort. **Affective engagement** refers to emotional reactions such as interest, enjoyment, anxiety, and identification with learning. These dimensions can diverge: a student may be behaviorally active (clicking, spending time) while cognitively shallow (passively accepting output), or affectively interested while behaviorally distracted. This multidimensionality is why engagement must not be equated with any single observable behavior. ### How student engagement appears in the research - **Engagement as an outcome of AI design:** [[genai-motivation-engagement-2026|GenAI motivation research]] shows that engagement in [[generative-ai]]-supported learning follows the satisfaction of basic psychological needs ([[self-determination-theory|autonomy, competence, relatedness]]) — engagement is the downstream result of motivational support, not of technology availability alone. Engagement is often measured by [[self-report-measures|self-report]] while behavioral engagement comes from interaction logs, and the two are not interchangeable. - **Engagement gain without a learning gain, in a cluster-randomized classroom trial:** [[domain-specific-chatbot-stem-enthusiasm-2025|Rücker and Becker-Genschow (2025)]] randomized 195 ninth-grade classes (experimental group 102 students) to a domain-specific mathematics chatbot or to conventional differentiation materials for a single lesson on the Heron method of estimating square roots. Situational interest rose substantially in the chatbot condition (M = 2.63 vs 2.43, p = 0.00005, Cohen's d = 0.63) and acceptance on all four [[technology-acceptance-model|Technology Acceptance Model]] dimensions was high, yet the pre-post performance comparison found no significant group-by-time interaction (F(1194) = 2.84, p = 0.094) and extrinsic cognitive load was slightly higher. The study is one of the few cluster-randomized tests of a customized chatbot in [[k-12|secondary]] [[math-education|mathematics]], and its split result is the point: interest and [[learning-gains|achievement]] moved on different schedules, so an engagement finding is not evidence of learning. - **Quality over quantity:** [[critical-engagement-code-completion|Critical engagement in AI code completion]], [[icap-cognitive-engagement-llm-agents|cognitive-engagement discourse analysis]], and [[scaffolding-critical-engagement-genai-minority-students|scaffolding critical engagement]] show that *deep* (cognitive) engagement with AI predicts learning, while *shallow* (behavioral) engagement predicts the [[cognitive-offloading|Over-Reliance]] and learning displacement that dominate the knowledge base's risk literature. - **Fragile and context-dependent:** [[polished-artifacts-fragile-engagement-2026|Polished artifacts, fragile engagement]] and [[genai-tutor-engagement-patterns|multi-institution engagement patterns]] find engagement varies by task, context, and learner — an AI tool that engages one student deeply may produce shallow, output-chasing behavior in another. - **Motivational antecedents:** [[ai-availability-student-motivation|AI availability and motivation]] shows that knowing AI is available can reduce the perceived value of effortful engagement, particularly for novice learners — engagement is shaped by expectancy, value, and perceived competence as much as by tool features. **[[wang-goal-setting-ai-engagement-2026|Wang & Wang (2026)]]** extend this with a goal-setting-theory account of **758 university English learners** in AI-assisted learning, showing that **teacher support** directly enhances engagement and operates through students' **mastery-approach and performance-approach goals** (rather than avoidance goals). Engagement in AI contexts is therefore not only an individual or design outcome — it is also **socially scaffolded** by the teacher and by the goal orientations learners are encouraged to adopt. - **Competency and emotion as engagement drivers:** [[chatbot-engagement-genai-competency-emotion-2026|Zhao et al. (2026)]] model **871 university students** interacting with an [[llm]] chatbot, finding that **GenAI competency** predicts chatbot engagement both directly and indirectly through **positive emotions** (the affect pathway), and that both competency and positive emotion predict engagement and positive learning emotions. Engagement is thus jointly a *skill* and an *affective* outcome — learners who lack [[teacher-ai-competency|AI competency]] and experience anxiety or frustration disengage, which has implications for [[ai-literacy]] training as an engagement intervention rather than merely a skill goal. - **Discipline-associated cognitive engagement in student-AI chat.** [[student-ai-conversations-cognitive-engagement-2026|Chang and Li (2026)]] show that student prompts to AI encode ~62% higher-order cognitive demand on average, but Bloom-level engagement profiles differ sharply by discipline ([[stem-education|STEM]] Apply-prevalent 20.8%, language Understand-prevalent 31.7%, social science Create-prevalent 33.8%). Using a within-person design, they found the same students produced significantly more higher-order prompts in social science than STEM courses (p < .001), with course-level variation exceeding student-level variation — evidence that cognitive engagement with AI is shaped by disciplinary context, not just individual style. - **[[ai-feedback-quality|AI feedback]] sustains behavioral activation.** [[gpt4-feedback-student-activation-2026|Geschwind et al. (2026)]]'s semester-long lab-in-the-field experiment found that students receiving individual GPT-4 feedback on open-ended tasks sustained the highest participation across eight weekly tasks (~50% by the last, vs ~30% for lecturer- or peer-feedback groups) and wrote ~29 more characters per answer — individual AI feedback activated engagement on both the extensive margin (participation) and the intensive margin (effort per response), even though students rated the [[peer-assessment|peer feedback]] slightly higher. - **Predicting academic AI use from learning constructs.** An exploratory [[reinforcement-learning|machine learning]] framework analyzed survey data from 166 university students to identify learning-related constructs associated with intended academic ChatGPT use, using SHAP analysis to maintain [[explainable-ai|interpretability]]. Findings inform how engagement, learning support, and other constructs shape students' incorporation of AI tools into academic work. ### Measuring engagement: the metric-choice problem Engagement is operationalized through a range of observable signals. **Behavioral metrics** measure what learners *do* (time-on-task, activity counts, interaction frequency, persistence); **cognitive metrics** measure how learners *think* (depth of processing, critical engagement, discourse analysis); **affective metrics** measure how learners *feel* (emotion, motivation, interest); and **contextual metrics** capture multitasking and attention. AI-education research increasingly combines these and treats engagement as a mediating mechanism between AI tool design and [[learning-gains|learning outcomes]], rather than an outcome in itself. - **Six AI application families + multi-method measurement:** [[ai-student-engagement-online-learning-review-2025|Zhou's (2025) systematic review]] of 24 WoS studies maps six AI applications for engagement — [[conversational-ai|chatbots]] in course design, emotion/facial/voice recognition and eye tracking, ML for data analysis, teacher–student interaction support, personalized feedback/recommendations, and AI-powered bots in smart learning environments. It finds that integrating multiple AI modalities and data sources yields more accurate, real-time insight into cognitive, emotional, and behavioral engagement than single-source approaches — reinforcing the metric-choice problem above. The choice of metric is definitional: a study that measures engagement as *time-on-task* may conclude an AI tool enhances engagement when students spend more time interacting with it, while a study that measures engagement as *critical processing* may reach the opposite conclusion for the same tool. This is why the knowledge base's research distinguishes engagement (participation) from learning (actual cognitive gain) — see [[genai-performance-vs-learning|performance vs. learning]] — and why engagement metrics must be validated against what they claim to measure. - **Engagement as a fragile, situation-dependent signal:** [[polished-artifacts-fragile-engagement-2026|Polished artifacts, fragile engagement]] and [[genai-tutor-engagement-patterns|multi-institution engagement patterns]] find that engagement varies by context, task, and learner — the same tool produces strong engagement for some students and shallow, output-chasing behavior for others. - **Behavioral telemetry from learning platforms:** [[engagement-forecasting-its|Effort and progress forecasting]], [[learning-engagement-assistant-lea|Learning Engagement Assistant]], [[engagement-assessment-video|video engagement assessment]], and [[interactive-learning-dashboards-engagement|learning dashboards]] translate behavioral and physiological signals (attention, activity, persistence) into engagement metrics used for adaptive feedback and instructor intervention. - **Physiological sensing adds a modality — and a baseline problem.** [[e3sense-multimodal-learner-engagement-sensing-2026|E3Sense]] co-locates dry-electrode EEG, eye-tracking glasses, and forehead electrodermal electrodes on the head and predicts 450 segment-level engagement ratings from 30 university participants on a five-level ordinal scale: AdaBoost over the fused [[multimodal]] representation reached 75.0% balanced within-one-level accuracy against 63.0% for always predicting the most common rating. The narrowness of that gap is the point — within-one credit hands a sensor-free baseline most of its score on skewed ratings — and asking learners what engagement means to them moved the same measure from 64.6% to 71.5%, evidence that the [[self-report-measures|self-report]] label, not only the sensor, decides what such analytics can claim. - **Engagement as a learner-modeling signal:** [[engagement-intensity-learner-modeling|Engagement intensity as a learner-modeling signal]] uses engagement strength to inform adaptive AI systems, positioning engagement metrics as inputs to [[student-modeling]] and [[adaptive-learning]] rather than merely evaluation outputs. ### Engagement vs. learning A central theme in the knowledge base's research is that engagement and learning must be distinguished. AI tools that generate high engagement (time on task, interaction volume) may not produce learning if that engagement is passive or substitutes for the [[cognitive-offloading|cognitive work]] of understanding — see [[genai-performance-vs-learning|performance vs. learning]]. Conversely, productive struggle and [[desirable-difficulties|desirable difficulty]] can produce learning even when surface engagement feels lower. Engagement is therefore best treated as a *mechanism* — valuable insofar as it reflects or enables meaningful [[cognitive-psychology|cognitive processing]] — rather than a terminal outcome. The distinction is not academic. [[pramod-agentic-ai-motivational-pathways-2026|Pramod and Patil (2026)]] place engagement at the center of their PLS-SEM model, between motivation and social presence on one side and *perceived* performance on the other — the largest coefficient in their model, and still a perception rather than a measure of learning. ### Pedagogy mediates AI's effect on engagement A systematic synthesis of [[higher-ed|AI in higher education]] ([[long-ai-higher-ed-engagement-teaching-methods-2026|Long et al., 2026]]) emphasizes that the **[[teacher-role|teaching]] method an AI tool is embedded in is the decisive mediator** of whether it engages students. Chatbots, adaptive systems, and predictive analytics enhance engagement most when deployed within interactive pedagogies — flipped classrooms, [[project-based-learning|project-based learning]], and scaffolded [[feedback|feedback loops]] — rather than as standalone tools. The review formalizes this as the **PMAISE model** ([[pedagogy|Pedagogical]] Mediation of AI for Student Engagement), mapping the alignment between AI [[ai-technologies|technologies]], pedagogical strategies, and the affective, behavioral, and cognitive dimensions of engagement. The implication is that engagement outcomes are co-produced by the tool *and* the surrounding [[learning-design|instructional design]]: the same AI can amplify engagement in one pedagogy and inhibit it in another. ### Connections to related concepts Student engagement connects to [[motivation]] and [[self-determination-theory]] as its psychological drivers, and to [[student-experience]] as the lived context. Its measurement relies on [[learning-analytics]] and [[educational-measurement]], which supply the [[quantitative-research|quantitative]] tools for operationalizing the dimensions above. The distinction between deep and shallow engagement ties directly to [[self-regulated-learning]] (self-regulated learners engage strategically), [[cognitive-offloading]] and [[cognitive-offloading|Over-Reliance]] (shallow reliance as the failure mode), and [[metacognition]]. In system design, engagement signals feed [[student-modeling]] and [[adaptive-learning]], and engagement outcomes feature in [[research-methods-aied]] evaluations of AI-education interventions. - **Learner characteristics moderate TTS dialogue-based lessons (2026):** In LLM+TTS-generated teacher–student, student–student, and teacher–teacher dialogue lessons, [[experiential-learning]] style and critical-thinking disposition significantly interacted with dialogue format for ARCS-based motivation, indicating that AI-generated dialogue content is differentially motivating depending on learner profile ([[tts-dialogue-lessons-learner-characteristics-2026]]). - **Dimension-specific gains at the primary level (2026):** A nine-week GenAI-supported L2 [[writing-education|writing]] program with 301 Grade 5 and 6 students raised behavioral and emotional engagement but left cognitive and metacognitive engagement unchanged, and its authors name reduced self-monitoring during writing as a standing risk ([[genai-writing-program-primary-l2-motivation-engagement|Lu et al., 2026]]). The pattern is a concrete instance of the engagement-versus-learning distinction above: more activity and more enjoyment did not translate into deeper processing. - **The dissociation can run the other way (2026):** In a 12-week vocational interior-design course, an immersive VR studio with an embedded LLM teaching assistant raised cognitive (d = 0.90) and behavioral (d = 0.75) engagement over traditional [[project-based-learning|project-based]] instruction while affective engagement did not differ significantly (d = 0.38) — the inverse of the L2 writing case above ([[ai-ive-pbl-vocational-design-creativity-2026|Jin et al., 2026]]). Cognitive and behavioral gains here came with *lower* reported cognitive load, which the authors attribute to the assistant absorbing search and cross-disciplinary integration effort. Read against the writing case, the two studies suggest that which engagement dimension an AI-supported intervention moves is a property of the design — discourse-heavy immersive [[collaborative-learning|collaboration]] versus solo writing support — rather than of AI assistance in general, and that an affective advantage cannot be assumed from high-fidelity or intelligent feedback. - **AI literacy works on engagement through psychological resources (2026):** A moderated mediation study of 1,198 undergraduates in Zhengzhou, China ([[ai-literacy-learning-engagement-psych-capital-2026|Wang, 2026]]) modeled engagement as an outcome of [[ai-literacy]] rather than a by-product of tool use. AI literacy predicted learning engagement directly and also indirectly by building psychological capital, with the indirect route carrying roughly half of the total effect — partial mediation, so a technological competency converts into engagement only partly through the psychological resources it generates. Professional commitment, an identity-based variable, moderated the psychological-capital-to-engagement link without any direct effect of its own, and the translation of psychological capital into engagement was markedly stronger for students who saw themselves as headed into the profession. The pattern is the clearest available instance of the point above that learner characteristics condition how AI affects engagement. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[community-of-inquiry]] — Community of Inquiry (agentic engagement as a CoI dimension) - [[eportfolio]] - [[online-teaching-and-learning]] — Online Teaching and Learning - [[motivation]] - [[self-determination-theory]] - [[student-experience]] - [[learning-analytics]] - [[educational-measurement]] - [[self-regulated-learning]] - [[cognitive-offloading]] - [[metacognition]] - [[student-modeling]] - [[adaptive-learning]] - [[research-methods-aied]] - [[higher-ed]] - [[framing-ai-use-for-students]] - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) - [[self-report-measures]] - [[productive-failure]] ## Connected Articles - [[e3sense-multimodal-learner-engagement-sensing-2026]] — Head-confined EEG, eye tracking, and EDA predict five-level engagement ratings, while learners' own definitions shift the mapping (Anupkrishnan et al. 2026) - [[student-attention-estimation-fairness-2026]] — Fairness-Aware Multimodal Transformer Modeling for Real-Time Student Attention Estimation - [[ai-student-engagement-online-learning-review-2025]] - [[ai-online-education-engagement-satisfaction-2026]] - [[long-ai-higher-ed-engagement-teaching-methods-2026]] — AI in higher ed: systematic review of engagement + mediating role of teaching methods - [[genai-motivation-engagement-2026]] — Impact of Generative AI on Student Motivation and Engagement - [[critical-engagement-code-completion]] — To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion - [[icap-cognitive-engagement-llm-agents]] — Measuring Cognitive Engagement in Collaborative Discourse - [[genai-tutor-engagement-patterns]] — Not All Students Engage Alike: Multi-Institution Patterns - [[polished-artifacts-fragile-engagement-2026]] — Polished Artifacts, Fragile Engagement - [[ai-availability-student-motivation]] — "Why Put in This Much Effort?": How AI Availability Shapes Motivation - [[genai-performance-vs-learning]] — Distinguishing Performance Gains From Learning - [[scaffolding-critical-engagement-genai-minority-students]] — Scaffolding Critical Engagement With GenAI - [[engagement-intensity-learner-modeling]] — Engagement Intensity as a Learner-Modeling Signal - [[learning-engagement-assistant-lea]] — Learning Engagement Assistant - [[engagement-assessment-video]] — Engagement Assessment in Video Learning - [[engagement-forecasting-its]] — From Heuristics to Analytics: Forecasting Effort and Progress - [[interactive-learning-dashboards-engagement]] — Interactive Learning Dashboards and Engagement - [[young-people-learning-generative-ai-rapid-review-2026]] — Affective gains common but weak indicators of learning - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions - [[wang-goal-setting-ai-engagement-2026]] — Goal-setting theory: teacher support, achievement goals, and engagement in AI-assisted English learning (758 Chinese students) - [[student-motivation-need-satisfaction-genai-sdt-2026]] — Student motivation and need satisfaction in GenAI classrooms (Schweder, Hagenauer & Raufelder 2026) - [[chatbot-engagement-genai-competency-emotion-2026]] — GenAI competency and emotion as drivers of chatbot engagement (Zhao et al. 2026) - [[student-ai-conversations-cognitive-engagement-2026]] — Discipline-associated Bloom-level cognitive engagement in student-AI conversations (Chang & Li 2026) - [[determinants-chatgpt-use-higher-education-2026]] — ML/SHAP determinants of future ChatGPT use in higher education - [[gpt4-feedback-student-activation-2026]] - [[genai-writing-program-primary-l2-motivation-engagement]] — Dimension-specific engagement gains at the primary level (Lu et al. 2026) - [[ai-ive-pbl-vocational-design-creativity-2026]] — Cognitive and behavioral engagement up, affective engagement flat, in an immersive VR PBL studio (Jin et al. 2026) - [[domain-specific-chatbot-stem-enthusiasm-2025]] — Cluster-randomized secondary mathematics trial: situational interest rose with a customized chatbot while test performance did not (Rücker & Becker-Genschow 2025) - [[ai-literacy-learning-engagement-psych-capital-2026]] — AI literacy drives engagement directly and through psychological capital, amplified by professional commitment (Wang 2026) --- ## [Help-Seeking](https://edtechdev.github.io/aied/concepts/help-seeking/) > **Help-Seeking** — the learner's process of recognizing a need for assistance and strategically requesting it, and how that process plays out in AI-supported learning environments. In [[ai-education|AI in education]], help-seeking is central to whether AI tools support or undermine learning: the *quality* of help-seeking (when, how, and what learners ask for) strongly shapes outcomes, and AI tutors, hints, and [[pedagogy|pedagogical]] agents are designed precisely to elicit productive help-seeking rather than answer-seeking.([[lak2026-hint-button-unproductive-use]])([[ai-fallibility-warning-help-seeking]]) ## Questions to Consider - When you get stuck, do you tend to ask for a direct answer or for guidance that helps you figure it out yourself? What do you think each choice does to what you actually retain? - Research shows students often intend to learn with AI but default to asking for the answer — an 'intention-behavior gap' linked to worse performance. Why might good intentions so easily collapse into answer-seeking? - A persistent 'hint button' can turn a learning task into a copying exercise by signaling that help is always there. Can you recall a time when having help too easily available made you skip the thinking you needed to do? - One study found that simply warning students an AI could make mistakes actually increased their help-seeking. How might healthy skepticism change how students engage with a tutor versus blind trust? - Struggling students are often the least likely to seek help unprompted. If the students who most need support don't reach out, how should AI tools and instructors respond? - The page proposes delaying hints and moving the design question from 'whether' to 'how' to provide help. What would a well-designed help experience look like for your learners — and what would make them actually take it up? ## Introduction Help-seeking is a well-established construct in learning research, closely tied to [[self-regulated-learning]] and [[metacognition]]: it requires learners to monitor their own understanding, recognize a gap, decide help is needed, and formulate an effective request. With the rise of [[generative-ai|generative AI]] tutors, help-seeking has taken on new importance — and new failure modes. Learners often *intend* to use AI for learning but default to asking for direct answers, a gap that research in this knowledge base documents across domains and age groups.([[regulating-ai-tutor-adolescent-srl]])([[guided-llm-scaffolding-independent-learning]]) ## Productive vs. unproductive help-seeking The central distinction in the literature is between help-seeking that supports learning and help-seeking that bypasses it. ### Unproductive help-seeking behaviors Research in this knowledge base identifies concrete, observable patterns of unproductive help-seeking, especially in [[intelligent-tutoring|intelligent tutoring systems]]: - **Premature hint requests** — requesting help before making any solution attempt. Even uncertain students learn more by attempting first.([[lak2026-hint-button-unproductive-use]]) - **Superficial hint reading** — advancing through hints too rapidly to read them (flagged at a ~4 words/second benchmark), often jumping straight to the bottom-out hint that reveals the answer.([[lak2026-hint-button-unproductive-use]]) - **Answer-seeking over learning-seeking** — asking the AI to produce the answer rather than to explain or guide. In a study of 98 Grade-9 students using a GenAI tutor, interactions were dominated by instrumental requests with almost no monitoring or evaluation of their own learning — despite students having chosen scaffolded support beforehand. This **intention-behavior gap** was associated with *lower* post-test performance and higher extraneous cognitive load.([[regulating-ai-tutor-adolescent-srl]]) - **Consulting AI before any independent attempt or human source.** [[uneven-impact-generative-ai-student-learning-2026|Manikonda et al. (2026)]] measure this ordering directly as **early reliance** — consulting GenAI before independent thinking, a traditional search, or reaching an instructor — and find it associated with greater negative impact (β = .402, p = .004) as well as academic benefit (β = .301, p < .001) among 118 students in AI-related courses. The association with harm was absent at low [[ai-literacy|evaluation literacy]] and strongest at high evaluation literacy (b = .688 at +1 SD, p < .001), so students most able to judge AI output reported the most cost from consulting it first: the choice of *whom to ask first* carries a downside that skilfulness at evaluating the answer does not offset. It also shows that using AI for organizing, evaluating, and decomposing problems — **cognitive** rather than early reliance — is the pattern associated with positive impact, so the help-seeking failure mode is one of sequencing rather than of asking at all. - **Struggling students are least likely to seek help unprompted** — the engagement side of help-seeking. In [[one-click-away-khanmigo-two-year-school-experiment-2026|a two-year Khanmigo RCT (Oreopoulos & Low 2026)]], even with free access and mandatory practice time, the median struggling student messaged the AI tutor in only ~17% of mistake sessions, mostly with bare answers or clicks — consistent with the economics-of-education finding that initiative-dependent interventions reach fewest of the students who would benefit most. [[virtual-tutoring-computer-assisted-learning-takeup-2026|TWiK (Oreopoulos et al. 2026)]] shows take-up is highly responsive to reducing friction (first-session take-up rose 45%→83% after simplifying enrollment), but entry ≠ sustained participation (attendance stayed intermittent). ### Why unproductive help-seeking hurts learning The **affordance perspective** explains a key mechanism: when an interface makes help constantly and saliently available (e.g., a persistent "hint button"), it signals to learners that help is always there, creating an unintended affordance that can collapse the task into a copying exercise. Rapidly accessing bottom-out hints circumvents the active schema construction that learning requires.([[lak2026-hint-button-unproductive-use]]) ### The quality of help-seeking is measurable Two simple, interpretable indicators — premature hint requests and superficial hint reading — are computable from standard tutoring logs and are consistently associated with reduced [[learning-gains|learning gains]] across semesters, even after controlling for [[prior-knowledge|prior knowledge]]. This makes them practical for [[learning-analytics]] [[visualization|dashboards]] and real-time intervention, unlike complex machine-learned "gaming the system" detectors.([[lak2026-hint-button-unproductive-use]]) ## Designing AI systems to promote productive help-seeking ### Scaffolding how students ask Explicit training in **reasoning-focused help-seeking** — requesting stepwise hints and verification rather than final answers — produces better outcomes than uncritical reliance. In a quasi-experimental undergraduate statistics study, guided LLM access (with training on reasoning-oriented help-seeking) led to stronger independent performance and better self-assessment calibration than unrestricted LLM access. The lesson: **LLM access alone is an incomplete intervention**; the design challenge is to scaffold *how* students use AI so it functions as a reasoning partner rather than an answer-getting tool.([[guided-llm-scaffolding-independent-learning]]) Interaction cost is part of the same question. [[penquiry-pen-based-llm-qa-2026|Rhee et al. (2026)]] identify a **Referential Barrier** and an **Expressive Barrier** that stop pen-based learners from asking an [[llm|LLM]] anything at all: pointing at a diagram region or an equation term cannot be expressed in typed prose, and the effort of formulation lands exactly when a question is most fragile. Their Penquiry system resolves reference by snapping ink marks to document elements and expands sparse ink keywords into full queries through autocompletion; two iterative studies of 16 participants each found the cognitive and physical overhead of inquiry fell significantly. Whether lower asking cost produces *better* help-seeking or merely more of it is left open, and the authors propose temporally adaptive autocompletion — foundational verification early in a session, higher-level prompts later — as a route from reduced friction to [[scaffolding|fading support]] rather than a permanent crutch. A third lever on the cost of asking is *where* the help comes from. [[course-specific-rag-help-seeking-higher-ed-2026|Gray and Hobbs (2026)]] built Beacon, a course-specific [[rag|retrieval-augmented]] assistant grounded in one programming module's approved materials, and evaluated it with 15 computing students and four academics. 89% of participants rated its answers highly aligned with course materials and 66.7% said it supported rather than replaced their learning, though only around half to 60% reported gains in understanding or confidence. The motivation is the barrier this section documents: 62.5% of those students said they sometimes avoided asking for help when they needed it and 75% reported anxiety when a topic did not make sense, so a private, module-grounded channel is offered as a first rung before approaching a lecturer. The academics interviewed kept the counter-argument alive — they valued that Beacon withheld full solutions and worried that unrestricted tools let students skip a development stage — which is why the design earns its place by refusing to complete the work. ### Calibrating trust through transparency A classroom experiment with 252 students found that **warning students about AI fallibility increased help-seeking** in a math tutoring system. Transparency about potential system errors improved learners' engagement with the system — connecting help-seeking to [[trust-calibration]] and [[hallucination-risk]].([[ai-fallibility-warning-help-seeking]]) ### Rethinking hint and scaffold delivery Rather than removing help, research recommends re-engineering how it is delivered: - **Delayed hint availability** — requiring minimum engagement time or solution attempts before hints (especially bottom-out hints) are accessible.([[lak2026-hint-button-unproductive-use]]) - **Moving from *whether* to *how*** — the key design question is how to structure hint delivery aligned with productive-struggle principles, not whether to provide hints at all.([[lak2026-hint-button-unproductive-use]]) ### The uptake problem in LLM tutors Real-world students frequently **bypass a [[conversational-ai|chatbot]]'s [[scaffolding]]** — not necessarily harmfully, but often because there is a mismatch between the chatbot's pedagogical framing and the student's own learning goals. Evaluation pipelines must therefore measure not just whether a tutor scaffolds, but whether students *take up* that scaffolding, rather than assuming they will.([[rethinking-scaffolding-llm-tutors]]) ## Help-seeking and self-regulated learning Help-seeking is an integral part of [[self-regulated-learning]]: productive help-seeking requires learners to monitor understanding, judge when help is needed, and select appropriate sources. In GenAI contexts, this becomes even more demanding, since students must also exercise [[agency]] over the AI and maintain epistemic vigilance rather than deferring to it. Research in this knowledge base supports the need for [[scaffolding|scaffolds]] that promote more [[agentic-ai|agentic]] and epistemically proactive AI use, and highlights the risk of [[cognitive-offloading|Over-Reliance]] and [[cognitive-offloading]] when help-seeking degrades into unconditional answer-seeking.([[regulating-ai-tutor-adolescent-srl]])([[guided-llm-scaffolding-independent-learning]]) ### LLM-mediated help-seeking as a four-stage process [[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026|Viberg et al. (2026)]] show that, in everyday STEM study, LLM help-seeking is not a single act but a layered, context-dependent process with four stages: (1) *deciding whether help is needed* — students try tasks independently first to preserve learning value; (2) *choosing whom to ask* — ChatGPT as a low-barrier first step, then peers for conceptual negotiation, then instructors for complex or high-stakes issues; (3) *determining the type of help* — from hints and explanations to scaffolding [[problem-solving]], streamlining routine work, and extending learning; and (4) *judging the help received* — exercising selective trust and verifying AI outputs against coursework or with humans. Crucially, students favored **instrumental help-seeking** (enhancing understanding) over **executive help-seeking** (obtaining solutions), a distinction that the authors propose adapting into new SRL-for-LLM measurement items. In fully online [[english-education|composition]], availability of the tool is not the bottleneck. [[reed-resource-literacy-genai-composition-2026|Reed (2026)]] observed that students who struggled were not the ones without support but the ones who did not recognize when help was needed, which resource fit the task, or how to judge feedback once it arrived — and fluent [[generative-ai]] output is easily mistaken for authoritative support. Her response was to make help-seeking *structured* rather than merely available: required touchpoints that decode task demands, map resources with justification, compare feedback sources, and close the loop with reflection. ### Making behavioral context visible: TutorTrace [[tutortrace-learner-behavioral-states-2026|Barron et al. (2026)]] tackle the behavioral precursor to help-seeking in [[cs-education|AI-assisted programming education]]: human tutors adapt to learners' observable behavior, not just their explicit requests, but AI tutors lack that context. **TutorTrace** is a dataset and pipeline that makes learners' behavioral context computable in real time from low-level IDE telemetry (four deployments, N=480; ~180K events, 13,633 behavioral segments, 27 metrics), deriving a taxonomy of learner activity *before* the first AI query, *between* consecutive queries, and *across* the session. This enables systems to classify whether a query reflects **guided** help-seeking (preceded by independent work) or **dependent** help-seeking (no independent work) — AUROC=.717 on held-out prediction — and to predict imminent queries (AUROC=.726). A preliminary classroom evaluation found that behavior-aware prompts reduced intervals between queries with no independent work from 50.0% to 20.7%. This connects [[learning-analytics]] telemetry to [[intelligent-tutoring|adaptive tutoring]], showing that behavioral context can be operationalized at scale to scaffold *how* students seek help rather than merely respond to their explicit questions. ## Implications for design and research 1. **Design help-seeking affordances deliberately.** Persistent, salient help buttons can enable bypass strategies; delay access and structure delivery to support [[desirable-difficulties|productive struggle]].([[lak2026-hint-button-unproductive-use]]) 2. **Scaffold the help-seeking itself.** Train learners in reasoning-focused requests (stepwise hints, verification) rather than assuming access equals good use.([[guided-llm-scaffolding-independent-learning]]) 3. **Use transparency to calibrate trust.** Warning about AI fallibility can increase appropriate help-seeking and engagement.([[ai-fallibility-warning-help-seeking]]) 4. **Measure uptake, not just scaffolding.** Evaluate whether students actually engage with pedagogical framing, not only whether the tutor provides it.([[rethinking-scaffolding-llm-tutors]]) 5. **Support monitoring and agency.** Help-seeking scaffolds should strengthen [[metacognition]] and [[self-regulated-learning]], guarding against [[cognitive-offloading|Over-Reliance]]. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[self-regulated-learning]] - [[metacognition]] - [[scaffolding]] - [[intelligent-tutoring]] - [[student-experience]] - [[cognitive-offloading]] - [[learning-analytics]] - [[k-12]] - [[higher-ed]] - [[socratic-method]] - [[pedagogical-agent]] - [[ai-literacy]] - [[trust-calibration]] - [[affective-tutoring]] - [[feedback]] - [[active-learning]] - [[agentic-ai]] ## Connected Articles - [[penquiry-pen-based-llm-qa-2026]] — Penquiry: A Pen-based Interactive In-situ Q&A System Leveraging LLMs - [[tutortrace-learner-behavioral-states-2026]] - [[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026]] — LLM-mediated help-seeking in STEM: layered, instrumental, and verified - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — Virtual tutoring with CAL: an experiment in take-up and learning - [[studychat-student-dialogues-chatgpt-ai-course-2026]] — The StudyChat dataset of student–LLM dialogues in an AI course - [[lak2026-hint-button-unproductive-use]] — Premature hint requests and superficial hint reading predict lower learning gains in an ITS - [[ai-fallibility-warning-help-seeking]] — Warning about AI fallibility increases help-seeking in a math tutoring system - [[regulating-ai-tutor-adolescent-srl]] — The intention-behavior gap in adolescent GenAI help-seeking and self-regulated learning - [[guided-llm-scaffolding-independent-learning]] — Guided LLM scaffolding improves reasoning-focused help-seeking and independent learning - [[rethinking-scaffolding-llm-tutors]] — The scaffolding/student-uptake mismatch in real-world LLM tutor deployments - [[surfacing-isolated-learners]] — Using AI to surface learners who need help, mediating teacher-student feedback - [[halani-designing-for-reach-2026]] — Designing for reach: the student alone with AI and access to help - [[uneven-impact-generative-ai-student-learning-2026]] — Early reliance: consulting GenAI before independent thought, search, or an instructor predicts both benefit and harm (Manikonda et al. 2026) - [[reed-resource-literacy-genai-composition-2026]] — Resource literacy in online composition: the bottleneck is recognizing when help is needed (Reed 2026) - [[course-specific-rag-help-seeking-higher-ed-2026]] — Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education - [[adaptive-scaffolding-contingency-comet-tutor-2026]] — Adaptive Scaffolding Needs Contingency: An AI Tutor That Escalates and Fades on What the Learner Does --- ## [Social-Emotional Learning](https://edtechdev.github.io/aied/concepts/social-emotional-learning/) > **Social-emotional learning (SEL)** — the process of developing the competencies that enable individuals to synchronize thoughts, emotions, and actions to foster positive interactions with oneself and others: self-awareness, self-management, social awareness, relationship skills, and responsible decision-making (the CASEL framework). In AI in education, SEL is increasingly recognized as critical because the rapid integration of [[generative-ai|generative AI]] into learning raises questions about students' [[well-being]], sociability, empathy, and trust — and because technical AI literacy alone is insufficient for navigating AI-mediated learning environments. ## Questions to Consider - The CASEL framework lists five competencies — self-awareness, self-management, social awareness, relationship skills, and responsible decision-making. Which of these do you think an [[intelligent-tutoring|AI tutor]] could support, and which do you suspect it cannot? - The page distinguishes social-emotional learning from emotional intelligence. Before reading on, how would you describe the difference, and why might the distinction matter for how schools approach them? - If technical AI literacy alone is 'insufficient' for navigating AI-mediated learning, what do you think is missing — and what does that imply for how you prepare students (or yourself)? - How might heavy reliance on AI reshape a student's sociability, empathy, or sense of trust — and are those changes something education should actively design for? - The [[research-methods-aied|research]] suggests SEL supports the relational dimensions of learning that AI must 'complement rather than replace.' Where have you seen technology strengthen a human relationship, and where has it quietly substituted for one? ## Introduction Social-emotional learning is closely related to, but distinct from, emotional intelligence (EI): SEL/SEC (social-emotional competencies) encompasses the ability to synchronize thoughts, emotions, and actions for positive interactions, while EI is an individual's capacity to process emotional information (conceptualized through ability models — reasoning and [[problem-solving]] — or trait models — emotional dispositions and behaviors). In the AI era, SEL matters because AI can reshape learning in ways that affect students' relational and emotional development, and because educators need both technological skill and emotional intelligence to support learners effectively. ### How SEL appears in the research - **Integrating SEC into AI literacy:** [[sec-ai-literacy-narrative-review-2026|The narrative review by Palmquist et al.]] proposes an integrated framework that combines AI literacy with social-emotional competencies, arguing that technical proficiency alone is insufficient — educators and students need both technological and emotional intelligence to navigate AI-mediated learning environments, fostering [[personalized-learning|personalized learning]], collaboration, and ethical [[student-engagement|engagement]]. - **Teachers and relational practice:** Research on [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|AI literacy frameworks]] and [[mind-the-trust-gap-teacher-student-views-control-agency-k12-classroom-ai|teacher-student trust]] emphasizes that SEL supports the relational dimensions of learning (teacher-student and student-student relationships), which AI must complement rather than replace. - **Well-being and AI's affective impact:** examine how the increasing use of generative AI affects students' socio-emotional skills, well-being, sociability, and sense of trust and empathy — concerns that motivated the OECD's call for AI literacy grounded in humanistic, social, and emotional values. - **Affective dimensions of AI:** SEL connects to [[affective-computing]] and [[well-being]] research, examining how AI systems can support or undermine emotional and relational learning. - **Policy deficit in AI × SEL:** A [[meta-analysis-systematic-review|systematic review]] of 65 papers at the AI–SEL intersection ([[policy-deficit-ai-sel-2026|Tran, Liu & Nguyen 2026]]) finds a substantial "policy deficit": nearly three-quarters of studies state no policy implications, and the few that do often lack actor-specific guidance. The review links policy engagement to publication venue and warns of a "techno-solutionist" trap in which technical potential is foregrounded while the institutional conditions for responsible implementation remain under-specified. It proposes a "WH-question" framework (Who, What, Why, When/Where, How) to move from "implication-as-afterthought" to "implication-as-methodology," connecting [[ai-education|AI-for-SEL]] innovation to [[educational-policy-ai|educational policy]] and [[governance]]. ## Potential subtopics of SEL in the AI-era research The knowledge base's research clusters SEL into several distinct subtopics, each with its own evidence base: ### Self-efficacy and confidence A core SEL competency (self-awareness/self-management) that strongly conditions how students interact with AI. [[student-dependency-on-ai-literacy-self-efficacy-2026|Student dependency on AI]] (478 Israeli HE students) found that while skill-based AI literacy dimensions were *positively* associated with AI dependency, both academic and AI-specific **self-efficacy** and effort [[regulation]] were *negatively* associated — meaning AI literacy alone does not protect against dependency; [[self-efficacy]] does. [[self-efficacy-tutoring-learning|Cen et al. (EC-TEL 2026)]] found that students with lower baseline self-efficacy achieved greater learning gains regardless of practice format, and that favor toward the tutor mattered more for tutor-based practice — underscoring the value of tailoring practice to motivational profiles. Connected concepts: [[self-efficacy]], [[motivation]], [[agency]]. ### Motivation and the "AI availability" effect Motivation intersects with SEL's responsible-decision-making and self-management. Research on [[ai-availability-student-motivation|AI availability and student motivation]] and [[ai-perceptions-students-teachers-motivation-2026|student/teacher motivation and self-efficacy]] shows that the mere availability of AI can reshape students' motivational orientation and perceived effort — relevant to how AI might undermine or support intrinsic motivation. [[student-dependency-on-ai-literacy-self-efficacy-2026|Self-efficacy research]] frames effort regulation as the counterweight to AI dependency. ### Emotion regulation and affective support Emotional regulation is a key SEL competency with direct learning consequences, forming a direct SEL→[[learning-gains|achievement]] link. [[affective-computing]] tools like [[kar-mathbuddy-affective-math-tutoring-2025|MathBuddy]] model student emotions to shape [[pedagogy|pedagogical]] responses, and [[ai-campus-wellbeing-tools|AI campus well-being tools]] (e.g., PsychoGPT, AURA) span prevention and intervention. ### Shame, guilt, and emotional responses to AI use Emotions regulate *how* students make AI use visible. [[shame-guilt-ai-regulation-computing-education|"Stuck in a Spiral"]] (19 computing students) found that shame and guilt act as social regulators of AI use, driving hiding behaviors and selective disclosure and creating cycles of reduced agency. [[ai-anxiety-strategic-regulation-writing-2026|AI anxiety]] can be transformed into strategic regulation of AI as a learning resource. These connect SEL's social awareness and self-management to [[academic-integrity]] and responsible AI use. ### Trust and belonging SEL supports relational learning and social cohesion. [[finkelstein-principled-ai-education-2025|Principled AI education]] and [[ai-chatbot-collective-efficacy-collaborative-learning|AI chatbots for collaborative learning]] connect to **belonging** and collective efficacy — the shared belief in a team's ability to accomplish tasks. The Brookings premortem emphasizes that overreliance on AI threatens social-emotional well-being, teacher-peer relationships, and student [[privacy]]/safety — dimensions of belonging and connectedness. [[mind-the-trust-gap-teacher-student-views-control-agency-k12-classroom-ai|Teacher-student trust]] is central to whether AI is perceived as supportive. ### Persistence, mindset, and productive struggle SEL overlaps with the effortful dimension of learning. [[framing-5-percent-problem-teachers-persistence|Framing the 5% problem]] identifies low student persistence as a recurring challenge in educational technology, shaped by motivation/buy-in, cognitive roadblocks, resilience under challenge, and connection. [[substitution-to-scaffolding-ai-harm-cycle-2026|Favero et al.]] warn that AI that substitutes for effort erodes the very capacities education builds — aligning with [[metacognition]] and the value of productive struggle over [[cognitive-offloading|over-reliance]]. ## Implications for instructors and instructional design - **Treat SEL as a first-class design goal, not an add-on.** Integrate social-emotional competencies into [[ai-literacy]] curricula (per [[sec-ai-literacy-narrative-review-2026|Palmquist et al.]]), so students learn not just *how* to use AI but *when* and *why* — with attention to their own emotions, effort, and relationships. - **Design for self-efficacy and agency.** Because skill-based AI literacy can *increase* dependency while [[self-efficacy]] and effort regulation protect against it, instruction should deliberately build students' confidence and self-management alongside technical skill — e.g., scaffolded practice, low-stakes successes, and prompts that require students to verify and own AI output. - **Use AI to support, not supplant, relational learning.** AI tools should complement [[collaborative-learning|teacher-student and student-student relationships]], and [[teacher-role|educators]] should retain the relational role — monitoring affect and belonging — rather than delegating it entirely. - **[[affective-tutoring|Emotion-aware]] and affect-sensitive design.** [[affective-computing|Affect-aware systems]] can detect and respond to emotional states (anxiety, frustration, confusion), but evidence shows benefits are **not universal** — moderating by learner proficiency and profile — so emotional design must be tailored, not assumed. - **Address the emotional side of AI use.** Design against shame/guilt spirals and AI anxiety by normalizing discussion of AI use, fostering transparency, and reducing surveillance-based responses that erode trust and agency. - **Support persistence and productive struggle.** [[scaffolding|Scaffold]] rather than substitute (per [[substitution-to-scaffolding-ai-harm-cycle-2026|Favero et al.]]), so AI deepens rather than bypasses effortful learning — aligning with [[metacognition]] and [[desirable-difficulties|productive difficulty]]. - **Prepare teachers' SEL competency.** Educators need both technological skill and emotional intelligence ([[teacher-ai-competency]]); teacher [[educational-development|professional development]] should build capacity to support students' SEL in AI-mediated settings. ## Connections to learning gains and other measures - **Self-efficacy moderates gains.** [[self-efficacy-tutoring-learning|Cen et al.]] found lower-baseline-self-efficacy students achieved the *largest* learning gains, and that tutor-favorability predicted gains in tutor-based practice — showing motivational profiles shape who benefits from which format. - **Well-being and engagement as intermediate outcomes.** SEL-related outcomes (motivation, [[well-being]], belonging, engagement, self-efficacy) often function as mediators of downstream achievement, and AI research increasingly measures them alongside — or in some cases instead of — raw test scores. - **Effects are conditional, not universal.** Research cautions that SEL-oriented interventions may help some learners (by profile/proficiency) and not others, so claims about SEL-based learning gains should be examined for moderator effects. - **The harm side of the ledger.** The cautions that AI-driven [[cognitive-offloading|overreliance]] threatens social-emotional well-being, relationships, and belonging — outcomes that, if eroded, can undermine the very foundations of long-term learning and achievement. ### Connections to related concepts SEL connects to [[ai-literacy]] (as a complement that makes AI literacy relational and ethical), [[affective-computing]] and [[well-being]] (the affective dimensions of AI), [[self-regulated-learning]] (self-management and effort regulation), [[self-efficacy]] and [[motivation]] (learner beliefs that moderate [[student-ai-interaction|AI interaction]]), [[agency]] (protecting learner control against dependency), [[ethics]] (responsible decision-making), [[teacher-ai-competency]] (educators' capacity to support SEL), [[student-experience]] (well-being and belonging), and [[learning-gains]] (the evidence that SEL supports achievement). It relates to [[higher-ed]] and [[k-12]] as the settings where SEL-infused AI literacy is cultivated. ## Connected Concepts - [[anxiety-and-stress]] - [[ai-literacy]] - [[affective-computing]] - [[well-being]] - [[self-regulated-learning]] - [[self-efficacy]] - [[motivation]] - [[agency]] - [[collaborative-learning]] - [[ethics]] - [[teacher-ai-competency]] - [[teacher-role]] - [[scaffolding]] - [[educational-development]] - [[student-experience]] - [[learning-gains]] - [[higher-ed]] ## Connected Articles - [[sec-ai-literacy-narrative-review-2026]] — Integrating Social-Emotional Competencies Into AI Literacy - [[mind-the-trust-gap-teacher-student-views-control-agency-k12-classroom-ai]] — Mind the Trust Gap: Teacher-Student Views - [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl]] — The Scaffolded AI literacy (SAIL) framework - [[teacher-education-ai-literacy-sdt-2026]] — Teacher Education for AI Literacy (SDT) - [[student-dependency-on-ai-literacy-self-efficacy-2026]] — Student Dependency on AI, Self-Efficacy, and Resource Management - [[self-efficacy-tutoring-learning]] — Self-Efficacy and Favorability Shape Learning from Tutoring - [[shame-guilt-ai-regulation-computing-education]] — Shame and Guilt as Social Regulators of AI Use - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding - [[ai-chatbot-collective-efficacy-collaborative-learning]] — AI Chatbots and Collective Efficacy - [[kar-mathbuddy-affective-math-tutoring-2025]] — MathBuddy: Affective Math Tutoring - [[ai-campus-wellbeing-tools]] — AI-Driven Campus Well-being Tools - [[policy-deficit-ai-sel-2026]] — The Policy Deficit in AI × SEL Research - [[human-ai-complementarity-social-emotional-learning-2026]] — Human–AI complementarity in early social-emotional learning (Raave et al. 2026) --- ## [Well-Being](https://edtechdev.github.io/aied/concepts/well-being/) > **Well-being** — the positive state of being mentally, physically, and socially healthy, encompassing emotional, psychological, and social dimensions. In AI in education, well-being has become a central concern because the rapid integration of [[generative-ai|generative AI]] into learning environments can affect students' and educators' mental health, motivation, belonging, anxiety, and sense of agency — raising questions about whether AI supports or undermines learners' well-being. ## Questions to Consider - Well-being in education is often treated as a 'soft' or secondary concern next to learning outcomes. But the page frames it as a central, designable dimension of AI integration. Do you think of learner and teacher well-being as something to design for, or as a byproduct you can check later? What might the field lose by treating it as an afterthought? - Think about how AI use has affected your own or your students' confidence, anxiety, sense of belonging, or agency — for better and worse. What's one concrete way AI has changed your emotional experience of learning or teaching, beyond just convenience? - The page notes AI use is intertwined with anxiety — about academic integrity, about whether relying on AI signals weakness, about over-dependence. Where have you seen AI induce anxiety rather than relieve it, and who in that situation was most affected? - One common intuition is that AI reduces stress by handling hard tasks. Where might the opposite be true — AI that lowers short-term effort yet increases anxiety about competence, integrity, or 'not really knowing how to do it'? How would you even measure a shift in well-being rather than just output? - Well-being spans emotional, psychological (purpose, [[agency|autonomy]], competence), and social (belonging, relationships) dimensions. Pick one of those. How could an AI tool quietly improve or undermine it in a learning setting — and what would you look for as evidence? - Teacher well-being — workload, anxiety about disruptive technology, capacity to offer emotional support — is as much at stake as students'. If you're an educator or leader, what would it take for AI integration to support educators' well-being rather than add to the pressure, and who should be accountable for that? ## Introduction Well-being in education is multifaceted: it includes emotional well-being (positive affect, low distress), psychological well-being (purpose, autonomy, competence), and social well-being (belonging, positive relationships). In the AI era, well-being matters because AI can reshape learning in ways that affect these dimensions — from reducing students' confidence and increasing anxiety about [[academic-integrity|academic integrity]], to fostering or undermining [[student-engagement|engagement]] and belonging. Concerns about AI's impact on students' socio-emotional skills, well-being, sociability, and sense of trust and empathy (raised by the OECD and others) have positioned well-being as a key consideration in responsible AI integration. ### How well-being appears in the research - **[[ai-literacy|AI literacy]] and social-emotional learning:** [[sec-ai-literacy-narrative-review-2026|Research integrating SEC into AI literacy]] argues that fostering educators' and students' emotional intelligence and well-being is essential for navigating AI-[[sociocultural-learning|mediated learning]] environments, connecting to [[social-emotional-learning]] and affective dimensions of AI. - **AI anxiety and student experience:** Studies on students' engagement with AI (e.g., [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination|SDT-based research]]) find AI use is intertwined with anxiety, trust, and confidence, with students' well-being affected by concerns about academic integrity, [[creativity]], and [[cognitive-offloading|Over-Reliance]]. - **AI anxiety, adaptation, and dependence as well-being signals:** [[zhang-ai-anxiety-academic-motivation-emotion-2026|Zhang et al. (2026)]] find AI anxiety is negatively tied to academic motivation partly through reduced [[metacognition|emotion regulation]] (moderated by gender) in a large Chinese sample; [[wu-psychological-adaptation-ai-japanese-learning-2026|Wu (2026)]] shows learners of [[language-learning|Japanese]] sort into maladaptive, moderate, and positive psychological-adaptation profiles driven by technostress and resilience that shift toward better adaptation over a semester; and [[yan-conversational-ai-engagement-dependence-synthesis-2026|Yan (2026)]] cautions that cross-sectional correlates of [[conversational-ai]] engagement (loneliness, anxiety, low well-being) should not be read as consequences, and that supportive and harmful experiences coexist. - **Culturally [[situated-learning|situated]] well-being support and its limits:** [[culturally-aware-student-stress-chatbot-2026|Bashir and Afzal (2026)]] describe Sukoon, a hybrid well-being system for Pakistani university students that pairs a Random Forest stress classifier (89.09% accuracy over three severity levels; 20 survey features) with an [[llm]] dialogue layer that escalates tone and support intensity across three tiers in line with the Stepped Care Model. It was built because Western-designed mental-health tools are English-language and culturally mismatched for students who express distress in Urdu or Roman Urdu and who face academic, financial, familial and relational stressors simultaneously; the authors are explicit that it is not a [[medical-education|clinical]] diagnostic or therapy tool, that high-distress responses point toward professional counselling, and that the chatbot layer has not yet been evaluated with students on cultural appropriateness or emotional safety. - **Teacher well-being and role:** AI's impact on [[teacher-role|teachers]] — including workload, anxiety about teaching with disruptive technology, and the capacity to provide emotional support — is a recurring concern, connecting to [[teacher-ai-competency]] and [[educational-development|professional development]]. - **Relational densification as the evaluative criterion for AI-supported teacher development:** [[ai-emotional-intelligence-teacher-development-2026|Aponte et al. (2026)]] argue that AI supports teachers' socio-emotional development only when it functions as relational infrastructure rather than a symbolic substitute for human accompaniment. They propose **relational densification** as the criterion for judging AI-supported professional-development initiatives — whether they strengthen trust, mentoring, peer support, collaboration, psychological safety, and reduced isolation — and note these relationships have downstream effects on students through classroom climate, [[pedagogy|pedagogical]] responsiveness, and socio-emotional support. The synthesis also cautions that affective data [[governance]] and the political economy of educational AI carry distinct ethical risks, framing [[ai-literacy|critical AI literacy]] as a socio-emotional competence. - **Ethics and responsible AI:** Well-being is a core ethical consideration in [[ai-education]], linking to [[ethics]] and the imperative to design AI that supports rather than harms learners' mental health and belonging. ### Well-being as a design consideration A recurring theme is that well-being should be a deliberate design consideration in AI in education, not an afterthought. This means: designing AI to support rather than replace human relationships; ensuring students can maintain agency and confidence rather than experiencing [[anxiety-and-stress|AI-induced anxiety]] or over-reliance; supporting educators' capacity and well-being as they integrate AI; and [[ai-ed-evaluation|evaluating AI]] systems not only for [[learning-gains|learning outcomes]] but also for their effects on students' and teachers' well-being. [[research-methods-aied|Research]] connects well-being to [[motivation]], [[self-regulated-learning]], and [[student-experience]] (belonging and engagement). [[daoism-ai-education-philosophy-2026|Xie (2026)]] argues for an educational telos to match: the Daoist "Zhenren" (真人) counter-ideal replaces frictionless optimization with "cultivated wholeness," reimagining learning as the harmonious integration of self, society and cosmos, and insisting that "no student is merely a dataset to be managed, but a whole being capable of achieving equanimity." ### Connections to related concepts Well-being connects to [[student-experience]] (as a dimension of learners' overall experience), [[social-emotional-learning]] and [[affective-computing]] (the emotional competencies AI intersects with), [[ethics]] (as a core ethical consideration), [[motivation]] and [[self-regulated-learning]] (well-being supports and is supported by these), [[teacher-ai-competency]] (educators' capacity and well-being), and [[higher-ed]] and [[k-12]] as the settings where AI shapes well-being. **Friction, meaning and loneliness as a signal.** [[zohar-bloom-inzlicht-against-frictionless-ai-2026|Zohar, Bloom and Inzlicht (2026)]] give the well-being case a mechanism: effort signals that our actions matter, so people who work toward a task feel more competent, value the product more and see it as more purposeful — and even on objectively meaningless tasks, adding friction raises appraised meaning. On relationships, they treat loneliness not only as an affliction (linked to cardiovascular disease, dementia, stroke and premature death) but as a **[[biology-education|biological]] signal** akin to hunger or pain: discomfort that motivates reaching out, accepting invitations, investing in existing relationships and tolerating difficult conversations. AI companions can soothe that discomfort, which the authors regard as genuine progress in some cases, while also silencing the signal that drives connection — and they temper the argument by stage, holding that for people isolated by circumstance rather than choice, denying access to such technology "would be cruel" ([[motivation]], [[social-emotional-learning]]). ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[anxiety-and-stress]] - [[student-experience]] - [[social-emotional-learning]] - [[affective-computing]] - [[ethics]] - [[motivation]] - [[self-regulated-learning]] - [[teacher-ai-competency]] - [[higher-ed]] ## Connected Articles - [[zohar-bloom-inzlicht-against-frictionless-ai-2026]] — Friction, meaning, and loneliness as a biological signal - [[zhang-ai-anxiety-academic-motivation-emotion-2026]] — The Relationship Between AI Anxiety and Academic Motivation - [[wu-psychological-adaptation-ai-japanese-learning-2026]] — Profiles and Transitions of Psychological Adaptation in AI-Assisted Japanese Language Learning - [[yan-conversational-ai-engagement-dependence-synthesis-2026]] — A Critical Narrative Synthesis of Conversational AI Engagement and Dependence - [[mindful-llm-math-tutoring-2026]] — Beyond Problem Solving: Large Language Models for Emotional and Reflective Support in Mathematics Learning - [[sec-ai-literacy-narrative-review-2026]] — Integrating Social-Emotional Competencies Into AI Literacy - [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination]] — Students' Engagement With GenAI (SDT) - [[teacher-education-ai-literacy-sdt-2026]] — Teacher Education for AI Literacy (SDT) - [[genai-motivation-engagement-2026]] — Generative AI, Motivation, and Engagement - [[ai-chatbot-collective-efficacy-collaborative-learning]] — AI Chatbots, Collective Efficacy, and Collaboration - [[sovereign-hive-titl-further-education-2026]] - [[aivaluate-anxiety-assessment-2026]] — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026) - [[policy-deficit-ai-sel-2026]] — The Policy Deficit in AI × SEL Research - [[emancipatory-ai-learner-flourishing-2026]] — Emancipatory vision oriented toward learner flourishing - [[culturally-aware-student-stress-chatbot-2026]] — An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning - [[daoism-ai-education-philosophy-2026]] — Alternative AI Philosophy: Daoism as Method for AI in Education - [[ai-emotional-intelligence-teacher-development-2026]] — Relational densification as the criterion for AI-supported teacher development --- ## [Creativity](https://edtechdev.github.io/aied/concepts/creativity/) > **Creativity** — the capacity to generate novel and valuable ideas, solutions, or artifacts. In the AI era, creativity is a central educational stake: [[generative-ai|generative AI]] can both amplify creative work (as a divergent-thinking partner) and undermine it (by homogenizing output and replacing the generative process). ## Questions to Consider - Creativity is often described as divergent thinking — generating many possibilities — versus convergent thinking that narrows to one right answer. Where does your own work or study sit on that spectrum, and which does AI most readily help with? - Generative AI is a statistical engine: it can propose many options, but it tends toward the average. If many students rely on the same model, what happens to the diversity of ideas across the class? - A 'think first, ChatGPT later' study found students who generated their own ideas before AI showed no immediate boost — yet outperformed everyone on a later unassisted creativity task. Why might protecting the independent-thinking phase produce learning that free AI use doesn't? - If AI can produce a polished artifact instantly, what is the learner's creative work actually worth — and how would you design an assignment so the generative process stays with the student? - Is an idea generated with AI's help 'yours'? How you answer might change whether you treat AI as a divergent-thinking partner or as something that replaces your creative process. ## Introduction Creativity spans the divergent-thinking end of the cognitive spectrum — generating multiple possibilities — in contrast to convergent thinking, which arrives at a single correct solution. AI systems are especially relevant to creativity because they are statistical generators: they can propose many options (supporting ideation) but also tend toward the average, producing the idea-level homogenization documented when many students rely on the same model. ### Creativity and generative AI - **Amplification:** AI can act as a divergent-thinking partner — brainstorming alternatives, generating counterarguments, and offering perspectives the learner might not consider. Role-specialized [[agentic-ai|multi-agent]] configurations can restore ideational diversity. - **Homogenization risk:** single-model assistance can reduce the diversity of ideas across students, so the same model produces convergent outputs. This is a direct threat to creativity in [[writing-education]] and [[design-education|design education]]. - **Protecting creative agency:** keeping the learner's generative process in the loop — draft-first routines, requiring original synthesis, and using AI to challenge rather than replace — preserves the creative work that produces durable learning. - **A K–12 field map: GenAI mostly *enhances* creativity and rarely *measures* it.** A PRISMA-guided [[meta-analysis-systematic-review|scoping review]] of 45 studies (2017–2025) sorts the K–12 literature into four uses — creativity [[assessment]] (4 studies), human–AI co-creativity (3), [[stakeholders]]' perceptions (14) and creativity enhancement (34) — and shows how narrow the field still is: language-based storytelling and writing crowd out nearly everything else, with music, [[embodied-learning|embodied]] and spatial work and [[equity-in-ai-education|equity]]-focused studies almost absent. Purpose-built [[scaffolding|scaffolds]] that retained authorship (voice input, storyboards, divergent "many options" prompts) outperformed off-the-shelf platforms, and [[llm]] models scored divergent thinking close to human raters; yet the authors' own concern is the same homogenization this page tracks — shrinking linguistic diversity, style convergence, AI-contaminated training data and [[metacognition|metacognitive]] laziness in co-writing — which they gloss with Runco's term "artificial creativity": output that looks creative without the human experience behind it. Over-reliance reappeared as a risk across perception, co-creativity and [[curriculum-design|curriculum]] studies alike ([[genai-creativity-k12-scoping-review-2026]]). - **Co-creativity is a network effect, not an individual trait:** [[trikonet-trivalence-co-creativity-2026|Ruhland (2026)]] draws on Actor-Network Theory to model creativity as an emergent property of socio-technical networks — a triadic interplay of stabilization, destabilization, and re-stabilization among human and non-human actors, with [[generative-ai|generative AI]] agents treated as equal, constitutive network participants. In a study where [[pedagogy|pedagogical]] avatars were co-designed in a creative network, the avatar designs were not individual creative acts but emerged through translation processes among students, [[research-methods-aied|researchers]], and design tools. TriKoNet frames the risk that ready-made AI avatars shift creative [[agency]] toward machine-induced convenience (the "Convenience Trap") and argues that co-constituting the AI's action structure with [[learners]] preserves co-creativity. - **Think-first collaboration sustains independent creativity:** Wong and Qiu (2026) found that students who generated their own ideas *before* using ChatGPT (a "think first, ChatGPT later" protocol) showed no immediate boost on the assisted task, yet outperformed both a free-AI group and a human-only group on a later unassisted creativity task. Freely using ChatGPT produced only transient performance that collapsed when assistance was removed — a form of [[cognitive-offloading|Over-Reliance]] rather than learning — whereas collaborative [[human-ai-collaboration|co-creation]] aimed at improving one's *own* ideas yielded durable gains in independent creativity. This gives direct experimental evidence that protecting creative agency is not merely desirable but is what converts AI-assisted work into learning. - **Reaffirming human creativity in the AI era.** The ascent of generative AI challenges educational fields to reaffirm the value of human creativity. A [[project-based-learning|project-based]] digital [[storytelling-in-education|storytelling]] framework for art and design education was designed explicitly to cultivate emotional, cultural, and narrative capacities that AI lacks, with students producing [[multimodal]] narratives from local cultural heritage — positioning creativity as the distinctly human contribution in AI-integrated learning. ### Creativity across domains Creativity is not a single monolithic faculty — it is realized differently in writing, visual art, [[math-education|mathematics]], and computing, and generative AI interacts with each domain's creative process in distinct ways. Understanding those differences matters for designing AI-assisted learning that protects rather than bypasses each domain's core creative act. - **Visual art and design — the competence paradox.** Text-to-image (T2I) tools compress the distance from idea to artifact, but ease does not simply liberate creativity. In a study of art and design students (417 surveyed), creative competence strongly predicted *intention* to use T2I tools — yet also predicted more *restrained, selective actual use*, as students negotiated authorship, originality, and skill preservation against efficiency ([[t2i-competence-paradox-2026]]). Ease and ready availability enable quick output generation rather than sustained [[student-engagement|engagement]], so the tool can quietly become a shortcut that erodes the iterative studio workflow it was meant to accelerate. - **Mathematics — creativity without transfer.** An AI-supported [[inquiry-based-learning|inquiry-based learning]] intervention in [[k-12|secondary]] mathematics (students averaging 12.79 years) significantly raised *creative mathematical performance* and attitudes toward math — but did **not** significantly improve critical [[problem-solving]] ([[mujib-ai-ibl-creative-math-2026]]). Creativity and convergent problem-solving are separable outcomes; AI-assisted inquiry can grow creative production while transferable analytical skills lag, reinforcing the wider [[cognitive-offloading|performance-learning gap]]. - **Creative computing — friction that protects iteration.** Novice creative coders learn by understanding and extending "found examples," which AI can helpfully scaffold or temptingly bypass. Flowcode, an AI-powered creative-computing environment, pairs a code-structure [[visualization|flowchart]] with a learning-oriented chat and deliberately-designed friction — shown across two studies to steer AI use toward understanding and extending code rather than [[vibe-coding]] around it ([[flowcode-ai-creative-coding]]). Productive difficulty here is a feature that preserves the learner's generative loop. - **Literary creation — the unit between theme and text.** [[incipit-axiom-grounded-scaffolding-literary-creation-2026|Incipit]] argues that a work is organized by *stated premises* rather than themes, and builds an intermediate level — 1,455 axiom records spanning 149 works, joined by 472 typed relations — at which a writer revising a configuration can change an organizing commitment, a reader can justify a reconstruction, and a critic can compare two works. Its own audit sharpens the measurement point above: the artifact records a single curation, 1,448 of the 1,455 axioms map to exactly one work, and the proposed validation (five raters on a 35-record sample) has not run, so the framework is a research program rather than evidence about creative learning. ### Putting creativity into practice The recurring design principle across these findings is to **keep the learner's own generative act in the loop** — whether that act is producing an idea, iterating an artifact, or extending a found example — and to use AI as a divergent-thinking partner that challenges and expands rather than replaces it. - **For [[teacher-role|instructors]]:** sequence assignments so students generate their own initial ideas *before* consulting AI (a "think first, ChatGPT later" protocol), then use AI to interrogate, extend, or play devil's advocate against those ideas. Grade process and original synthesis alongside polish, so effort-averse shortcuts (one-click T2I output, vibe-coded solutions) gain nothing. In art and design, make authorship and skill development the assessed object, not just the artifact. - **For developers and designers:** build tools that reveal and reward the iteration loop — show structure (as Flowcode's flowchart does), add [[desirable-difficulties|productive friction]] at the "ship the first output" moment, and offer alternatives and critiques rather than a single polished answer. When the tool makes the next best version effortless, the learner's creative decisions should still be the ones that matter. - **For researchers:** treat creativity as [[discipline-specific-aied|domain-specific]] and outcome-specific — measure whether gains [[transfer-of-learning|transfer]] to convergent or unassisted tasks, not just whether assisted output looks more creative. The measurement gap is the field's sharpest problem: within the 45-study K–12 corpus only the four assessment studies used a formal creativity measure, most studies never defined the construct at all, and the creative *process* — preparation, incubation, illumination, verification — went essentially unmeasured. Hence the review's prescription is measurement-first: score creativity unobtrusively inside open-ended tasks through evidence-centered design and stealth [[assessment]], and treat creative [[self-efficacy]] as part of the outcome rather than an implied side effect ([[genai-creativity-k12-scoping-review-2026]]). ### Connections Creativity connects to [[critical-thinking]] and to [[constructivist]] learning. It is protected by the same [[reducing-ai-misuse]] [[scaffolding|scaffolds]] that preserve learning, and by [[authentic-assessment]] designs that reward original reasoning over polished products. ## Connected Concepts - [[critical-thinking]] - [[constructivist]] - [[writing-education]] - [[student-experience]] - [[generative-ai]] - [[reducing-ai-misuse]] - [[authentic-assessment]] - [[collaborative-learning]] - [[design-thinking]] - [[problem-solving]] - [[inquiry-based-learning]] - [[cognitive-offloading]] - [[scaffolding]] - [[arts-design-and-media-education]] ## Connected Articles - [[incipit-axiom-grounded-scaffolding-literary-creation-2026]] — Axiom-grounded scaffolding for literary creation: premises as the unit between theme and text, validated only in plan (Liu & Zhao 2026) - [[powerful-learning-with-emerging-technology-2025]] — Scaffolding creativity instead of completing it - [[typology-generative-ai-tools-education-2026]] — Typology of Generative AI Tools for Education - [[jin-emergent-learner-agency-implicit-hai-2026]] — Emergent learner agency in implicit human-AI collaboration: supportive vs. contrarian personas - [[rana-genai-design-thinking-2025]] - [[chatgpt-critical-creative-thinking-review]] — ChatGPT and Critical and Creative Thinking: Systematic Review - [[think-first-chatgpt-later-2026]] — Think First, ChatGPT Later: Independent Human Creativity - [[ai-collaborative-learning-skills-impacts]] — AI and Collaborative Learning: Impacts on Creativity - [[enhancing-creative-writing-with-robot-llm-integration-the-interplay-of-embodimen]] — Robot-LLM Integration and Creative Writing - [[multi-agent-llm-social-learning]] — Beyond the AI Tutor: Social Learning with LLM Agents - [[genai-mindtool-generative-learning]] — GenAI as a Mindtool for Generative Learning - [[project-based-digital-storytelling-art-design-2026]] — Project-based digital storytelling framework for art/design education in the AI era - [[t2i-competence-paradox-2026]] — The competence paradox of text-to-image AI among art and design students - [[mujib-ai-ibl-creative-math-2026]] — AI-supported inquiry-based learning and creative mathematical performance - [[flowcode-ai-creative-coding]] — Flowcode: an AI-powered environment scaffolding iteration in creative computing - [[trikonet-trivalence-co-creativity-2026]] — TriKoNet: trivalence model of potential co-creativity in socio-technical networks - [[know-when-to-trust-ai-scoring-reliability-2026]] — Know When to Trust: scoring divergent-thinking responses with LLMs, and where it breaks down - [[genai-creativity-k12-scoping-review-2026]] — PRISMA scoping review of 45 K–12 studies: four GenAI uses for creativity, homogenization, and the measurement gap --- ## [Student-AI Interaction](https://edtechdev.github.io/aied/concepts/student-ai-interaction/) > **Student-AI interaction** — the patterns, processes, and cognitive work in how learners engage with [[generative-ai|generative AI]] systems during learning and [[problem-solving|problem solving]]. [[research-methods-aied|Research]] here characterizes what students ask of AI, how prompts and dialogues evolve, and how interaction quality relates to [[learning-gains|learning outcomes]], [[cognitive-offloading]], and [[agency]]. ## Questions to Consider - Think about the last few prompts you (or a student) wrote to an AI. Would you describe most of them as asking for the answer, or asking the AI to explain, probe, or evaluate? What do you suspect that pattern does to learning? - The research finds that a small subset of question types accounts for most student inquiries, and that the questions change as a task progresses. Why do you think students' questioning narrows, and what does that suggest about how they're using the tool? - The page claims shallow, answer-seeking prompts are associated with reduced learning and over-reliance, while reflective, verification-oriented interaction supports understanding. What do you think separates a 'good' prompt from a 'bad' one — and is that the student's responsibility or the tool's design? - If interaction quality is shaped by task context and scaffolding rather than being a fixed trait of the student, how might a course or tool be redesigned to invite a wider, more productive range of inquiry? - How would you know whether a student's fluent AI dialogue reflects genuine learning or just skilled delegation — and what would you check to find out? ## Introduction Student-AI interaction is the observable surface of learners' [[student-engagement|engagement]] with generative AI — the questions they pose, the prompts they write, the way they negotiate and verify AI output, and how those patterns shift across task stages and over time. It sits at the intersection of [[student-experience]], [[prompt-engineering]], and [[learning-analytics]], and is central to debates about whether AI use in education represents genuine learning or [[cognitive-offloading|over-reliance]]. Where [[human-ai-collaboration]] frames the high-level division of cognitive labor between people and models, student-AI interaction is the concrete, measurable enactment of that relationship — the specific inquiries, prompts, and negotiation moves learners make moment to moment. ### What students ask AI A core strand of research measures the **types and quality of student inquiries**. Studies apply taxonomies of question types — for example the Graesser et al. 18-type taxonomy — to classify student-AI interactions, often using few-shot classifiers to scale the analysis across hundreds or thousands of interactions. Findings indicate that a small subset of question types accounts for the majority of student inquiries, and that the questions students ask **change substantially as a task progresses** (e.g., [[student-ai-inquiry-types-cs2-2026]]). This task-dependence matters: interaction quality is not a fixed trait of the student but is shaped by problem context, [[scaffolding]], and the affordances of the AI tool. At the youngest ages, [[vahedian-children-attitudes-ai-chatbot-2026|Vahedian Movahed & Martin (2025)]] found children (ages 6–14) actively tested a [[conversational-ai|chatbot]]'s credibility by posing known-answer questions (e.g., "how big is a t rex") — an expression of epistemic self-agency — while the modal child asked only 1–3 questions and a standout first-grader asked 21, underscoring how developmental and individual variation shapes the questions learners pose. Complementing these taxonomy studies, [[yan-cognitive-outsourcing-genai-assessments-2026|Yan et al. (2026)]] characterize the *dialogue form* of student inquiries. Among 38 [[higher-ed|undergraduates]] completing unsupervised argumentative essays, 76.32% used a single-turn ask–get-answer–stop pattern — typically pasting the assessment title without specifying their needs and resubmitting identical prompts when dissatisfied — and 78.94% touched [[generative-ai|GenAI]] only at the start (ideas, background) or end (polishing, length) of a task, keeping it separate from reading and independent [[writing-education|writing]]; only 23.68% sustained iterative back-and-forth dialogue with follow-up questions and their own reasoning. The authors place these patterns on a spectrum from **cognitive outsourcing** to **cognitive reallocation** — the GenAI-era analogue of surface versus [[metacognition|deep approaches]] to learning — noting that most students conceived the tool as an upgraded search engine, which constrained them to the outsourcing end. ### Interaction quality and learning A complementary strand links the *form* of interaction to learning. Shallow or habitually narrow prompts (asking AI to produce the answer rather than to explain, probe, or evaluate) are associated with reduced learning and increased over-reliance, whereas reflective, verification-oriented interaction supports [[metacognition]] and durable understanding. This connects student-AI interaction directly to [[intelligent-tutoring]] design: systems can be built to invite a wider, more productive range of inquiry and to scaffold question-asking rather than merely answering. Interaction need not run through answer-seeking prompts at all — when AI critiques students' own work, the exchange becomes a reflective, verification-oriented dialogue: in [[oppenheimer-llms-collaborative-learning-partners-2026|Oppenheimer, Cash & Connell Pensky (2025)]], learners' responses to [[llm]] essay [[feedback]] showed reflection in 92.7%, acceptance in 93.6%, and active rebuttal of LLM claims in 87.8% of cases (inter-rater κs = 0.81–0.89), and their response-to-[[ai-feedback-quality|feedback quality]] improved across iterations as a learnable skill. The reflective side of interaction need not run through direct prompting — in [[breideband-community-builder-cobi-2026|CoBi]], students engaged with classroom-level AI visualizations of their own [[collaborative-learning|collaborative]] speech and deliberated over when the AI's classifications seemed off, turning apparent misclassifications into opportunities for calibrating their understanding of the AI's capabilities and limits ([[trust-calibration]]) rather than purely accepting its output. Bernstein and Sibia (2026) document an iterative-filter pattern in how CS2 students handle GenAI explanations ([[student-reception-genai-analogies-computing-2026]]): they cross-reference against lecture notes, demand provenance ("I would be a lot more doubtful... without one"), and probe with follow-up questions for inconsistency rather than issuing a single accept-or-reject judgment. Students also read explanations for whose knowledge and background they assumed — assumed [[prior-knowledge|prior knowledge]] beyond the syllabus, default sport and gaming references ("the more male-dominated side of computing"), and excessive repetition all functioned as signals about the imagined reader, with over-scaffolding read as condescending rather than merely inefficient. Consulting AI at the right *point* in a task matters as much as the wording of the individual prompt: in the same study, [[yan-cognitive-outsourcing-genai-assessments-2026|Yan et al. (2026)]] found the reallocation-oriented minority alternated independent work with GenAI consultation and reported unchanged total effort but a shifted focus — moving resources from searching to checking argumentative quality and balance, and writing reflection notes after sessions to counter shallow retention — whereas learners with mastery goals but weak [[ai-literacy|AI literacy]] fell into an "efficiency paradox," offloading "not by intention, but by default." ### From interaction to pedagogy Characterizing student-AI interaction informs [[learning-design]]: instructors can notice when students' questioning patterns are narrow or shallow and design interventions that broaden inquiry; [[teacher-role]] shifts toward coaching students to interact productively with AI. It also grounds [[ai-literacy]] curricula that treat effective prompting and verification as learnable skills rather than innate abilities. Non-use is itself an interaction pattern that [[pedagogy]] must plan for. [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al. (2026)]], studying 85 [[teacher-education|student teachers]] in three courses where GenAI use in [[assessment]] was explicitly permitted, found that 62.4% (53 of 85) declined to use it at all, far below the 79–83% adoption seen in comparable UK and Australian surveys, and that adopters' use was shallow and corrective rather than generative (proofreading 43.8%, clarity checks 34.4%, text generation only 18.8%). Their choices tracked assessment design and institutional culture rather than technical difficulty: 41.5% of non-adopters feared being wrongly accused of [[academic-integrity|plagiarism]], and nine of eleven interviewees read the permissive policy itself as a possible "trap." The gap between 32 survey-reported users and 28 self-declarations shows that students' *reported* AI interaction is shaped by graded consequences — a measurement caveat for learning-analytics accounts of student-AI interaction. ## Discipline and Cognitive Engagement in Student-AI Chat - **Discipline-associated cognitive engagement in student-AI chat.** Chang and Li (2026) analyze student prompts to AI across 116 courses with a within-person, cross-discipline design, showing that student-AI conversations reflect **discipline-associated** cognitive engagement rather than fixed individual interaction styles. Roughly 62% of prompts encoded higher-order cognitive demand overall, but Bloom-level profiles differed sharply by discipline: [[stem-education|STEM]] courses elicited Apply-prevalent prompts (20.8%), language courses Understand-prevalent (31.7%), and social science courses Create-prevalent (33.8%). Paired within-person comparisons confirmed the same students produced significantly more higher-order prompts in social science than STEM courses (pooled n = 16, p < .001), and course-level variation exceeded student-level variation — a strong argument that AI teaching assistants should be designed and evaluated with disciplinary context in mind. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[human-ai-collaboration]] - [[student-experience]] - [[prompt-engineering]] - [[learning-analytics]] - [[cognitive-offloading]] - [[intelligent-tutoring]] - [[metacognition]] - [[agency]] - [[generative-ai]] - [[llm]] - [[ai-literacy]] ## Connected Articles - [[yan-cognitive-outsourcing-genai-assessments-2026]] — Cognitive outsourcing vs. reallocation in unsupervised student–GenAI assessments (Yan et al. 2026) - [[zou-is-this-a-trap-student-teachers-genai-2026]] — “Is this a trap?”: student teachers’ non-adoption of GenAI in assessments (Zou et al. 2026) - [[tutortrace-learner-behavioral-states-2026]] - [[enright-staff-perspectives-genai-2026]] - [[student-ai-inquiry-types-cs2-2026]] — Analysis of Types of Inquiries in Student-AI Interaction - [[student-llm-interaction-taxonomy-review-2026]] — Student-LLM Interaction Taxonomy Review - [[teacher-authored-prompts-student-ai-dialogue]] — Teacher-Authored Prompts in Student-AI Dialogue - [[constructing-epistemic-ai-literacy-student-ai-co-programming]] — Constructing Epistemic AI Literacy - [[icap-cognitive-engagement-llm-agents]] — ICAP Cognitive Engagement with LLM Agents - [[dura-llm-cs2]] — Demystify, Use, Reflect, Assess (DURA): LLM Integration in CS2 - [[learnlm-improving-gemini-learning]] — LearnLM: scenario-guided learner-AI tutoring conversations - [[li-dbagent-llm-educational-agent-cs-2026]] — LLM-based educational agent (DBagent) in CS education - [[strydom-human-gai-paradigms-2026]] — Framing human-AI dynamics: seven GAI engagement paradigms (Strydom 2026) - [[chatgpt-qiskit-homework-autogradable-2026]] — ChatGPT solves Qiskit homework; autogradable design - [[llm-adaptive-programming-error-explanations-2026]] — LLM adaptive explanations of programming errors - [[isaza-chatgpt-engineering-prompting-2026]] — Logged prompting and integration behaviors - [[student-ai-conversations-cognitive-engagement-2026]] — Discipline-associated Bloom-level cognitive engagement in student-AI conversations (Chang & Li 2026) - [[breideband-community-builder-cobi-2026]] - [[oppenheimer-llms-collaborative-learning-partners-2026]] - [[vahedian-children-attitudes-ai-chatbot-2026]] - [[student-reception-genai-analogies-computing-2026]] — Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education - [[naim-bypass-offload-scaffold-llm-learning-2026]] — Bypass, Offload, or Scaffold: A Conceptual Model of How Large Language Models Shape Learning - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback - [[learning-analytics-genai-secondary-writing-2026]] — Using Learning Analytics to Support Secondary School Students' Writing with Generative AI - [[ai-tutor-modality-randomized-field-experiment-2026]] — When AI Tutors Speak: Evidence from a Randomized Field Experiment - [[context-prompts-physics-assignments-2026]] — Artificial Intelligence Driven Physics Assignments using Context Prompts --- ## [Problem Solving](https://edtechdev.github.io/aied/concepts/problem-solving/) > **Problem solving** — the process of formulating, analyzing, and resolving novel or complex challenges — is a core 21st-century competency that AI tools both amplify and threaten. Across the knowledge base's articles, generative AI functions as a scaffold, a dialogic partner, and an answer engine, and its educational value hinges on whether learners remain the primary decision-makers or outsource their reasoning to the machine. This tension between efficiency and deeper cognitive and regulatory [[student-engagement|engagement]] defines current [[research-methods-aied|research]] on AI and problem solving. ## Questions to Consider - Think of a genuinely hard problem you solved well. What made the struggle productive, and what would have been lost if someone had simply handed you the answer? - The page describes a 'cognitive debt': delegating reasoning to AI produces the right product at the cost of understanding. Have you ever got the right answer without really understanding it? What was the cost later? - AI tools help students explore multiple perspectives and generate their own problems—but can also encourage 'metacognitive laziness.' How do you tell the difference between AI as a thinking partner and AI as a substitute for thinking? - One experiment found reflective and hybrid feedback outperformed direct AI feedback on delayed, AI-free transfer. Why might feedback that makes you do more work lead to learning that lasts longer? - How does [[teacher-role|teaching]] students to *pose* their own problems (rather than only solve given ones) support transfer and self-study? When has generating a question taught you more than answering one? - If AI errors are treated as provocations to question and verify, a limitation becomes a learning opportunity. Have you ever learned more from an AI's mistake than from its right answer? ## Introduction Generative AI offers clear support for problem solving. [[generative-ai|LLM]] tools help students explore alternative solutions, incorporate interdisciplinary perspectives, and simulate authentic real-world scenarios — such as [[medical-education|clinical]] and [[ethics|ethical]] dilemmas or prototype testing in [[stem-education|STEM]] — while delivering scalable, timely feedback and shifting assessment toward scenario-based, competency-based evaluation. In [[ai-assisted-collaborative-learning-model-dbr|design-based research]], AI positioned as an "intelligent [[pedagogical-agent|learning partner]]" within structured collaborative tasks produced substantial [[learning-gains|gains]] in problem-solving performance, with students increasingly using AI to generate multiple perspectives rather than seek single answers. Yet the same tools can erode the very skill they claim to support. Research on [[hao-human-ai-collaborative-problem-solving-cognition|human-AI collaborative problem solving]] shows a striking trade-off: the interaction profile that achieves the highest task performance also shows the lowest self-[[regulation]], as students delegate reasoning to the AI and reap what the authors call a "cognitive debt" — obtaining the right product at the cost of the self-constructive process of understanding. Studies of [[ai-collaborative-learning-skills-impacts|collaborative learning]] similarly warn of "metacognitive laziness," where outsourcing cognitive effort undermines independent analysis and self-regulated learning. ### Problem solving in AI education research The knowledge base's articles approach problem solving through distinct but converging lenses. [[llm-critical-thinking-teamwork-review|A PRISMA systematic review]] finds that LLMs often produce incomplete or incorrect responses, prompting students to question, verify, and improve information — turning model imperfections into validation-and-correction cycles that strengthen [[critical-thinking|critical thinking]] and mental independence. Rather than treating problems as given, other work foregrounds *problem posing*: [[genai-assisted-problem-posing-physics-2026|training students to generate their own physics problems]] supports [[transfer-of-learning|transfer]] and self-study, and [[dai-chatbots-problem-posing-primary-2026|GenAI chatbots]] improved primary students' problem-posing quality and reduced cognitive load in [[inquiry-based-learning|inquiry-based learning]]. Scaffolding and assessment are also central. [[adaptive-ai-scaffold-collaborative-problem-solving-2026|Adaptive AI scaffolds]] derived from sequence-mining individual students' process patterns boost on-task performance in collaborative [[math-education|mathematics]] problem solving — though maximal scaffolding also increased scripting behavior. On the measurement side, [[llm-computational-thinking-physics-2026|LLMs can mirror human raters]] in detecting growth in [[computational-thinking|computational thinking]] across large [[physics-education|physics]] courses, while [[computational-thinking-educational-robotics-secondary-2026|computational thinking]] is proposed as the explicit conceptual glue that makes [[educational-robotics|educational robotics]] foster genuine problem solving rather than isolated technical exercises. ### How AI both helps and hinders The dual character of AI for problem solving is consistent across studies. On the helping side, [[conversational-ai|chatbots]] outperform search engines for problem posing by improving question quality and integrating cognitive networks; adaptive scaffolds improve performance; and collaborative AI designs raise both [[critical-thinking|critical thinking]] and problem-solving scores while students remain primary decision-makers. On the hindering side, over-delegation reduces regulatory engagement, direct feedback without reflection encourages passive uptake, and heavy reliance on AI to build consensus threatens interpersonal skill development and [[motivation|intrinsic motivation]]. Notably, [[genai-feedback-design-multisite-experiment|a multisite randomized experiment]] found that reflective and hybrid feedback designs outperformed direct [[ai-feedback-quality|AI feedback]] on delayed, AI-free transfer — evidence that AI's educational value depends less on access than on preserving [[agency|student agency]], evaluative judgment, and ownership during revision. Design principles emerging from the literature include prioritizing dialogic tension over seamless efficiency, enabling proactive co-regulation, and keeping metacognitive scaffolding and "mind-in-the-loop" oversight. ### Implications For educators, the balance of evidence points to structured, process-oriented integration: [[prompt-engineering|prompt engineering]] training shifts students from unfocused to productive [[student-ai-interaction|AI interaction]], embedding AI within collaborative inquiry tasks yields durable [[learning-gains|gains]], and deliberately using AI errors as provocations converts a limitation into a learning opportunity. For designers, the consistent finding that efficiency does not equal deep learning argues for AI that questions, challenges, and scaffolds rather than answers. For institutions, responsible integration requires pairing [[generative-ai|GenAI]] with explicit [[ai-literacy|AI literacy]] training and assessment rubrics that reward reasoning and justification over correct products. The recurring theme across [[cognitive-psychology]]-grounded work is that problem solving is learned through effortful engagement — and AI's role should be to preserve that effort, not erase it. ### Connections to other concepts Problem solving is the applied outcome of [[critical-thinking|critical thinking]] and [[computational-thinking|computational thinking]], and is cultivated through [[problem-based-learning|problem-based learning]] and [[inquiry-based-learning|inquiry-based learning]] frameworks. It depends on [[scaffolding]] that maintains cognitive demand, on [[self-regulated-learning|self-regulation]] and [[metacognition|metacognitive]] oversight to avoid [[cognitive-offloading|cognitive offloading]], and on [[collaborative-learning|collaborative]] structures in which human and AI work together. [[cognitive-psychology|Cognitive psychology]] and the [[transfer-of-learning|transfer of learning]] literature provide the theoretical grounding for why learner-generated problems and reflective feedback produce more durable [[learning-gains|problem-solving gains]]. The balance of worked examples versus practice for acquiring such skills is itself content-dependent: [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit, Koedinger & Carvalho (2025)]] show generalizable problem-solving skills are best learned through example-integrated practice (alternating worked examples with problems), because worked examples supply the inductive information needed to acquire and generalize a correct rule, whereas pure practice risks strengthening spuriously correlated features rather than the general procedure. ## Simulating Collaborative Problem Solving with LLM Agents - **Simulating collaborative problem solving with participant-specific [[llm]] agents.** Fang (2026) trains individual LLM agents on real participants' dialogue to reproduce collaborative problem solving, and validates the [[simulation|simulations]] with [[network-analysis|Epistemic Network Analysis]], showing simulated and real dialogues are statistically indistinguishable (ENA distance 0.17; permutation p = 0.65). This offers a scalable way to study and generate authentic collaborative problem-solving discourse — with implications for both research and the design of practice environments for this 21st-century skill. ### Automated CPS Skill Coding - Measuring [[collaborative-learning|collaborative problem solving (CPS)]] competence typically requires coding behavior from simulated-task process data into specific CPS skills. [[prompt-engineering|Context-aware prompting]] of pre-trained language models can automate this coding, modeling contextual dependencies and fusing cognitive and social abilities to achieve superior performance over strong baselines on CPS task datasets — addressing the labor-intensity of manual coding and enabling large-scale, real-time assessment. ## Connected Concepts - [[critical-thinking]] - [[computational-thinking]] - [[collaborative-learning]] - [[problem-based-learning]] - [[inquiry-based-learning]] - [[creativity]] - [[scaffolding]] - [[self-regulated-learning]] - [[cognitive-psychology]] ## Connected Articles - [[hao-human-ai-collaborative-problem-solving-cognition]] — Interaction profiles (delegated reasoning vs concerted interpretation) and the efficiency–regulation trade-off - [[ai-assisted-collaborative-learning-model-dbr]] — DBR model positioning generative AI as an intelligent learning partner - [[llm-critical-thinking-teamwork-review]] — PRISMA review of LLMs fostering critical thinking, teamwork, and problem solving - [[adaptive-ai-scaffold-collaborative-problem-solving-2026]] — Sequence-mined adaptive scaffolds for collaborative problem solving - [[genai-assisted-problem-posing-physics-2026]] — Problem posing as a self-regulated learning strategy with GenAI - [[genai-feedback-design-multisite-experiment]] — Reflective/hybrid feedback outperforms direct AI on delayed transfer - [[dai-chatbots-problem-posing-primary-2026]] — Chatbots improve primary students' problem posing in inquiry-based learning - [[llm-computational-thinking-physics-2026]] — LLMs as scalable assessors of computational problem solving in physics - [[llm-agents-collaborative-problem-solving-simulation-2026]] — Fine-tuned participant-specific LLM agents reproducing collaborative problem solving dialogues (Fang 2026) - [[rule-integrated-llm-tutoring-primary-math-2026]] — Rule-guided vs ad-hoc scaffolding in an LLM tutoring system for primary mathematics (Looi et al. 2026) - [[context-aware-prompting-cps-skill-identification-2026]] — Context-aware prompting for automated collaborative problem-solving skill coding - [[rachatasumrit-example-problem-ratio-2026]] - [[geovad-bench-visual-chain-of-thought-geometry-2026]] — Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving --- ## [Mastery Learning](https://edtechdev.github.io/aied/concepts/mastery-learning/) > **Mastery learning** — a [[pedagogy|pedagogical]] framework, formalized by Benjamin Bloom, in which [[learners]] advance only after demonstrating a defined threshold of competence on each unit, rather than moving on a fixed class schedule. It rests on the premise that most students can reach mastery given sufficient time, feedback, and instruction tailored to their current state. AI tutoring and adaptive systems are increasingly operationalizing this model by continuously modeling learner knowledge, selecting tasks, and sustaining practice until competence is demonstrated. ## Questions to Consider - Most schooling fixes time and lets achievement vary — everyone moves on after a set number of weeks. Mastery learning inverts this: achievement is held constant while time, feedback, and practice vary. Which model better matches how you've actually learned something difficult? - A critical caveat in the page: 'correctness is not mastery.' A learner can produce right answers while missing a key constraint, fooling the system into declaring mastery early. Can you think of a skill where being able to perform it correctly still didn't mean you truly understood when *not* to do it? - Mastery-based AI grants learners agency to choose their own tasks, but simulations show naive self-selection can produce massive overpractice. Where's the right balance between letting a learner choose and imposing constraints that keep progression efficient? - The page stresses durable retention, not just a single correct performance — that's why mastery should be followed by spaced practice. How might a learner appear to 'master' something today only to lose it within hours? - If an AI declares you have 'mastered' a topic, what would you want it to check before you believe it — beyond getting a few answers right? ## Introduction Mastery learning holds that achievement should be held constant while time and support vary: learners work through small, well-sequenced units and receive corrective feedback until they meet a mastery criterion, instead of being moved on regardless of what they have learned. Bloom's reframing makes frequent [[formative-assessment]] and an explicit definition of competence central to instruction, and it is exactly that combination — diagnosis, feedback, adaptive pacing — that [[adaptive-learning]] and [[intelligent-tutoring]] automate. AI therefore widens the feasibility of mastery approaches and sharpens the question of whether generated feedback is calibrated well enough to certify it. ## Origins and Core Idea Bloom's mastery learning reframed the goal of instruction from "sorting students by aptitude" to "ensuring competence before progression." Where conventional instruction treats time as fixed and achievement as variable, mastery learning inverts this: achievement is held constant and time, feedback, and practice are allowed to vary. Learners work through small, well-sequenced units and, crucially, receive corrective feedback when they fall short of the mastery criterion rather than being moved along regardless. This places [[formative-assessment]] at the heart of the model — frequent, low-stakes checks that diagnose whether a learner is ready to advance — and it presupposes a clear notion of [[assessment]] tied to observable performance rather than seat time. ## How AI Operationalizes Mastery The bottleneck for classical mastery learning was the teacher-side cost of diagnosing each learner's state and personalizing subsequent instruction. Modern AI systems attack this through [[student-modeling]] and [[knowledge-tracing]]: instead of a single aggregate score, the system maintains a dynamic representation of which knowledge components a learner has (or has not) mastered. The Responsible-DKT work on [[neural-symbolic-knowledge-tracing]] injects explicit mastery rules into a deep learner model — repeated correct responses raise predicted mastery, while repeated incorrect responses act as a stronger signal of non-mastery — producing interpretable and temporally reliable state estimates that [[intelligent-tutoring]] can act on. With a running model of mastery, the system's job becomes deciding *what to present next*. [[simulation|Simulations]] of learners' task-selection strategies show that naive autonomy (e.g., self-selected tasks, risk-averse weakness targeting) can produce substantial overpractice on complex multi-step problems, whereas targeted system constraints can correct maladaptive strategies with little penalty to efficient learners. This is precisely the trade-off that [[adaptive-learning]] and [[personalized-learning]] systems must balance: granting [[agency|learner agency]] where it helps while imposing constraints that keep progression toward mastery efficient. Such decisions also interact with learners' own capacity to regulate their effort, tying mastery learning to [[self-regulated-learning]]. **A critical caveat to mastery inference: correctness is not mastery.** [[deceptive-overgeneralization-adaptive-learning-2026|An, McLaren, and Stamper (2026)]] show that learners who overgeneralize a skill — producing correct actions while omitting a critical application constraint — can appear mastered, leading [[knowledge-tracing]]-based mastery stopping rules to end practice before they encounter a case where the action should be *withheld*. The remedy is to assess *when to withhold* the action, not just how to execute it: include "do-not-act" detector tasks before the mastery threshold triggers, paired with [[feedback]] that names the missing constraint. Mastery is better understood as discrimination of application constraints plus action execution, not correctness alone. **A second caveat concerns the evidence rule behind the threshold.** [[crediting-assisted-work-inflates-mastery-2026|Srivastava (2026)]] ran four update rules over identical event sequences from the ASSISTments 2012–13 mathematics logs — a confirmatory half of 12,716 students and 985,813 scored events — and found the declared mastery count moved with the rule rather than with the learners: crediting any completion put 93.9% of 113,428 student–skill pairs past the 0.95 posterior, against 72.8% when hinted or retried rows were read as failed first attempts. The pairs the lenient rule declared ahead of the strict rule went on to 70.9% unaided accuracy against 85.7% where the rules agreed, below the 0.744 base rate. A progression gate that counts assisted completions therefore certifies learners whose later independent work sits below average, which makes the treatment of [[help-seeking|help]] inside the update rule — not the numeric threshold itself — the decision that fixes what a mastery badge certifies. ## Practice, Retention, and the Limits of AI Support Mastery also depends on durable retention, not merely a single correct performance. Cognitive science on [[retrieval-spacing-interleaving|retrieval practice]] and the forgetting curve motivates spacing practice after the mastery threshold is reached. AI spaced-repetition systems such as Memdora generate practice materials at the point of reading and offer a taxonomy of cognitively grounded retrieval interactions, scheduled by state-of-the-art algorithms, so that achieved mastery is reinforced over time rather than lost within hours. These designs draw on [[cognitive-psychology]] and the principle of [[desirable-difficulties]] to make the effort of retrieval itself part of the learning process. Finally, the evidence warns against assuming AI-generated support is uniformly beneficial. In a multi-[[governance|institutional]] study of AI-generated animated traces for novice programmers, benefits were context-dependent and short-term, and mid-[[student-engagement|engagement]] learners experienced a performance decrement attributed to coordination costs — an expertise-reversal-style effect that underscores the need to personalize support to the learner's current state rather than blanket-apply a tool. Likewise, a developmental continuum of [[ai-literacy|AI literacy]] in [[higher-ed|higher education]] positions mastery as not merely adopting AI tools fluently but progressing through stages of informed and critical use, each with its own [[formative-assessment]] strategies. Together these findings frame AI-enabled mastery learning as a system that must be calibrated to individual learners, sustainably spaced, and assessed for genuine competence rather than fluent output. Standards-based grading is the assessment counterpart to mastery learning, and [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] identify it among five innovative practices aligned with "assessment for learning" that higher-education educators could adopt — but they find it virtually absent from management-education discourse. They attribute this to normative barriers: norm-referenced "grading on a curve," external signaling (rankings, internships, accreditation), and students' instrumental mindset all resist mastery-oriented, gradeless approaches. Their recommendation is incremental experimentation — for example, introducing standards-based rubrics for a single task before scaling — supported by program-level coordination and documented Scholarship of [[teacher-role|Teaching]] and Learning evidence. ## Connected Concepts - [[adaptive-learning]] - [[personalized-learning]] - [[intelligent-tutoring]] - [[knowledge-tracing]] - [[student-modeling]] - [[self-regulated-learning]] - [[formative-assessment]] - [[desirable-difficulties]] - [[retrieval-spacing-interleaving]] — retrieval and spacing as the practice engine inside mastery cycles ## Connected Articles - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[neural-symbolic-knowledge-tracing]] — Injecting mastery/non-mastery rules into deep learning for responsible, interpretable learner modeling - [[simulating-learner-task-selection]] — Simulating how learner task-selection strategies and system constraints shape mastery-learning efficiency - [[memdora-ai-spaced-repetition]] — Cognitively grounded, AI-powered spaced repetition for sustaining retention after mastery - [[ai-generated-traces-novice-programmers]] — Context-dependent, learner-moderated effects of AI-generated learning media on performance - [[ai-literacy-continuum-higher-education]] — A five-stage developmental continuum for moving students from uncritical tool use to critical AI competence - [[mesny-innovative-assessment-grading-management-2026]] - [[crediting-assisted-work-inflates-mastery-2026]] — Crediting assisted work inflates mastery: which evidence rule decides who is declared mastered (Srivastava 2026) --- ## [Anxiety and Stress](https://edtechdev.github.io/aied/concepts/anxiety-and-stress/) > **Anxiety and stress** — the negative emotional states that AI integration can induce in learners and educators (fear of being [[legal-issues-and-risks|falsely accused]], surveillance stress, worries about competence or [[academic-integrity|integrity]]), alongside the positive uses of AI to detect, monitor, and alleviate stress and anxiety. This concept sits within the broader [[well-being]] family and overlaps with [[social-emotional-learning]] and [[affective-computing]], but names the specific emotion-construct — and its productive as well as harmful sides — that AI-in-education [[research-methods-aied|research]] now studies directly. ## Questions to Consider - AI anxiety isn't one thing — it spans proctoring stress, learner-emotion anxiety, career fears, and AI used to relieve stress. Which of these have you felt or witnessed, and did it affect your learning or teaching? - A common assumption is that anxiety is always bad. But research finds AI anxiety can be productive — anxious learners verify and revise more carefully. Can you recall a time your own anxiety made you work more carefully? - Continuous surveillance and the fear of being falsely flagged raise test anxiety and can impair performance — and the stress falls hardest on tool-novice and already-vulnerable students. Is that an equity problem, or just a comfort issue? - Career anxiety — the fear that AI will displace or devalue your professional future — is forward-looking and identity-level. To what extent is that fear driving how you, or students you know, engage with AI? - Studies link higher career adaptability to lower AI anxiety and find that self-efficacy offers only limited buffering. If generic confidence isn't enough, what kind of support would actually reduce career-related AI anxiety? - Faculty anxiety about GenAI may mirror earlier moral panics over calculators and search engines. When is worry about a new technology a legitimate concern, and when is it a recurring pattern of resistance to change? ## Introduction AI anxiety is not a single thing. It spans at least four distinct directions, each with its own evidence base: (1) **stress induced by AI proctoring and integrity surveillance**, (2) **AI anxiety as a learner emotion** that can be either a barrier or a productive signal, (3) **career-related AI anxiety** — the fear that AI will displace or devalue one's professional future, and (4) **AI as a tool for detecting and relieving stress and anxiety**. Recognizing all of these — including the productive-anxiety finding that challenges the purely negative framing — is what distinguishes this concept from the broad [[well-being]] umbrella. ## Remote proctoring, false accusation, and surveillance stress The clearest and most-studied source of AI-induced stress is [[remote-proctoring]]. Continuous surveillance, the fear of being falsely flagged, and the pressure of being watched raise test anxiety and can impair performance. Key evidence: - [[academic-dishonesty-automated-proctoring-ai-2026|The automated-proctoring review]] documents **test-taker anxiety** (especially for users not proficient with online tools) and **proctor/test-taker proficiency gaps causing false malpractice accusations** — the acute stress of being wrongly accused of cheating with AI. - The [[remote-proctoring]] concept page details how **continuous surveillance and fear of false flags** raise test anxiety, and how stressed students perform worse — an equity and [[ethics|fairness]] problem, not just a comfort one. - Institutions adopting remote proctoring must pair AI monitoring with **accessible alternatives, clear communication, and support for test-taker anxiety**, and weigh it against [[authentic-assessment|authentic assessment]] alternatives that reduce surveillance. This cluster connects AI anxiety to [[privacy]], [[academic-integrity]], and [[equity-in-ai-education]]: the stress falls hardest on tool-novice and already-vulnerable students. ## AI anxiety as a learner emotion (barrier and signal) Beyond proctoring, AI use itself generates anxiety — about being replaced, about whether one's work is "really one's own," about competence. Crucially, this anxiety is **not purely negative**: - [[ai-anxiety-strategic-regulation-writing-2026|Kim (2026)]] finds **AI anxiety can be productive**: higher AI anxiety was positively associated with verification and revision behaviors (β=.24, p<.01), and evaluative capacity predicted active [[student-engagement|engagement]] (β=.46, p<.001). Kim's four regulatory types (Uncritical Reliance 18.7%, Selective Integration 34.6%, Evaluative Transformation 31.8%, Strategic Rejection 14.9%) show anxiety-driven scrutiny can transform students into more deliberate, self-regulated users of [[generative-ai|generative AI]] rather than passive adopters. This reframes AI anxiety from a barrier to a potentially useful signal that encourages closer scrutiny. - [[acceptance-ai-english-tools-2026|Acceptance studies]] show anxiety shapes whether learners adopt AI tools, and [[teacher-education-ai-literacy-sdt-2026|teacher-education research]] links AI anxiety to motivation and [[self-regulated-learning|self-regulation]]. - [[aivaluate-anxiety-assessment-2026|AIvaluate]] studies student anxiety during AI-mediated [[assessment|performance-based assessments]], showing assessment anxiety persists and must be designed for. - **Moral panic and educator anxiety:** [[moral-panic-genai-classroom|the moral-panic framing]] shows faculty anxiety about GenAI mirrors earlier panics (calculators, search engines) — a [[teacher-role|teacher]]-side stress response that shapes classroom policy. - **Teachers' "state of vulnerability" and feeling "stuck."** [[farazouli-navigating-uncertainty-teachers-genai-2026|Farazouli et al. (2026)]] capture the *educator* side of AI anxiety directly: 24 Swedish university teachers described the emergence of GAI as alarming and overwhelming, and reported a **state of vulnerability** — low confidence, insecurity, and discomfort driven by limited knowledge of GAI's capabilities, limited exposure, and fear of "not being ahead of students." Teachers felt "stuck" between utopian and dystopian discourses, burdened by amplified responsibility for [[bias-mitigation|fairness]] and quality, and worried about feeling incompetent when assessing student work potentially (co-)produced with AI. This frames teacher AI anxiety as a genuine emotional and professional response to role reconfiguration — not mere resistance — and argues for supporting teacher confidence and well-being, not just tool training. - **Anticipatory guilt: distress that precedes the experience.** [[vassallo-ai-guilt-complex-faculty-2026|Vassallo (2026)]] surveyed academic staff at a Maltese university (109 respondents) and found that guilt about AI attaches to fear of transgression rather than to its experience: the AI Guilt Index (α = 0.88) drew its strongest endorsement from worry that AI use undermines one's credibility (34.9%) and feeling like [[academic-integrity|cheating]] (25.7%), with remorse *after* use last (9.2%). The paradox is that the small group of non-users reported higher guilt (M = 3.25) than any user group (M = 2.32) — because avoiders never test their fears, avoidance can preserve the distress it was meant to prevent. Guilt was highest among early-career academics (M = 2.71) and lowest among senior ones (M = 2.03), and it tracked concealment rather than honesty, correlating with limiting AI use out of unease (r = .62) and with avoiding disclosure to colleagues (r = .50) but not with formal [[ai-use-disclosure|disclosure]] (r = .08). - **Fear of lost professional value, and a way through it.** [[chick-faculty-development-ethical-ai-2026|Chick, Morello & Staffey (2026)]] document an educator anxiety that is existential rather than technical. Ten faculty and staff entered a six-week institute with academic integrity as the leading concern (80%) and [[cognitive-offloading|over-reliance]] second (70%), describing generative AI early on as a "cheating machine", a "threat" or a "dehumanizing force"; a mathematics professor put the identity threat plainly — "I spent years developing expertise in my field. Now a machine can solve problems faster and explain solutions better than I can. What's my value anymore?" Structured, safe experimentation moved the vocabulary toward "curious assistant" and "creative partner" and left 90% reporting positive perceptions, and the authors are explicit that the resistance they observed "stems not from technophobia or stubborn traditionalism but from legitimate concerns about educational quality, equity, and human agency" — a direct counterweight to reading educator anxiety as mere moral panic. **Institutional support works on anxiety through appraisals, not reassurance.** A two-wave survey of 547 Chinese undergraduates ([[school-support-ai-learning-anxiety-control-value-2026|Jiang, Chen & Chen, 2026]]) traced how perceived school support relates to AI learning anxiety through control-value appraisals — the learner's sense of competence (control) and of the tool's usefulness (value). Support predicted lower anxiety directly, but most of its association ran through the appraisals rather than around them: via AI learning self-efficacy, via [[technology-acceptance-model|perceived usefulness]], and via a sequential route in which self-efficacy fed usefulness, which in turn lowered anxiety. A first-order model showed the support *dimensions* were not independently doing the work — only informational support retained a significant path to self-efficacy — and an artificial neural network cross-validation ranked self-efficacy and perceived usefulness as the most stable predictors of anxiety. The implication is that institutional encouragement aimed at anxiety only lands if it changes what students believe about their own capability and the tool's usefulness; general reassurance does not alter the appraisals that generate the anxiety. **Faculty forecasts as a distributed form of AI anxiety.** [[watson-rainie-ai-challenge-faculty-survey-2026|Watson & Rainie (2026)]]'s survey of 1,057 US college and university faculty registers educator anxiety as expectations rather than symptoms: 95% expected generative AI to increase students' over-reliance on the tools, 94% more academic integrity concerns, 90% diminished [[critical-thinking|critical thinking]], 83% shorter attention spans and 81% wider [[equity-in-ai-education|digital inequities]], while 39% believed the tools would diminish the role of faculty and 47% feared the long-term employment impact in their disciplines would be negative. The same respondents were not uniformly pessimistic — 61% still expected improved and customized learning — but 73% had personally dealt with an academic integrity case involving students' generative AI use, which is where a forecast turns into workload. The report is explicitly a non-scientific sample that is not generalizable, so it documents the sector's expressed fears rather than measured effects. This direction connects AI anxiety to [[motivation]], [[ai-literacy]], [[student-experience]], and [[self-regulated-learning]]. ## Career-related AI anxiety A distinct and increasingly studied dimension is **career anxiety** — the fear that AI will displace jobs, erode employability, or devalue one's professional future. Where proctoring anxiety is situational and learner-emotion anxiety is about in-task competence, career anxiety is forward-looking and identity-level, and it is tightly linked to [[career-development-and-readiness]]. The [[kim-ai-anxiety-comprehensive-analysis|AI Anxiety comprehensive analysis]] identifies the **fear of replacement by AI** as the primary contributor to AI anxiety, alongside uncontrolled AI growth, privacy, misinformation, and bias. A growing empirical literature now quantifies how this fear affects students: - **Career adaptability is a protective factor.** [[wang-career-adapt-abilities-ai-anxiety-english-2026|Wang (2026)]] shows career adapt-abilities significantly and negatively predict AI anxiety among English majors, with core self-evaluations partially mediating the relationship; the low-adaptability group had the highest AI anxiety. - **AI anxiety impairs career decisions.** [[duan-ai-anxiety-career-decisions-college-2026|Duan et al.]] use structural equation modeling to show AI anxiety directly and negatively predicts career decisions, and does so largely by undermining **career adaptability** (accounting for 63.35% of the total effect); [[self-efficacy]] offered only limited buffering. - **AI anxiety predicts job-search anxiety at scale.** [[ustun-ai-anxiety-job-finding-anxiety-2026|Üstün & Danacıoğlu]] (1,057 students) and [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al.]] (821 health-sciences students) find AI anxiety and negative AI attitudes predict job-finding/job-search anxiety, with women, social-science majors, and lower-income students most affected. **Practical implication:** building [[career-development-and-readiness|career readiness]] — adaptability, core self-evaluations, and employer-valued AI skills — is a validated intervention for reducing career-related AI anxiety, more than generic self-efficacy alone. Institutions should universalize [[ai-literacy]] and career-planning support, and address the [[equity-in-ai-education|equity]] patterning of career anxiety. ## AI for detecting and relieving stress and anxiety The positive side: AI systems increasingly detect and help alleviate stress and anxiety. - [[ai-campus-wellbeing-tools|AI campus well-being tools]] span prevention (improved feedback collection) and intervention (advancing mental-health detection). - [[affective-text-wearable-student-health|Affective text + wearable sensing]] (a year-long study of 458 students with Oura rings) shows ultra-brief naturalistic text can complement wearable physiological sensing for longitudinal student health monitoring — a concrete AI-enabled stress-detection pathway. - This links to [[affective-computing]] and [[affective-tutoring]], where AI reads and responds to emotional state. - **Which stressors the model weighs — and why context matters.** [[culturally-aware-student-stress-chatbot-2026|Bashir and Afzal (2026)]] offer a [[machine-learning]] window on which stressors actually drive student distress in a [[global-south|non-Western context]]. Feature-importance analysis on 1,100 survey responses put blood pressure first (15.6%) and **teacher-student relationship second (10.0%)** — ahead of sleep quality (9.3%), depression (8.3%) and social support (7.6%) — while anxiety level ranked ninth at 4.8%, which the authors read as evidence that student stress is multi-dimensional rather than driven by a single psychological indicator. They attribute the salience of the teacher-student relationship to the comparatively hierarchical educational environment in Pakistan and present it as a hypothesis for locally collected data, underlining that stress models and their feature weights are context-dependent and cannot be assumed to transfer across student populations. ## Why this is distinct from well-being [[well-being]] is the broad positive state (emotional, psychological, social health) that AI can support or undermine. **AI anxiety and stress** is the specific, measurable emotion-construct within that space — it names the discrete negative affect and its productive uses, and it has its own dedicated evidence base (proctoring stress, productive AI anxiety, AI-driven stress detection). Rather than rivaling well-being, this page develops the anxiety/stress dimension in depth and cross-links heavily to [[well-being]], [[social-emotional-learning]], [[affective-computing]], and [[remote-proctoring]]. ## Practical guidance - **Design for anxiety, not just integrity.** Proctoring and AI-monitoring systems should minimize false accusations and surveillance stress, especially for novice and vulnerable students; pair monitoring with accessible alternatives and clear communication. - **Leverage productive anxiety.** Rather than only reducing AI anxiety, support the verification, revision, and self-[[regulation]] that anxious-but-engaged students already exhibit. - **Use AI to detect and relieve stress** — via affective computing, wearables, and campus well-being tools — while guarding privacy. - **Consider educator anxiety** in AI adoption and policy, not just student experience. ## Connected Concepts - [[career-development-and-readiness]] — career readiness as a protective factor against AI anxiety - [[well-being]] — the broad umbrella this concept develops in depth - [[remote-proctoring]] — the primary source of surveillance/false-accusation stress - [[social-emotional-learning]] — the emotional competencies involved - [[affective-computing]] — AI reading emotional state - [[affective-tutoring]] — AI responding to affect - [[academic-integrity]] — the integrity/cheating-accusation dimension - [[privacy]] — the surveillance concern underlying proctoring stress - [[student-experience]] — anxiety as part of the learner experience - [[motivation]] — anxiety's effect on adoption and engagement - [[self-regulated-learning]] — the productive-anxiety link - [[ai-literacy]] — building capacity that reduces unfounded anxiety - [[equity-in-ai-education]] — disproportionate stress on vulnerable students - [[ethics]] — the fairness of surveillance and false accusations - [[teacher-role]] — educator-side anxiety - [[assessment]] — AI-mediated assessment anxiety - [[social-norms-ai-use]] — the affective cost of visibility ## Connected Articles - [[ai-anxiety-strategic-regulation-writing-2026]] — AI anxiety as a productive strategic-regulation signal - [[aivaluate-anxiety-assessment-2026]] — Student anxiety in AI-mediated performance-based assessment - [[ai-campus-wellbeing-tools]] — AI-driven tools for campus well-being - [[affective-text-wearable-student-health]] — Affective text + wearable sensing for student health - [[academic-dishonesty-automated-proctoring-ai-2026]] — Automated proctoring, test-taker anxiety, false accusations - [[automated-online-exam-proctoring-decade-review-2026]] — Decade-long review of automated exam proctoring - [[moral-panic-genai-classroom]] — Faculty anxiety and moral panic about GenAI - [[acceptance-ai-english-tools-2026]] — Anxiety shaping AI tool acceptance - [[teacher-education-ai-literacy-sdt-2026]] — Teacher education, AI literacy, and anxiety - [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination]] — AI use intertwined with anxiety, trust, confidence - [[ortiz-bonnin-chat-or-cheat-chatgpt-dishonesty-2025]] — Risk perception and academic-dishonesty anxiety suppress ChatGPT use - [[qu-wang-disclose-or-not-genai-2026]] — The disclosure dilemma as a source of student stress - [[conijn-fear-big-brother-proctored-exams-2022]] — Proctored exams raise test anxiety without reducing cheating - [[kim-ai-anxiety-comprehensive-analysis]] — Comprehensive analysis of AI anxiety and interventions - [[wang-career-adapt-abilities-ai-anxiety-english-2026]] — Career adapt-abilities reduce AI anxiety - [[duan-ai-anxiety-career-decisions-college-2026]] — AI anxiety impairs career decisions via career adaptability - [[ustun-ai-anxiety-job-finding-anxiety-2026]] — AI anxiety and attitudes predict job-finding anxiety - [[dag-ai-perceptions-career-anxiety-health-2026]] — AI anxiety predicts job-search anxiety in health sciences - [[farazouli-navigating-uncertainty-teachers-genai-2026]] — University teachers' experiences and perceptions of GAI: vulnerability, rethinking assessment, student learning at risk (Farazouli et al. 2026) - [[culturally-aware-student-stress-chatbot-2026]] — An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning - [[school-support-ai-learning-anxiety-control-value-2026]] — Perceived school support lowers AI learning anxiety mainly through control-value appraisals (69.5% mediated); ANN cross-validation (Jiang, Chen & Chen 2026) - [[vassallo-ai-guilt-complex-faculty-2026]] — The AI Guilt Complex: anticipatory guilt exceeding post-use remorse among academic staff (Vassallo 2026) - [[watson-rainie-ai-challenge-faculty-survey-2026]] — AAC&U/Elon survey of 1,057 US faculty: forecasts of over-reliance, integrity concern and a diminished faculty role (Watson & Rainie 2026) - [[chick-faculty-development-ethical-ai-2026]] — From fear to curiosity: a six-week faculty institute moving instructors through identity threat (Chick, Morello & Staffey 2026) - [[beyond-the-algorithm-academic-developers-digital-mediators-2026]] — Academic developers' impostor feelings and ethical discomfort as affective labor in AI-mediated work --- ## [Student Experience](https://edtechdev.github.io/aied/concepts/student-experience/) > **Student experience** — how learners perceive, interact with, and are affected by AI tools in educational settings. With over 85 articles in the knowledge base, student experience is one of the most-researched dimensions of [[ai-education|AI in education]]. AI impacts students in both positive and negative directions, and the same tool can help and harm depending on how it is designed and used. ## Questions to Consider - The page claims AI affects students in both positive and negative directions, often simultaneously — the same tool can help and harm depending on design and use. Can you give an example from your own experience where AI helped and hurt at the same time? - It describes a 'performance-learning gap': students do better with AI but worse on later unassisted tasks. How do you think that gap arises, and what would it take to close it? - If over-reliance on AI means delegating the reasoning you actually need to practice, where would you draw the line between legitimate help and offloading that erodes learning? - The [[research-methods-aied|research]] asks whether simply knowing AI is available changes student effort. Do you think awareness of AI makes students work harder, less hard, or differently — and how would you test your belief? - Given that AI access and effectiveness vary across student populations, what equity concerns do you think matter most when a course adopts an AI tool, and who is responsible for addressing them? ## Introduction ### How student experience is studied - **Large-scale surveys:** [[ai-in-the-wild-college|AI in the Wild]] analyzes authentic interactions of thousands of college students, while [[genai-availability-grades-satisfaction|availability and satisfaction studies]] correlate AI access with student outcomes. These are [[self-report-measures|self-report measures]]: they capture perceptions and intentions well and behavior only approximately. - **Interaction patterns:** [[tracing-genai-literacy-interaction-patterns|Tracing GenAI literacy]] maps how students engage with AI across assignments. [[misiejuk-cognitive-offloading-prompting-2026|Prompting analysis]] reveals cognitive engagement levels through prompt structure. - **Motivation and agency:** [[ai-availability-student-motivation|AI availability and motivation]] examines whether knowing AI is available changes student effort. [[aied-unfinished-mission-bypass|AIED's unfinished mission]] frames [[agency]] and [[motivation]] as central challenges. - **Perceptions and attitudes:** [[genai-usage-design-students-survey|GenAI usage surveys]] and [[student-mental-models-genai|mental model studies]] investigate how students understand and trust AI. - **Equity dimensions:** student-experience intersects with [[equity-in-ai-education]] — AI access and effectiveness vary across student populations. ## Ways AI impacts students AI affects students across cognitive, motivational, [[affective-computing|affective]], identity, social, and equity dimensions. The research points to **positive and negative impacts in each dimension**, often simultaneously — the direction depends on design and use. ### Cognitive impacts - **Positive:** AI can [[scaffolding|scaffold]] learning with [[feedback]], hints, and explanations, supporting understanding, practice, and [[help-seeking]]. Students can use AI to [[ai-literacy|learn how to use AI well]], and well-designed tools keep the learner doing the cognitive work (see [[does-ai-help-students-learn|Does AI help students learn?]]). - **Negative:** AI can drive [[cognitive-offloading|over-reliance and cognitive offloading]], where students delegate the reasoning they need to practice. This is the **performance–learning gap**: students do better *with* AI but worse on later unassisted tasks (see [[ai-misuse-learning-harm|AI Misuse and Learning Harm]]). Overuse may also erode [[metacognition]] and [[self-regulated-learning]]. ### Motivational impacts - **Positive:** AI can raise [[motivation]] and [[student-engagement|engagement]] by providing personalized, immediate, and low-stakes support — helping students persist and feel competent (see [[self-determination-theory|self-determination]] perspectives on autonomy, competence, and relatedness). - **Negative:** Knowing AI is available can reduce student effort and motivation to struggle productively ([[ai-availability-student-motivation|AI availability and motivation]], [[wang-safety-gap-productive-struggle-2026|the safety gap]]). Over-reliance can erode [[agency]] and the sense of accomplishment that comes from doing work oneself. ### Affective and well-being impacts - **Positive:** AI can offer low-pressure, on-demand help and reduce anxiety about asking questions, supporting [[well-being]] and confidence. - **Negative:** AI use is associated with [[anxiety-and-stress|anxiety and stress]], including fears about being replaced, uncertain assessment, and the pressure to keep up. Studies such as [[kim-ai-anxiety-comprehensive-analysis|a comprehensive analysis of AI anxiety]] and [[aivaluate-anxiety-assessment-2026|AIvaluate]] document these affective costs. [[shame-guilt-ai-regulation-computing-education|Shame and guilt]] around AI use can drive hiding and selective disclosure, harming honest engagement and [[social-emotional-learning|social-emotional]] well-being. These pressures are local and social rather than written down: see [[social-norms-ai-use|social norms of AI use]]. Assessment conditions add their own pressure: [[harerimana-remote-proctoring-nursing-scoping-2026|a scoping review of remote proctoring in nursing assessment]] finds students anxious about connectivity and about being wrongly accused of cheating, first-time users of an invigilation app describing it as anxiety-inducing and reporting difficulty concentrating while watched, and roughly a fifth hitting browser-extension or connectivity failures despite preparatory resources. ### Identity impacts - **Positive:** AI can support [[learner-identity|identity formation]] by scaffolding disciplinary belonging, confidence, and professional aspirations — e.g., helping students see themselves as capable practitioners. - **Negative:** AI can threaten [[learner-identity|learner identity]] through authorship loss and competence doubt — when AI produces the work, students may stop feeling it is "theirs." The [[t2i-competence-paradox-2026|competence paradox]] in creative fields shows ease-of-use undermining the craft-based identity students derive from authorship. ### Academic-integrity and fairness impacts - **Negative:** AI enables new forms of [[academic-integrity|academic dishonesty]] (AI-generated essays, unauthorized completion), driving debates about [[reduce-ai-cheating|detection and reduction]]. This interacts with [[equity-in-ai-education|equity]]: unequal access to, and understanding of, AI tools can widen gaps between students. - **Positive/constructive:** AI can support [[authentic-assessment|authentic, process-oriented assessment]] and reflective practice (e.g., [[pedlow-genai-selfassessment-2026|guided self-assessment]]), turning integrity concerns into opportunities for [[ai-literacy]] and responsibility. Student accounts of integrity are less settled than the dishonesty framing suggests. [[mulisa-students-genai-integrity-perspectives-2026|Mulisa and Mezgebu (2026)]] interviewed 27 undergraduates at an Ethiopian university and found the student body divided against itself: almost all used GenAI or watched peers use it and most credited it with raising their achievement, a minority called coursework use outright misconduct, and the sharpest and most widely shared complaint was fairness — AI users scoring above students who worked honestly, which some described as killing their sense of diligence and left one participant unsure "whether we are benefiting or suffering from the use of AI." The procedure side matters too: [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] coded 1,162 GenAI misconduct cases and found that the evidence most often cited — detector output, similarity reports, AI-typical content patterns — carried the weakest probative value, and that with no minimum evidentiary threshold in the pipeline students with thin cases were pushed toward appeals. Over-inclusive definitions broaden that exposure: [[wright-transcription-not-generation-2026|Wright (2026)]] shows that prohibitions aimed at "[[generative-ai|generative AI]]" can catch tools that merely convert the format of work a student already authored, an over-inclusion that falls hardest on disabled and [[equity-in-ai-education|equity]]-exposed students. [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] points the constructive way out, treating integrity as a [[pedagogy|pedagogical]] practice enacted through [[evaluative-judgment|judgment]] — annotated decision trails, verification, oral defense, version history — rather than compliance secured through surveillance. ### Social and relational impacts - **Positive:** AI can mediate [[collaborative-learning|collaboration]] and [[student-ai-interaction|human-AI interaction]], supporting teamwork, peer interaction, and access to diverse perspectives. - **Negative:** AI can reduce genuine peer and instructor interaction, create [[ai-sycophancy|sycophantic]] dynamics, and — when operating as an undisclosed teammate — reshape group discourse and [[agency]] in ways learners cannot see or contest (see [[jin-emergent-learner-agency-implicit-hai-2026|emergent learner agency in implicit HAI]]). ### Long-run / capability impacts - **Positive:** AI can help students build transferable skills for an AI-integrated workplace — [[lodge-adaptive-capabilities-genai-future-2026|adaptive capabilities]] such as [[ai-literacy]], [[distributed-cognition]], and [[metacognition]] — and raise expectations about AI-skills readiness ([[ithaka-sr-ai-skills-college-graduates-2026|AI-skills expectations for graduates]]). - **Negative:** An over-reliant or unreflective AI experience can leave students less able to perform without AI, less practiced at independent reasoning, and uncertain of their own capabilities (see [[ai-misuse-learning-harm|AI misuse and learning harm]]). **Overall:** the same AI tool can support or undermine students depending on design and use. The guardrail throughout is to keep the learner doing the cognitively important work while using AI for support ([[scaffolding|scaffold, do not substitute]]), and to attend to the full range of impacts — not just performance. One configuration shifts where the experience begins: when AI generates the course readings themselves rather than helping with homework, students become auditors of their own [[curriculum-design|curriculum]]. In Sidorkin's (2026) graduate course, students valued the contextual specificity and adjustability of the generated texts and 75 percent agreed they learned more than in a comparable course without an AI companion, yet they had to infer source quality from context because Wikipedia links and peer-reviewed citations appeared in the same lists without labels, and four of 24 survey respondents used dependence language, including one describing themselves as "somewhat codependent on the AI for reassurance and structure." ## Connections Student experience connects to [[cognitive-offloading|Over-Reliance]] (excessive AI dependence), [[ai-literacy]] (skills for effective use), [[cognitive-offloading]] (how AI changes cognitive work), and [[student-engagement|engagement]] (how AI systems measure and respond to student behavior). It is the learner-facing member of the [[stakeholders]] umbrella, and the home for summarizing all the ways AI impacts students. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[learner-identity]] — evolving disciplinary, professional, creative, and academic learner identities - [[agency]] - [[well-being]] - [[anxiety-and-stress]] - [[social-emotional-learning]] - [[remote-proctoring]] - [[generative-ai]] - [[llm]] - [[higher-ed]] - [[ai-literacy]] - [[cognitive-offloading]] - [[equity-in-ai-education]] - [[k-12]] - [[scaffolding]] - [[metacognition]] - [[self-regulated-learning]] - [[framing-ai-use-for-students]] - [[academic-integrity]] - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) - [[self-report-measures]] - [[student-engagement]] — how AI systems measure and respond to student behavior - [[social-norms-ai-use]] — how AI use becomes visible to peers ## Connected Articles - [[shame-guilt-ai-regulation-computing-education]] — Shame and guilt as social regulators of AI use - [[t2i-competence-paradox-2026]] — The competence paradox: creative identity in text-to-image GenAI use - [[best-response-student-ai-dialog-2026]] - [[chatgpt-perception-online-learning-engagement-2026]] - [[ai-tools-academic-work-cheating-2026]] - [[genai-student-experiences-uk-he-survey-2026]] - [[metacognitively-discordant-completion-genai-2026]] - [[ai-generated-interactive-fiction-education-2026]] - [[ai-in-the-wild-college]] - [[genai-availability-grades-satisfaction]] - [[student-rationalization-ai-writing]] — Student rationalization of AI use in academic writing (Kim et al. 2026) - [[tracing-genai-literacy-interaction-patterns]] - [[misiejuk-cognitive-offloading-prompting-2026]] - [[ai-availability-student-motivation]] - [[aied-unfinished-mission-bypass]] - [[student-mental-models-genai]] - [[genai-usage-design-students-survey]] - [[spritz-ai-disciplinary-mediation-student-teams-2026]] - [[student-llm-interaction-taxonomy-review-2026]] - [[ithaka-sr-ai-skills-college-graduates-2026]] — AI-skills expectations for college graduates vs. institutional readiness - [[student-ai-inquiry-types-cs2-2026]] — Analysis of Types of Inquiries in Student-AI Interaction - [[lnenicka-secondary-students-genai-stem-2026]] — What secondary students do with GenAI tools across STEM - [[dai-chatbots-problem-posing-primary-2026]] — GenAI chatbots and problem posing in primary science - [[zuo-instructor-power-genai-writing-2026]] — Power relations perceived by college instructors grappling with GenAI in writing (Zuo, Xu & Dunning 2026) - [[bassett-ai-detectors-education-2026]] — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026) - [[wang-safety-gap-productive-struggle-2026]] — The Safety Gap: Restoring Productive Struggle - [[aivaluate-anxiety-assessment-2026]] — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026) - [[kim-ai-anxiety-comprehensive-analysis]] — A comprehensive analysis of AI anxiety - [[student-perceptions-ai-study-productivity-2026]] — Students' Perceptions of Artificial Intelligence Tools for Study Productivity and Learning: An Exploratory Survey Study - [[teo-ai-adoption-tertiary-meta-analysis-2026]] — Post-secondary adoption perspective - [[pedlow-genai-selfassessment-2026]] — Raising ethical awareness of GenAI use through student self-assessment - [[dollinger-equitable-assessment-ai-2026]] — Equitable assessment in an AI era - [[longitudinal-ai-usage-ethics-policy-teacher-education-2026]] — Longitudinal GenAI usage, ethics, and policy in teacher education (Parker et al. 2026) - [[genai-use-usefulness-student-experience-australia-2026]] — Student experience of GenAI usefulness in Australian higher ed (Chung et al. 2026) - [[genai-decision-capability-cognitive-load-2026]] — GenAI and students' perceived decision capability (cognitive-load account) - [[genai-professionalization-metaphors-2026]] — GenAI conceptualizations and student professionalization - [[sidorkin-ai-generated-course-readings-2026]] — Students as auditors of their own AI-generated curriculum (Sidorkin 2026) - [[mulisa-students-genai-integrity-perspectives-2026]] — Students on whether GenAI is a cheating tool or a learning partner - [[munoz-misconduct-allegation-evidence-2026]] — What misconduct allegation files actually contain as evidence - [[wright-transcription-not-generation-2026]] — Over-inclusive AI rules and the students they catch - [[sharma-judgment-visible-genai-assessment-2026]] — Integrity as evaluative judgment rather than compliance - [[harerimana-remote-proctoring-nursing-scoping-2026]] — Remote proctoring's emotional and equity costs for students --- ## [Social Norms of AI Use](https://edtechdev.github.io/aied/concepts/social-norms-ai-use/) > **Social norms of AI use** — the informal, locally enforced rules that decide when learners and instructors may use generative AI, how openly they can admit to it, and what counts as [[academic-integrity|cheating]] in practice rather than in policy. The corpus is thin on norms as a topic and thick on the mechanisms that produce them. Learners cannot reliably infer each other's AI use: the median correlation between a perceiver's estimates and their own self-reported usage profile was r = .73 at both waves, and rank accuracy for the *kind* of AI use was r = .34 at the second wave. [[shame-guilt-ai-regulation-computing-education|Shame and guilt regulate visibility rather than use]], producing hiding and selective disclosure instead of behavior change. Disciplinary norms diverge sharply: AI-integrated assignments appeared in 27% of business syllabi and 5% of humanities syllabi across 31,000 courses. And the formal [[ethics]] literature routes the fewest and least actionable norms to end users, meaning the people subject to classroom norms are largely absent from the documents that discuss them. ## Questions to Consider - If students cannot tell who is using AI, what are the informal rules actually based on? - Where do AI norms come from when an institution has no policy, or when the policy contradicts what instructors say in class? - Why would a student who has never used AI report more guilt about it than a student who uses it weekly? - What does shame accomplish that a policy cannot, and at what cost to [[help-seeking|asking for help]]? - Does surveillance change behavior, or only how visible that behavior is? - Whose norms dominate in a group assignment, and who absorbs the cost of the group's shared AI habits? - Should students have a hand in writing the norms they are held to? ## Introduction Formal [[regulation]] of AI in education arrives as policy: acceptable-use statements, syllabi clauses, [[ai-detection|detection]] procedures, [[remote-proctoring|proctoring]] systems, honor codes. Everything students actually experience, though, sits underneath that layer. A classmate who never mentions using ChatGPT, an instructor who says "I don't care how you write it" and then bans it in the syllabus, a study group that has quietly settled on what is fair, a department where everyone assumes the grading curve has already shifted. These are norms: shared expectations about acceptable behavior, enforced by approval, disapproval, and the risk of being seen. This page collects what the evidence base says about that layer. It is worth separating carefully, because "hidden curriculum" is often used loosely to mean anything unofficial. Here the subject is narrower: the operative rules about AI use, how they form, how they are enforced, and how weak the connection is between them and the written rules. The corpus has almost no study that measures norms directly, which is itself a finding. What it has is repeated evidence about the mechanisms: inference, disclosure, emotion, and the disciplinary variation that shapes what any local norm can become. ## What students believe about each other Norms require information about other people's behavior, and that information is worse than most instructors assume. [[student-perception-ai-use-collaboration|A study of student pairs]] tracked how accurately learners judged a partner's AI use across pair-programming sessions. Against each partner's own self-reported usage profile, the median correlation was r = .73 at both waves, high enough to feel like knowledge and low enough that many individual estimates were wrong. Rank accuracy for the type of AI use was r = .34 at the second wave, and 25% of teams became worse at judging their teammates over time. Two details matter for norms. First, misalignment did not shrink with repeated face-to-face work. Familiarity did not produce accuracy, which undermines the assumption that norms stabilize as groups get to know each other. Second, the authors' proposed fix was structural rather than attitudinal: shared prompt histories, AI-use annotations, or peer-visible records of AI-supported work. Making AI use visible is a different project from making students more honest or more perceptive about it, and only the first one has evidence behind it. ## Visibility, shame, and selective disclosure If peers cannot see AI use accurately, then the rules that develop around it are enforced largely through social risk. [[shame-guilt-ai-regulation-computing-education|An interview study with 19 computing students]] examined shame and guilt as social regulators and found that they govern disclosure rather than use. Students described hiding behaviors and selective disclosure, and they reported shaming themselves, their peers, and their faculty. The important negative result is that these emotions coexisted with continued AI use, generating cycles of reduced [[agency]] and moral tension rather than prompting anyone to stop. That pattern relocates the norm from "use or don't use" to "admit or don't admit." A classroom can have a strong anti-AI norm on paper, widespread use in practice, and an equilibrium of silence in between, without anyone being confused about what the rule says. It also predicts where the cost lands: on learners who are least able to manage the social risk, and on [[equity-in-ai-education|learners whose visible use carries different consequences]]. The instructor side of the same emotion is measurable. [[vassallo-ai-guilt-complex-faculty-2026|A survey of faculty]] using a four-item AI Guilt Index (Cronbach's alpha = 0.88) found anticipatory guilt exceeding experienced remorse: the mean index score was 2.39 (SD = 1.01), with 22.9% of participants above the midpoint. Non-users of AI reported *higher* guilt (M = 3.25, SD = 1.04, n = 8) than users (n = 101, M = 2.32, SD = 0.97), t(8.39) = 2.00, p = .078, and the remorse item separated the groups more sharply, t(7.79) = 2.44, p = .041. Cluster analysis produced four profiles: Comfortable Adopters (26.6%), Guilty Non-Users (29.4%), Cautious Users (28.4%) and Morally Distressed Avoiders (15.6%), with differences confirmed on guilt, F(3, 105) = 22.13, p < .001, eta-squared = 0.39. A norm that makes people who have not acted feel guilty is doing its work through anticipation. That is a real regulatory mechanism, and it is invisible to any policy audit that looks only at what faculty and students do. ## The local rulebook: disciplines and classrooms Norms are local, and the evidence on how local is unusually concrete. [[chirikov-regulate-ai-syllabi-2026|An analysis of 31,000 course syllabi]] found that AI-integrated assignments ranged from 27% in business to 5% in the humanities. The range is not a gap in policy compliance; it reflects different disciplinary judgments about what the work is, what practice is worth, and which skills a credential is supposed to certify. The authors' recommendation follows from that: grant instructors autonomy within [[governance|disciplinary norms]] rather than issuing one-size-fits-all mandates, because the mandate would have to be written for an average that does not exist. The classroom layer can also contradict the document layer in the same course. [[sobo-cheating-competing-ai-marketing-literacy-2025|Interviews with marketing students]] found peer norms and "keeping up" driving adoption, with word of mouth described as the biggest promotion AI has, sustained by fear of falling behind. The same study recorded a contradiction that instructors will recognize: some classes required AI while the syllabus banned it, and an instructor's in-class encouragement did not always survive into the written policy. When those two disagree, students treat the spoken norm as the operative one, and the syllabus as a liability to be managed. [[group-work|Group work]] is where informal norms do the most work, because the unit being graded is collective. [[chen-zou-genai-group-assessment-agency-2026|A study of 52 pre-service teachers]] across 15 focus groups identified distinct patterns rather than one enthusiasm-to-avoidance line, with groups negotiating what counted as legitimate assistance in the absence of a shared rule. Any student who has watched a group quietly settle on one member's AI use has seen a norm being written in real time, without a policy process and usually without a record. ## Formal ethics, lived norms The professional literature on AI ethics and the classroom norm it is supposed to reach are drifting apart, and the gap is documented. [[agarwal-ethical-values-norms-aied-2026|A systematic review of ethical values and norms]] extracted norms from 15 articles and mapped them onto four stakeholder sets from Smuha (2022). Developers received the most norms, followed by educational institutes and regulators. End users received the fewest and the least actionable, with no norms addressing students directly on non-discrimination, data stewardship, or educational aptness, and student voices essentially absent from the material. The result is a literature that specifies obligations for the people building and procuring AI, and leaves the people living with it in the classroom to work out their own rules. That is a plausible explanation for why the norms students describe, as in the shame and guilt study, are so often about visibility and reputation rather than about the values the ethics documents name. The two systems are regulating different things. ## Why surveillance does not settle norms When institutions try to replace informal norms with formal monitoring, the evidence is consistent and unflattering. [[conijn-fear-big-brother-proctored-exams-2022|A four-wave study of 1,760 students across 105 courses]] found that online proctoring significantly increased test anxiety, and had no effect on the temptation to cheat. It also did not change perceived difficulty or exam performance. The anxiety cost was concentrated among already-vulnerable students, including those with weaker home environments and less reliable technology. Two further studies explain why surveillance is a poor tool for norm-setting. [[harerimana-remote-proctoring-nursing-scoping-2026|A scoping review of remote proctoring]] found that monitoring was widely perceived as deterring misconduct, yet South African lecturers reported continued dishonesty despite active monitoring, a situation Khalil et al. (2022) call "subterranean ethics," and frequent minor alerts generated false positives and faculty review work. Surveillance buys compliance through fear of detection rather than commitment to a shared standard. [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|A model of assessment under imperfect information]] sharpens the mechanism: deterrence depends on a detector's ability to discriminate between hidden use and legitimate work rather than on its catch rate, and when false positives rise faster than true positives, students disengage from the system instead of complying with it. Its further result is the one that bears directly on norms: monitoring lowers the attractiveness of hidden use, but it does not increase [[ai-use-disclosure|disclosure]]. Students move toward openness only when reporting is safer or more valuable than concealment, and a prohibition can make concealment more attractive than compliance. [[teichmann-detecting-undetectable-misconduct-2026|The procedural-justice argument]] draws the conclusion: because skilled or lightly edited AI use is undetectable in the general case, a misconduct procedure built on detection produces unfairness without effectiveness, and the answer is better [[assessment]], not better surveillance. ## What this means for practice - **Make use visible instead of inferring it.** Students cannot judge each other accurately, and repeated collaboration does not fix it. Lightweight shared records of AI-supported work, annotations, or prompt histories give a group something factual to reason about. - **Expect the spoken norm to beat the written one.** If what is said in class contradicts the syllabus, students will follow the classroom. Alignment between the two is worth more than a stricter clause. - **Treat disclosure as a design problem, not a virtue problem.** Disclosure becomes likely when reporting is safer or more valuable than hiding. That means making disclosed use legible in the assessment itself, rather than asking for honesty as a character trait. - **Respect disciplinary variation.** A 27% to 5% spread across departments is not inconsistency to be normalized away, it is different judgments about what the credential certifies. Institution-wide mandates tend to be written for a discipline that does not exist. - **Do not reach for monitoring to set norms.** Proctoring raised anxiety without changing the temptation to cheat in the strongest study available, and it concentrates that cost on the least advantaged students. - **Ask who is missing from the norm.** The ethics literature assigns students almost no obligations and takes little account of their judgment, while the classroom assigns them a great deal and enforces it socially. Closing that gap is a curriculum question, not a compliance question. ## Connected Concepts - [[academic-integrity]] — the formal rulebook these norms grow alongside - [[ai-use-disclosure]] — the practice norms regulate most directly - [[framing-ai-use-for-students]] — how a classroom states its expectations - [[reducing-ai-misuse]] — the instructional counterpart to norm-setting - [[learner-identity]] — who a learner is taken to be when AI use becomes visible - [[ai-detection]] — the technical response that mostly fails - [[remote-proctoring]] — surveillance as an attempted substitute for norms - [[trust]] — what norms and disclosure both depend on - [[equity-in-ai-education]] — who absorbs the social cost - [[help-seeking]] — the behavior shame suppresses first - [[group-work]] — where group norms are negotiated - [[student-experience]] — the lived side of all of this - [[anxiety-and-stress]] — the affective residue - [[governance]] — the formal layer these norms sit under - [[ethics]] — the literature that has least to say to students ## Connected Articles - [[student-perception-ai-use-collaboration]] — Students' perception accuracy of partners' AI use and its relation to collaboration performance - [[shame-guilt-ai-regulation-computing-education]] — Shame and guilt as social regulators of AI use in computing education - [[chirikov-regulate-ai-syllabi-2026]] — How instructors regulate AI in college: evidence from 31,000 course syllabi - [[vassallo-ai-guilt-complex-faculty-2026]] — The AI guilt complex: moral emotions and ethical dilemmas in academic technology adoption - [[agarwal-ethical-values-norms-aied-2026]] — Identifying the ethical values and norms for artificial intelligence in education - [[conijn-fear-big-brother-proctored-exams-2022]] — The fear of Big Brother: the potential negative side-effects of proctored exams - [[harerimana-remote-proctoring-nursing-scoping-2026]] — Under surveillance: mapping remote proctoring practices in nursing assessment - [[mohamed-temimi-assessment-imperfect-information-disclosure-2026]] — Assessment design under imperfect information: generative AI, disclosure, and student response - [[teichmann-detecting-undetectable-misconduct-2026]] — Detecting the undetectable: misconduct procedures after generative AI - [[sobo-cheating-competing-ai-marketing-literacy-2025]] — Cheating or competing? AI in marketing education - [[chen-zou-genai-group-assessment-agency-2026]] — Agency in GenAI-supported group assessment --- ## [Career Development and Readiness](https://edtechdev.github.io/aied/concepts/career-development-and-readiness/) > **Career development and readiness** — the processes and capacities that prepare learners to build, adapt, and sustain a career in an AI-disrupted labor market: career adaptability, employability, workforce readiness, and the skills (including [[ai-literacy]]) that employers value. In the AI-in-education context this concept is increasingly important because AI both reshapes the skills graduates need and generates **career-related AI anxiety** about job displacement — making career readiness a protective factor for student [[well-being]]. ## Questions to Consider - Career readiness here is framed as adaptability — the capacity to navigate, adjust, and thrive across changing roles — rather than just credentialing. How does that reframing change what you think a career-focused education should actually build? - [[research-methods-aied|Research]] consistently shows career adapt-abilities significantly reduce AI anxiety about job displacement. Why might being more adaptable to career change make a student feel less anxious about a technology that threatens their current job prospects? - One key claim is that AI literacy is necessary but not sufficient for career readiness — students also need adaptability and positive self-evaluations. Can you think of someone who is highly AI-literate yet still anxious or unprepared for the workforce? What were they missing? - An ITHAKA S+R report documents a skills-prioritization gap: instructors emphasize critical, responsible use of AI, while employers favor workflow automation and human-AI teaming skills — and they agree on only one of 26 AI skills. Why do you think classroom and workplace priorities diverge so sharply, and who should adapt? - The report finds most institutions lack both a consensus on what AI skills look like and an assessment framework for them. If you had to define and assess 'workforce-ready AI skills' for a graduating student, what would you measure and how? - Fear of replacement by AI is identified as a primary driver of AI anxiety, and career readiness is positioned as a protective factor for student well-being. How should an education program address the anxiety itself, rather than only adding skills? ## Introduction As AI transforms occupations, education's role in career development has broadened from credentialing toward building **adaptability** — the capacity to navigate, adjust, and thrive across changing roles. This concept connects education to employability and links to [[anxiety-and-stress]]: students with stronger career adapt-abilities experience less AI anxiety. ## How career development and readiness appears in the knowledge base - **Career adaptability reduces AI anxiety.** [[wang-career-adapt-abilities-ai-anxiety-english-2026|Wang (2026)]] shows career adapt-abilities significantly and negatively predict AI anxiety among English majors, with core self-evaluations partially mediating the relationship; the low-adaptability group showed the highest AI anxiety. [[duan-ai-anxiety-career-decisions-college-2026|Duan et al.]] confirm the mechanism with SEM: AI anxiety impairs career decisions largely through eroded career adaptability (63.35% of the total effect), and self-efficacy offered limited buffering. [[ustun-ai-anxiety-job-finding-anxiety-2026|Üstün & Danacıoğlu]] add that AI anxiety and negative AI attitudes predict post-graduation job-finding anxiety across 1,057 students, with women, social-science majors, and second-years most affected. [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al.]] extend this to health-sciences students (821, r = 0.233). Career readiness is thus an empirically validated buffer against [[anxiety-and-stress|career-related AI anxiety]]. - **AI literacy is necessary but not sufficient.** [[ai-literacy-career-adaptability-business-2026|Testa et al.]] argue AI literacy alone is not enough for career readiness — students also need adaptability and positive self-evaluations, directly linking [[ai-literacy]] to career outcomes. - **Employer and graduate perspectives.** [[ithaka-sr-ai-skills-college-graduates-2026|The ITHAKA S+R report]] (500 US four-year-college instructors, compared against 200 US employers) documents a **systematic skills-prioritization gap** between instructors and employers that signals the workforce demands shaping [[higher-ed|higher education]] curricula. Instructors and employers agree on the importance of only one of 26 AI skills (setting realistic expectations for AI-augmented work): instructors prioritize a *critical, responsible-use* orientation (attribution, human accountability, limits of AI), while employers favor *workflow, automation, and human–AI teaming* skills. The report finds only three of 26 skills are taught by half or more instructors — the under-taught categories (workflow redesign, automation, technical integration) are precisely where employer demands diverge most — and that most institutions lack both a consensus on what AI skills look like and an assessment framework for them. For career development, this means graduates' readiness depends on closing a real, measurable gap between what employers value and what curricula teach, not just on adding AI literacy. - **Course policy as a workforce-competency decision.** [[mccorkle-aligned-genai-course-policy-2025|McCorkle's (2025)]] design case makes the trade-off explicit at task level: for each step of a semester project the instructor pairs an emerging [[generative-ai|GenAI]] workforce competency (prompting for objectives, generating images, writing scripts, text-to-speech narration) against the need to assess a foundational skill, and permits AI only where the competency wins — a concrete way to build the workflow skills employers value ([[prompt-engineering]], evaluation of [[llm|LLM]] output) into existing assignments rather than adding a separate AI course. - **Sector-specific readiness frameworks.** [[workforce-readiness-smart-manufacturing-wrl-2026|Workforce readiness for smart manufacturing]] and [[ai-engineering-computing-workforce-grey-literature-2026|the future of the engineering/computing workforce]] translate general employability into [[discipline-specific-aied|discipline-specific]] competency frameworks. - **Workforce transitions.** and [[post-covid-ict-career-aspirations|ICT career aspirations]] examine how students' career intentions shift in response to technological change. - **Theoretical grounding.** The [[kim-ai-anxiety-comprehensive-analysis|AI Anxiety comprehensive analysis]] identifies the **fear of replacement by AI** as a primary driver of AI anxiety — the career dimension this concept addresses head-on. ## Career readiness as a protective and developmental goal A recurring theme is that career development in the AI era should be a **deliberate educational goal**, not an afterthought: building career adapt-abilities, core self-evaluations, [[self-efficacy]], and employer-valued AI skills, while directly addressing the anxiety students feel about AI displacement. This links career development to [[professional-training]], [[self-efficacy]], [[motivation]], and [[anxiety-and-stress]], and positions education as both a skills pipeline and a source of psychological readiness. ## Connections to related concepts Career development and readiness connects to [[professional-training]] (the vocational skills dimension), [[ai-literacy]] (the AI-competence dimension), [[self-efficacy]] and [[motivation]] (the psychological resources that support adaptation), [[anxiety-and-stress]] (career anxiety as a key component), [[higher-ed]] and [[k-12]] (the settings where readiness is built), and [[student-experience]] (career concerns as part of the learner experience). ## Connected Concepts - [[professional-training]] - [[ai-literacy]] - [[self-efficacy]] - [[motivation]] - [[anxiety-and-stress]] - [[higher-ed]] - [[student-experience]] - [[well-being]] ## Connected Articles - [[mccorkle-aligned-genai-course-policy-2025]] — Assessment-vs-workforce-competency trade-offs decided task by task (McCorkle 2025) - [[wang-career-adapt-abilities-ai-anxiety-english-2026]] — career adapt-abilities reduce AI anxiety - [[ai-literacy-career-adaptability-business-2026]] — AI literacy and career adaptability in business education - [[ithaka-sr-ai-skills-college-graduates-2026]] — AI skills for college graduates: instructor and employer priorities - [[workforce-readiness-smart-manufacturing-wrl-2026]] — workforce readiness for smart manufacturing - [[ai-engineering-computing-workforce-grey-literature-2026]] — AI and the future of the engineering/computing workforce - [[post-covid-ict-career-aspirations]] — ICT career aspirations after COVID-19 - [[kim-ai-anxiety-comprehensive-analysis]] — AI anxiety and the fear of replacement - [[duan-ai-anxiety-career-decisions-college-2026]] — AI anxiety impairs career decisions via career adaptability (63.35% mediation) - [[ustun-ai-anxiety-job-finding-anxiety-2026]] — AI anxiety and attitudes predict job-finding anxiety (1,057 students) - [[dag-ai-perceptions-career-anxiety-health-2026]] — AI anxiety predicts job-search anxiety in health sciences (r=0.233, 821 students) --- ## [Lifelong Learning](https://edtechdev.github.io/aied/concepts/lifelong-learning/) > **Lifelong learning and AI** — how AI supports continuous education and skill development beyond formal schooling, and how it reshapes adult and workplace learning. AI can personalize, scaffold, and make learning-on-demand more accessible for adults, while also raising questions about autonomy, self-direction, and who controls the learning process in AI-mediated environments. ## Questions to Consider - Think about the last new skill you learned outside formal schooling — on the job, online, or self-taught. Who decided what to learn, how to learn it, and when you were 'done'? How would AI change that picture? - A central question in the page is autonomy: as AI participates in identifying your learning needs, setting your goals, and even producing content, what remains genuinely self-directed in your own learning? - Formal education treats time as fixed and curricula as sequential. The page proposes 'Learnity graphs' — interconnected units of knowledge learners navigate flexibly across a lifetime. What do you gain, and what do you risk losing, by freeing learning from a fixed sequence? - AI can both support adult learners and change which skills adults must continuously update. If AI keeps reshaping your field, what capabilities would you invest in that no tool can simply give you? - Lifelong learning is often framed as an individual's responsibility. But access across the lifespan is also an equity question. Who gets the time, resources, and support to keep learning as an adult — and how might AI widen or narrow that gap? ## Introduction Lifelong learning refers to continuous education throughout life — upskilling, reskilling, [[professional-training|professional development]], and informal learning beyond formal degrees. AI is increasingly central to this, both as a tool that supports adult learners and as a force that changes the skills adults must continually update. ### How lifelong learning appears in the research - **Design guidelines for adult learning:** [[ai-adult-learning-guidelines-dis2026|Research from the AI-ALOE institute]] synthesizes empirically grounded design guidelines for AI-powered adult-learning technologies, emphasizing practical, context-sensitive support. - **AI-guided skill acquisition:** [[ai-guided-learning-audiovideo-2026|AI-guided audio/video learning]] shows how AI can scaffold the consume–understand–imitate stages of skill acquisition, adapting media to the learner. - **From fixed curricula to adaptive structures:** [[learnity-graphs-lifelong-learning-framework-2026|Learnity graphs]] propose rethinking fixed higher-education curricula as interconnected units of knowledge that learners can navigate flexibly across a lifetime. - **Autonomy and self-direction:** [[andragogy-cognitive-delegation-genai-2026|Andragogy and cognitive delegation]] revisits adult-learning theory under AI-mediated cognitive delegation, asking what remains self-directed when AI participates in identifying needs, setting goals, and producing content. - **Policy and community:** [[ai-lifelong-learning-policy|AI in lifelong-learning policy]] and [[community-centered-ai-education-adults|community-centered AI education]] address the institutional and equity dimensions of adult AI learning. ### Connections Lifelong learning connects to [[adult-learning]] and [[professional-training]] (its primary contexts), [[personalized-learning]] and [[adaptive-learning]] (the AI mechanisms that support it), [[self-regulated-learning]] and [[metacognition]] (the learner processes involved), and [[equity-in-ai-education]] (access across the lifespan). It also intersects with [[educational-development]] when educators themselves are the lifelong learners. ## Connected Concepts - [[self-directed-learning]] - [[adult-learning]] - [[professional-training]] - [[personalized-learning]] - [[adaptive-learning]] - [[self-regulated-learning]] - [[metacognition]] - [[equity-in-ai-education]] - [[higher-ed]] - [[educational-development]] - [[scaffolding]] - [[intelligent-tutoring]] - [[cognitive-offloading]] ## Connected Articles - [[ai-adult-learning-guidelines-dis2026]] — Guidelines for designing AI technologies to support adult learning - [[ai-guided-learning-audiovideo-2026]] — AI-guided learning for skill acquisition - [[learnity-graphs-lifelong-learning-framework-2026]] — Rethinking higher education with Learnity graphs - [[andragogy-cognitive-delegation-genai-2026]] — Andragogy and cognitive delegation in AI-mediated learning - [[ai-lifelong-learning-policy]] — AI in lifelong learning: opportunities and challenges in adult-education policy - [[community-centered-ai-education-adults]] — Co-designing community-centered AI education for adults - [[cognitive-commons-ai-expertise-regeneration]] — The tragedy of the cognitive commons: AI and expertise regeneration - [[lodge-adaptive-capabilities-genai-future-2026]] — Adaptive capabilities for assuring quality learning in a gen AI-integrated future (Lodge et al. 2026) - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning --- ## [Workplace Learning](https://edtechdev.github.io/aied/concepts/professional-training/) > **Workplace learning** — the use of AI for workforce development, corporate learning, and professional skill acquisition. Professional training extends [[ai-education|AI in education]] beyond formal schooling into workplace and [[lifelong-learning|lifelong learning]] contexts. ## Questions to Consider - Think of a skill you learned on the job rather than in a classroom. What made that workplace learning effective, and how might an AI coach replicate or improve it? - The page's [[career-development-and-readiness|Workforce Readiness]] Level framework suggests the highest competency stages are 'gated by industry-embedded experience rather than coursework.' What does that imply for how we should train people—and for the limits of AI [[simulation]]? - Virtual patients and training simulators let professionals practice safely. What kinds of judgment and interpersonal skills might a simulator struggle to capture, no matter how realistic? - The 'Dual Train Problem' is the tension between rapidly changing AI skills and the slower pace of policy and [[curriculum-design|curriculum]]. If you could choose durable competencies to prioritize for learners today, what would they be? - Adult learners balance work and study, often through screens. How might AI-powered professional training both enable and complicate that balancing act—especially around data, trust, and time? ## Introduction ### AI in professional training - **Simulation and practice:** [[adaptive-virtual-patient-psychotherapy-training|Virtual patient training]] and [[astra-atco-training-simulator|ATCO training simulators]] create AI-powered professional practice environments. In [[teacher-education|teacher education]], AI role-play simulation extends this into practice-based teaching: [[zhuang-zhang-chatgpt-math-teacher-education-2026|Student GPT]] simulated a [[misconceptions|misconception]]-holding middle-school math student so preservice teachers could practice diagnosing and remediating student errors, aligning with the "approximations of practice" of practice-based [[teacher-role|teacher]] learning as an affordable complement to costly platforms like TeachLivE. - **Lifelong learning integration:** [[lifelong-learning]] and [[adult-learning]] [[research-methods-aied|research]] connect professional training to continuous education. - **Public sector:** [[ai-adoption-training-public-sector|Public sector AI adoption]] examines training in government contexts. - **Workforce readiness frameworks:** [[workforce-readiness-smart-manufacturing-wrl-2026|Smith et al.]] propose a Workforce Readiness Level (WRL) framework that adapts the Technology Readiness Level scale into nine competency stages scored across four pillars (digital/[[ai-literacy|AI literacy]], cyber-physical fluency, [[human-ai-collaboration|human-machine collaboration]], data-driven decision making), under a "no-thin-pillar" rule. Evidence from smart-manufacturing capstones shows the highest readiness stages are gated by industry-embedded experience rather than coursework — pointing to work-integrated learning as essential to professional AI training. - **[[discipline-specific-aied|Domain-specific]] PD evidence is thin.** A [[li-language-educators-genai-review-2026|systematic review of language educators]] (Li et al. 2026) found only three of 23 studies reported structured [[educational-development|professional development]], yet those that did converged on gains in knowledge, confidence, and identity — evidence that structured, domain-specific training (pairing technical skill with practical wisdom) is both scarce and effective, and that PD should move from awareness-raising and ethics through hands-on tool mastery to co-design of AI-enhanced lessons. - **Teacher educator professional development as model-building:** [[adaptive-ai-model-teacher-educators-2025|Eyal (2025)]] ran a year-long 180-hour course in which 22 higher-education teacher educators co-designed the Adaptive Artificial-Intelligence-Literacy Model, replacing fixed competency ladders with three inter-related axes (context fit, professional needs, dynamic development) and a 20-item reflective self-assessment questionnaire. The design premise is that AI literacy is situational: a pre-service teacher in a resource-limited setting, a subject teacher, and a principal need different competencies, so professional development should target role-specific judgment rather than a standardized rubric. - **What training should target:** the field's measurement base lags the technology it describes. In a systematic review of 33 teacher AI literacy instruments, [[assessing-teachers-ai-literacy-measurement-tools-2026|Zainal, Mohd Matore and Maat (2026)]] found that 29 (87.9%) targeted general AI concepts while only four (12.1%), all published in 2025, addressed generative AI. If instruments track what training is meant to build, that distribution marks generative-AI competence in teaching as the least measured and most urgent target for professional development. - **Workforce forecasting:** [[ai-engineering-computing-workforce-grey-literature-2026|Fletcher et al.]] review U.S. gray literature on AI and the engineering/computing workforce, framing the "Dual Train Problem" (rapid change vs. urgent policy) and recommending that [[higher-ed|higher education]] prioritize durable AI competencies, [[ethics]] and [[governance]], and skill-based credentials aligned with emerging roles (e.g., [[prompt-engineering]], AI auditing, [[educational-policy-ai|AI policy]]) to sustain human-centered work in an automated economy. - **Oral assessment for workplace capability.** A TVET design study addresses a long-standing mismatch between text-heavy assessment and the verbal, situational capabilities that professional qualifications certify, using an [[llm]] to support interactive oral assessment. Across four cohorts the voice format was rated realistic by 21 of 33 learners with no dissenting response on its advantage over a written [[eportfolio|portfolio]], and the system ran fully offline on one laptop for up to 12 simultaneous learners, deleting recordings after 90 days and leaving scoring to assessors ([[ai-supported-oral-assessment-tvet-2026]]). It is a concrete example of AI widening the range of assessable competence in professional-training rather than only automating existing written formats. Its cohorts, though, were Level 3 automotive and engineering classes, so the evidence sits in [[vocational-education|initial vocational provision]] rather than in the workplace upskilling this page covers. - **Institutional conditions, not national context, explain readiness gaps.** A comparative survey of 568 university faculty in Chinese (n = 340) and Kazakhstani (n = 228) faculty development centers ([[faculty-development-centers-genai-training-optimization-2026|Bi, Araily, Lyu & Xiu, 2026]]) found Kazakhstani faculty ahead on all seven [[generative-ai|GenAI]] readiness dimensions at baseline, with the largest gaps in disciplinary transfer, prompt design, and AI-supported assessment. Hierarchical regression dismantled the national explanation: the country difference fell sharply once prior GenAI use and recent training entered the model and became non-significant — an 80% reduction — once institutional support, perceived permission to experiment, [[multilingual-learning|multilingual]] resource access, policy clarity, and risk sensitivity were added. Exposure was unevenly distributed (a clear majority of Kazakhstani faculty had AI training in the past six months, against roughly a quarter of Chinese faculty), and Chinese faculty reported both higher policy clarity and higher risk sensitivity. Structured prompt-task training outperformed conventional GenAI familiarization on post-test prompt design by a wide margin. Faculty readiness, on this evidence, is produced by provision and permission — what a center offers and what it allows — rather than by the national system it sits in. - **Brief training shifts judgments, not intentions.** A three-hour pre-post pilot with 100 German teachers ([[mesenhoeller-teachers-ai-differentiation-acceptance-2026|Mesenhöller and Böhme, 2026]]) found that perceived usefulness (d = .32) and perceived ease of use (d = .25) rose significantly after a short practice-oriented session on AI for differentiation, while behavioral intention did not move from an already high baseline (M = 3.07). The authors read the gap as a sign that acceptance depends on conditions a session cannot supply, such as time, infrastructure and clear institutional rules. Short formats are worth running, but pairing them with those conditions is what turns favorable judgment into use. - **Sustained monitoring, not one-off workshops, for initial training.** A systematic review of 11 studies of AI in initial teacher training for primary mathematics ([[pinto-ai-initial-teacher-training-mathematics-review-2026|Pinto et al., 2026]]) found that nine interventions were a single session or a few sessions embedded in existing courses, that attitudes were usually sampled once after the fact, and that ethics appeared in only three studies. The authors argue AI competency needs to develop across the whole training sequence, from a preparatory phase into the practicum and the early years of practice, with monitoring that follows that progression. - **Exposure and repetition, not demographics, track favorable perceptions.** In a mixed-methods study of 302 Turkish primary mathematics teachers ([[cigerci-primary-teachers-perceptions-ai-mathematics-2026|Ciğerci and Uygun, 2026]]), prior AI training (t = 3.661) and frequency of AI use (F = 41.280) were the variables most consistently associated with positive views, while willingness (M = 3.98) and attitudes (M = 3.82) sat well above personal experience (M = 3.00). The authors treat prior training as the clearest lever schools can act on, which points to repeated, sustained use rather than a single introduction. - **Role rotation as the load-bearing structure.** A design-based study of 62 pre-service educational psychologists in Kazakhstan ([[kenzhebayeva-ai-role-rotation-pedagogical-model-2026|Kenzhebayeva et al., 2026]]) rotated students through four professional positions over eight weeks, with generative AI supplying preliminary ideas. The authors argue the value lay in the rotation rather than in the tool, since each role framed the same case differently, and later cycles showed more requests for theoretical justification. They report engagement rather than measured gains, having collected no pre-post competence measures. ### Distinct from academic education Professional training differs from academic education in its focus on applied skills, immediate workplace relevance, and adult learner characteristics. [[adult-learning]] theory and [[adult-learning]] principles inform professional AI training design. Its other boundary is [[vocational-education|vocational education and training]]: VET admits people who do not yet hold the occupation and closes with a trade or technical qualification, so it carries initial occupational preparation and the qualification frameworks that certify it, whereas professional training starts from an existing role — reskilling, continuing professional education or vendor certification — and assumes the competence VET awards. **Expertise regeneration as a training concern.** The Cognitive Commons framework ([[cognitive-commons-ai-expertise-regeneration|Lovett 2026]]) argues that HRD must move beyond organizational reskilling to profession-level stewardship: eliminating entry-level developmental positions in AI-exposed sectors can deplete the shared expertise pool on which all organizations depend, with a time-delayed effect that appears only after 5–20 years. This reframes professional training from individual competency development to collective commons maintenance. ## Connected Concepts - [[lifelong-learning]] - [[adult-learning]] - [[vocational-education]] - [[educational-development]] - [[ai-literacy]] - [[simulation]] - [[higher-ed]] - [[generative-ai]] - [[llm]] - [[adaptive-learning]] - [[personalized-learning]] - [[virtual-and-augmented-reality]] — where immersive practice is most established ## Connected Articles - [[ai-engineering-computing-workforce-grey-literature-2026]] — AI and the Future of the Engineering and Computing Workforce - [[workforce-readiness-smart-manufacturing-wrl-2026]] — Workforce Readiness Level framework for smart manufacturing in the AI era - [[cdpk-pedagogy-benchmark-llms]] — LLM pedagogical-knowledge benchmark (CDPK + SEND) - [[ai-interior-design-malaysia-2026]] - [[crewscaler-ai-upskilling-framework]] - [[ai-coaching-rl-skill-development]] - [[adaptive-virtual-patient-psychotherapy-training]] - [[astra-atco-training-simulator]] - [[ai-adoption-training-public-sector]] - [[genai-pd-ai-pck-learning-gain-2026]] - [[cyberagents-gamified-cybersecurity-learning-2026]] - [[hdr-brachytherapy-agentic-ai-simulation-2026]] - [[residencyrl-clinical-rl-training-2026]] - [[ithaka-sr-ai-skills-college-graduates-2026]] — HiBob AI Skills Framework validated with instructors and employers - [[cognitive-commons-ai-expertise-regeneration]] — The tragedy of the cognitive commons: AI and expertise regeneration - [[reflective-triangle-model-teacher-ai-2026]] — Reflective Triangle Model: AI as cognitive mediator - [[utility-value-intervention-teach-responsibly-genai-2026]] — Utility-value intervention effects in learning to teach responsibly with GenAI (Boos, Eder & Lachner 2026) - [[li-language-educators-genai-review-2026]] — Language educators' practices and development with GenAI - [[zhuang-zhang-chatgpt-math-teacher-education-2026]] - [[faculty-development-centers-genai-training-optimization-2026]] — Comparative survey showing institutional conditions, not national context, explain faculty GenAI readiness gaps (Bi et al. 2026) - [[adaptive-ai-model-teacher-educators-2025]] — Teacher educators co-design an adaptive AI literacy model and a reflective questionnaire (Eyal 2025) - [[assessing-teachers-ai-literacy-measurement-tools-2026]] — Review showing teacher AI literacy instruments lag generative AI, marking the training target (Zainal et al. 2026) - [[mesenhoeller-teachers-ai-differentiation-acceptance-2026]] — Three-hour PD raised German teachers' perceived usefulness and ease of use, but not intention to use AI for differentiation - [[pinto-ai-initial-teacher-training-mathematics-review-2026]] — Review of AI in initial teacher training for primary mathematics: brief tool-focused training is not enough - [[cigerci-primary-teachers-perceptions-ai-mathematics-2026]] — Survey linking prior AI training and frequency of use to Turkish primary teachers' positive perceptions - [[kenzhebayeva-ai-role-rotation-pedagogical-model-2026]] — Role rotation as the structuring mechanism for AI-supported professional preparation --- ## [Technologies](https://edtechdev.github.io/aied/concepts/ai-technologies/) > **Technologies** — the models, architectures, and methods that power AI education systems, and the umbrella concept for the knowledge base's coverage of the technical layer. Where [[pedagogy]] and [[learning-theories]] concern *how teaching and learning happen*, and [[ai-ed-evaluation]] concerns *whether AI works*, this page anchors the *technical* strand: the AI systems ([[llm|large language models]], [[generative-ai|generative AI]], [[multimodal|multimodal models]], [[educational-robotics|robots]]) and the techniques used to build, control, and deploy them ([[prompt-engineering]], [[rag|retrieval-augmented generation]], [[reinforcement-learning]], [[educational-nlp]], [[knowledge-graph|knowledge graphs]], [[agentic-ai|agentic orchestration]]). ## Questions to Consider - You can be an excellent educator without being able to build an LLM — but this page argues your technical choices still shape what AI can and can't do in your classroom. What is one way the underlying technology of an AI tool might quietly change how your students learn, even if you never see the code? - A common assumption is that the model is the whole story — but techniques like retrieval-augmented generation (RAG) and prompt engineering exist precisely to control and ground LLM output. Before reading further, when you ask an AI to 'be more accurate' or 'use this source', what do you think is actually happening under the hood? - RAG is described as a core technique for reducing hallucination and improving safety. Why do you think fetching relevant knowledge to 'ground' an AI's answer would matter more for education than for, say, casual chat — and what could go wrong if that grounding fails? - The page claims that technical choices embody pedagogical assumptions: a tutor built on Socratic prompting reasons with learners, while an answer-generating model may just hand over solutions. Can you recall an AI tool you've used that seemed to 'assume' a particular teaching philosophy — and did that align with how you actually wanted to teach or learn? - Beyond raw accuracy, this page suggests AI systems should be evaluated on reliability, pedagogy, and equity. What headline metric do you suspect most people (including many educators) default to when judging whether an AI tool 'works', and why might that metric hide more than it reveals? - Agentic AI is described as shifting AI 'from a prompt-responding tool into a proactive collaborator.' How might a system that initiates and orchestrates multi-step workflows on its own change what you, as an instructor or learner, are responsible for — and who holds it accountable? ## Introduction [[ai-education|AI in education]] runs on a specific technical stack, and understanding it matters for [[teacher-role|educators]] and researchers even when they do not build systems themselves — because technical choices shape what AI can and cannot do in the classroom, the risks it carries, and how to evaluate it. This page organizes the knowledge base's technical-concept coverage: the AI systems, the techniques that adapt and control them, and how the technical layer connects to pedagogy, assessment, and evaluation. ## AI systems in education - **Large language models (LLMs).** The computational backbone of most modern [[ai-education|AIED]] — [[llm|LLMs]] generate human-like text for tutoring, assessment, and content generation, and are the most-referenced technology in the knowledge base. [[pedagogical-llm-training|Pedagogical training]] adapts general LLMs for educational use. - **Generative AI.** The broader category of systems that produce text, code, images, and other content — [[generative-ai|generative AI]] (driven chiefly by LLMs) is the technology behind the current wave of [[ai-education|AIED]] research. See also [[multimodal|multimodal models]] (text, image, audio) and [[simulation]]. - **Robots and embodied systems.** [[educational-robotics|Robots in education]] add an embodied and often social presence — programmable kits for computational thinking and humanoid/social robots for tutoring, storytelling, and role-play. Robotics is a distinct technical strand that overlaps [[agentic-ai|agentic AI]] and [[human-in-the-loop-ai|human-in-the-loop]] design. - **Knowledge-based systems.** [[knowledge-graph|Knowledge graphs]] and [[educational-nlp|educational NLP]] represent and process domain knowledge, increasingly combined with LLMs for grounded, explainable tutoring. ## Techniques and methods - **Prompt engineering.** [[prompt-engineering|Prompt engineering]] is how educators and developers shape LLM outputs — the primary mechanism through which offloading and control are enacted in LLM interactions. - **Retrieval-augmented generation (RAG).** [[rag|RAG]] grounds LLM outputs in retrieved knowledge, reducing hallucination and improving accuracy — a core technique for [[pedagogical-safety|safe]] educational deployment. - **Reinforcement learning.** [[reinforcement-learning|Reinforcement learning]] trains agents to optimize behavior over time, used in [[adaptive-learning|adaptive systems]] and [[game-based-learning|game-based learning]]. - **Agentic orchestration.** [[agentic-ai|Agentic AI]] systems plan and execute multi-step workflows — often orchestrating multiple specialized agents (see [[agentic-ai|multi-agent systems]]) — and are reshaping AI from a prompt-responding tool into a proactive collaborator. - **Model training and adaptation.** [[pedagogical-llm-training|Training and fine-tuning LLMs for pedagogy]], [[educational-llm-alignment|educational alignment]], and [[cstutorbench-slm-tutors|small-language-model adaptation]] make general models education-specific. ## How the technical layer connects to the field The technical strand is inseparable from the knowledge base's other themes: - **Pedagogy:** technical choices embody pedagogical assumptions — a [[intelligent-tutoring|tutor]] built on [[socratic-method|Socratic prompting]] reasons with learners, while an answer-generating model may default to direct provision (see [[pedagogy|pedagogies and teaching strategies]]). - **Assessment and evaluation:** [[ai-ed-evaluation]] and [[benchmark|benchmarks]] determine whether AI systems actually work; [[assessment]] and [[automated-assessment]] use the technical stack to grade and generate. - **Responsible use:** technical techniques are central to [[reducing-ai-misuse|reducing AI misuse]] — [[rag]] grounding, guardrails, [[prompt-engineering]] [[scaffolding]], and [[human-in-the-loop-ai|human oversight]] shape whether AI supports or undermines learning ([[cognitive-offloading]], [[hallucination-risk]]). ## Implications for AI in education - **Technical literacy supports critical use:** understanding the underlying models and techniques helps educators and learners use AI well and evaluate it critically (see [[ai-literacy]]). - **Choose technology by pedagogical intent:** the AI system and technique should follow the [[pedagogy|teaching strategy]], not the reverse. - **Evaluate the technical layer:** [[ai-ed-evaluation]] and [[benchmark]] research assess AI systems on reliability, pedagogy, and [[equity-in-ai-education|equity]], not just headline accuracy. - **Robots and agents are part of the stack:** [[educational-robotics|embodied]] and [[agentic-ai|agentic]] systems extend the technical repertoire beyond text — and bring their own design and safety considerations. ## Connected Concepts - [[llm]] - [[generative-ai]] - [[multimodal]] - [[reinforcement-learning]] - [[educational-nlp]] - [[knowledge-graph]] - [[simulation]] - [[educational-robotics]] - [[agentic-ai]] - [[prompt-engineering]] - [[vibe-coding]] - [[rag]] - [[pedagogical-llm-training]] - [[ai-ed-evaluation]] - [[benchmark]] - [[pedagogy]] - [[learning-theories]] - [[ai-literacy]] - [[adaptive-learning]] - [[personalized-learning]] ## Connected Articles - [[typology-generative-ai-tools-education-2026]] — Typology of Generative AI Tools for Education - [[agentic-ai-education-scoping-review]] — Scoping review of agentic AI in education - [[genai-meta-analysis-programming-learning]] — Meta-analysis of GenAI's effect on productivity and learning in programming - [[cstutorbench-slm-tutors]] — Small language model tutoring benchmarks - [[educational-llm-alignment]] — Aligning LLMs for education - [[eduguard-safe-rag-llm-tutor]] — Guardrailing RAG-based LLM tutors - [[hazra-safetutors-pedagogical-safety-2026]] — AI tutor safety and harms - [[elbench-education-llm-benchmark-2026]] — Education LLM benchmark - [[knowledge-based-design-generative-social-robots-2026]] — Knowledge-based design for generative social robots - [[teachy-mini-generative-social-robot-higher-ed-2026]] — Teachy Mini generative social robot - [[white-wu-robotics-ai-education-2026]] — Robotics and AI in education - [[benzion-ai-physics-simulations-virtual-lab]] — LLM-generated physics simulations for the classroom - [[teo-ai-adoption-tertiary-meta-analysis-2026]] — Factors in adopting AI tools --- ## [Machine Learning](https://edtechdev.github.io/aied/concepts/machine-learning/) > **Machine learning** — the technical foundation of AI in education: algorithms that infer patterns, predictions, and policies from educational data rather than from hand-coded rules. It spans supervised learning (classifying at-risk students, predicting grades), unsupervised learning (discovering learner clusters), reinforcement learning (inducing tutoring and scaffolding policies), and deep learning (neural models for sequences, courses, and visual behavior). [[generative-ai|Generative AI]] is the latest and most visible subset, but it sits atop a much older stack of predictive and adaptive machinery. ## Questions to Consider - You've likely encountered recommendation systems or risk scores that 'learned' from data. How comfortable are you with a model deciding something about you (like an at-risk flag or a recommended path) based purely on patterns in historical data? - The page draws a sharp line between predicting a risk and actually intervening — a risk score tells you a student might fail, but not what to do about it. Where have you seen a prediction offered as if it were already a solution? - Machine learning can 'hack' its own reward: an AI tutor that optimizes for engagement can keep students entertained while [[teacher-role|teaching]] them little. If a system looks measurably successful but is pedagogically harmful, what measures should we watch besides the number it optimizes? - Predictive models trained on historical grades can encode systemic bias, and automated proctoring raises false-positive risks that wrongly flag normal behavior. When data carries the biases of the past, how much should we trust an AI that uses it to make high-stakes educational decisions? - Some AI systems outperform human-interpretable methods but remain opaque — you can see they work but not why. In education, when is interpretability a 'nice to have' and when is it non-negotiable? - Generative AI is often treated as entirely new, but the page frames it as the latest subset of machine learning. How does seeing ChatGPT as part of the same predictive and adaptive machinery change what risks and limits you'd expect it to inherit? ## Introduction Machine learning is what turns educational traces into actionable intelligence. It powers the [[student-modeling|student models]] behind [[adaptive-learning|adaptive systems]], the [[learning-analytics]] dashboards that flag at-risk learners, the [[intelligent-tutoring|intelligent tutors]] that choose the next problem or scaffold, and the automated proctoring systems that monitor remote exams. Across the articles synthesized here, the pattern is consistent: collect data on learners, learn a predictive or decision model from it, and act on that model — whether the action is early-warning, course recommendation, [[scaffolding|adaptive scaffolding]], or exam surveillance. ## What machine learning does in education **Predictive modeling for student success.** Supervised classifiers — Logistic Regression, Random Forest, SVM, K-Nearest Neighbors — identify [[at-risk-students-ml-prediction|at-risk students]] before they withdraw, using [[learning-gains|academic performance]], demographic, and enrollment records. More advanced architectures go further: the [[trace-course-grade-prediction-2026|TRACE]] transformer jointly predicts the courses a student will take and the grades they will receive next semester, modeling the concurrency of co-taken courses rather than treating history as a flat sequence. Optimizer-plus-sequence hybrids, such as the DMO-GRU framework in [[interactive-online-learning-ai-2025|interactive online learning]], combine automatic feature selection and hyperparameter tuning with recurrent nets to reach high accuracy and low error on [[student-engagement|engagement]] and performance prediction. A shared ambition is [[precision-education-student-digital-twins-2026|"precision education"]]: continuous risk stratification and student digital twins that anticipate failure and align pathways with outcomes, shifting institutions from reactive remediation to preventive support. Glass-box models can hold their own here: [[zhang-ml-student-progress-programming-2026|Zhang, Jeffries & Koprinska (2025)]] show that intrinsically interpretable decision trees — pruned to just 3–5 leaf nodes by feature selection — predict module-level student progress in large-scale online [[cs-education|programming]] courses as accurately as black-box random forests and SVMs (85–91% accuracy), evidence that interpretability need not be traded away for predictive power in early-warning applications. Tree-based models also prove effective for estimating item difficulty: [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] fed LLM-extracted cognitive and linguistic features into random forests and gradient boosting machines to predict the difficulty of K-5 math and reading items (N = 5170), reaching correlations up to r = 0.87 with lower RMSE/MAE than direct LLM estimates, dummy regressors, TF-IDF baselines, and metadata-only models. The tree-based models also yielded interpretable [[explainable-ai|feature importance]] — grade level and word count were top predictors — showing how structured features plus interpretable learners can outperform a single holistic judgment. [[multimodal]] fusion extends this predictive toolkit: Bird (2026) fuses a fine-tuned ELECTRA transformer with a searched deep neural network over computational-linguistics features to classify English literature by UK Key Stage at an F1 of 0.996 — far above the best unimodal transformer (BERT, 0.75) and the best linguistic-feature network (0.392) — a concrete case where combining model families outperforms any single approach. Sukoon ([[culturally-aware-student-stress-chatbot-2026|Bashir & Afzal, 2026]]) is a compact applied case of supervised classification in education-adjacent [[well-being]] support: a Random Forest (100 estimators) over 20 features from the psychological, physiological, environmental, academic and social dimensions of a 1,100-response student stress survey, trained on a stratified 70/15/15 split with a scaler fitted on training data only to avoid leakage, and benchmarked against an SVM on the same splits (89.09% vs. 88.48% accuracy). Two touches are worth noting: the authors interpret the confusion matrix by *error direction* — the model's main mistake was classifying low stress as moderate, which in a wellness deployment routes a student to more support rather than less — and they use native feature importances as a substantive finding (blood pressure 15.6%, teacher-student relationship 10.0%), while conceding that a single stratified split, without k-fold intervals, and a non-representative dataset limit the confidence of those weights. **Adaptive instruction and tutoring.** Machine learning closes the loop between modeling and teaching. In [[adaptive-scaffolding-cognitive-engagement-its|intelligent tutoring systems]], both a [[knowledge-tracing|Bayesian Knowledge Tracing]] heuristic and a deep [[reinforcement-learning]] policy adaptively selected worked-example types to elicit different levels of cognitive engagement, significantly improving posttest performance relative to non-adaptive control — while the two policies diverged by learner [[prior-knowledge|prior knowledge]], raising interpretability questions. Reinforcement learning also underlies the [[pedagogical-safety-rl|pedagogical-safety]] agenda: because an RL tutor optimizes a proxy reward, it can "hack" that reward (boosting engagement while teaching little), and architectural constraints on prerequisite enforcement and minimum cognitive demand are needed to keep it safe. Lighter adaptive programs, such as the rule-based [[stem-education|STEM]] system in [[bin-bakheet-adaptive-ai-stem-deep-learning-2026|sixth-grade science]], show that machine learning can be used sparingly — for monitoring rather than direct trajectory control — while still supporting deep learning. On the [[knowledge-tracing]] side, [[pradeesh-outcome-knowledge-tracing-affinity-2026|Pradeesh et al. (2026)]] show a recurrent model can be tuned to a [[curriculum-design|curriculum]]'s own structure: their OKT model treats Outcome-Based-Education course outcomes as knowledge concepts, couples expert-validated OBE affinity mappings (relations between course and program outcomes) with a Memory Augmented Neural Network for cross-outcome impact, and pairs a GRU backbone with domain-adaptive BERT embeddings — reaching 89.81% AUC and beating DKT, DKVMN, EKT, and SimpleKT on live engineering-program data. **Diagnosis from tabular numerical answers.** [[yin-arthur-ai-teaching-assistant-engineering-econ-2026|Yin et al. (2026)]] show gradient-boosted trees solving the *diagnosis* side of tutoring in a domain long closed to AI: for Engineering Economics Calculated Formula Questions, where students' handwritten solutions lack structured digital data, they train a dedicated [[intelligent-tutoring|XGBoost]] multi-label backbone per question to map submitted intermediate and final numerical answers to instructor rubric mistake labels (average precision 0.81, recall 0.79, accuracy 0.65). Random-masking data augmentation — masking input features to "NaN" at varying probabilities — preserves the logical dependencies of tabular solutions better than interpolation methods like SMOTE and significantly improves diagnosis, and including intermediate answers helps most. It is evidence that tree-based models plus structure-aware augmentation can bootstrap tutoring feedback where generative methods lack training data. A field-level view of the [[reinforcement-learning|RL]] subset is provided by the [[riedmann-reinforcement-learning-education-review-2026|Riedmann, Schaper & Lugrin (2025) systematic review]] of 89 RL-in-education studies: it confirms RL is a major ML lineage for inducing adaptive tutoring and scaffolding policies, but finds classical (Q-learning) methods more consistently effective than Deep RL (61% vs 36% significant superiority) and flags that over half of studies conduct no statistical testing — a [[research-methods-aied|methodological]] caution that echoes the validation concerns below. **From prediction to action and accountability.** A persistent limitation is the gap between a risk score and a feasible intervention. The [[sc2r-counterfactual-recourse-educational-2026|SC2R]] framework couples a calibrated predictive model with integer-programming recourse generation and semantic validation, producing intervention plans that respect timing, budget, and availability constraints rather than merely being model-valid. This move "beyond prediction" toward constraint-respecting, machine-checkable recommendations is the field's response to the charge that predictive [[ai-education|AI in education]] can recommend actions institutions cannot or should not take. ## Machine learning in proctoring A distinct application is automated exam proctoring. Deep-learning systems — CNNs and RNNs/LSTMs analyzing eye movements, head posture, and facial expressions — detect cheating more reliably than traditional monitoring, but the [[automated-online-exam-proctoring-decade-review-2026|decade-long systematic review]] finds persistent dataset limitations, single-model evaluations, reproducibility gaps, and false-positive risks that can wrongly flag normal behavior. The companion [[academic-dishonesty-automated-proctoring-ai-2026|review of academic dishonesty]] documents the cheating methods such systems must counter (identity spoofing, browser use, copy-paste) and the practical burdens — cost, connectivity, test-taker anxiety — that bound their equity. Together they caution that ML proctoring must be paired with [[privacy]]-preserving, context-aware design and accessible alternatives. ## Limits and ethical concerns The synthesized literature is candid about machine learning's limits. Gains often come from [[benchmark]] or single-institution data and may not generalize without retraining; [[sc2r-counterfactual-recourse-educational-2026|SC2R]] and [[trace-course-grade-prediction-2026|TRACE]] both acknowledge this. Predictive models trained on historical grading can encode systemic bias, and risk scores applied to sensitive student data raise [[privacy]] and [[equity-in-ai-education|equity]] concerns that demand [[governance]] and [[human-in-the-loop-ai|human oversight]]. Interpretability is a recurring tension — [[reinforcement-learning|RL]] policies can outperform interpretable heuristics while remaining opaque. And [[pedagogical-safety-rl|reward hacking]] shows that an ML system can be measurably successful while pedagogically harmful, which is why architectural safety constraints and audit infrastructure matter. Validation practice is another source of overconfidence. [[schuetze-knowledge-tracing-forgetting-2026|Schuetze, Yan, and Carvalho (2025)]] show that knowledge-tracing models (BKT, BKT-with-Forgetting, AFM) appear to capture learning trends when fit retroactively to all available sessions, yet under **time-based (walk-forward) cross-validation** — training on earlier sessions to predict later ones, mirroring real deployment — they overestimate future performance, miss the spacing effect, and mis-order practice conditions; forgetting-free models even matched forgetting-augmented ones, indicating forgetting was absorbed into other parameters rather than learned. When the training and deployment distributions are temporally separated — the norm for longitudinal student data — a strong in-sample or retrospective fit is no guarantee of predictive validity. ## Generative AI as a subset [[generative-ai|Generative AI]] — large language models and related [[llm|LLM]] systems — is best understood as a subset of machine learning: the same neural and training foundations, but applied to *generating* content (explanations, feedback, dialogue) rather than classifying or predicting. It inherits the field's validity, bias, and safety concerns while adding new ones such as hallucination. For the purposes of this knowledge base, machine learning is the broader technical umbrella; generative AI is its most visible contemporary branch. ## Teacher education and machine-learning literacy Machine learning also appears in education as a *subject*. In [[microbit-robotics-machine-learning-teacher-training-2026|initial teacher training]], hands-on coding and robotics interventions using the Micro:bit and supervised image-classification projects significantly improved preservice teachers' knowledge of computational concepts and introductory machine learning, and their attitudes toward teaching it. As [[ai-literacy|AI literacy]] enters curricula, equipping [[teacher-education|teachers]] with a working grasp of machine learning becomes a precondition for teaching it to students. ## Connected Concepts - [[student-modeling]] - [[learning-analytics]] - [[intelligent-tutoring]] - [[adaptive-learning]] - [[personalized-learning]] - [[reinforcement-learning]] - [[generative-ai]] - [[llm]] ## Connected Articles - [[at-risk-students-ml-prediction]] — Supervised ML classification to identify students at risk of withdrawal - [[trace-course-grade-prediction-2026]] — Transformer jointly predicting courses and grades (TRACE) - [[precision-education-student-digital-twins-2026]] — AI-powered student digital twins for preventive, career-aligned pathways - [[sc2r-counterfactual-recourse-educational-2026]] — Semantics-constrained counterfactual recourse for actionable intervention - [[interactive-online-learning-ai-2025]] — DMO-GRU hybrid for interactive online-learning prediction - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive scaffolding of cognitive engagement in an ITS (BKT vs DRL) - [[pedagogical-safety-rl]] — Formal framework for pedagogical safety in educational reinforcement learning - [[automated-online-exam-proctoring-decade-review-2026]] — Decade-long review of deep-learning automated proctoring - [[pradeesh-outcome-knowledge-tracing-affinity-2026]] — Outcome-based knowledge tracing with affinity mapping - [[bird-multimodal-educational-literature-2026]] — Multimodal fusion for classifying educational literature - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[schuetze-knowledge-tracing-forgetting-2026]] - [[zhang-ml-student-progress-programming-2026]] - [[riedmann-reinforcement-learning-education-review-2026]] - [[yin-arthur-ai-teaching-assistant-engineering-econ-2026]] - [[culturally-aware-student-stress-chatbot-2026]] — An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning --- ## [Generative AI](https://edtechdev.github.io/aied/concepts/generative-ai/) > **Generative AI** — AI systems capable of producing text, code, images, and other content, most prominently large language models like GPT-4 and Claude. Generative AI is the technology driving the current wave of [[ai-education|AI in education]] [[research-methods-aied|research]]. ## Questions to Consider - Generative AI produces fluent, confident-sounding content on demand. Does fluency equal correctness, and where have you seen a confident-sounding but wrong output — what made it hard to catch? - Unlike earlier rule-based or retrieval-based systems, generative models create new content rather than retrieving stored answers. How does that shift change the risks — hallucination, over-reliance, academic integrity — compared to a search engine? - With 80+ articles, generative AI is the largest thread in this knowledge base, spanning tutoring, assessment, content generation, and safety. Which application do you think is the most promising for learning, and which the most dangerous — and why? - The same technology that can generate a Socratic tutorial can also produce a 'correct-answer trap' that encourages copying. What design choices might separate generative AI that scaffolds learning from generative AI that short-circuits it? ## Introduction ### What makes generative AI different for education Unlike earlier rule-based or retrieval-based systems, generative AI produces fluent, contextually appropriate content on demand. This creates both unprecedented opportunities and novel risks: - **Content generation:** [[llm|LLMs]] can create instructional materials, examples, and explanations. [[book-level-synthetic-textbook-organization|Synthetic textbooks]], [[courseblueprint-adaptive-video-generation|adaptive videos]], and [[ai-generated-instructional-videos-computing-ed|instructional videos]] show the range of educational content generation. - **Lesson-planning drafts that are platform- and language-dependent.** Expert [[ai-ed-evaluation|evaluation of AI]]-generated science lesson plans shows content quality is neither uniform nor neutral. [[karaismailoglu-ai-lesson-plans-science-experts-2026|Karaismailoglu, Surmeli and Yildirim (2026)]] had eleven [[science-education]] specialists score ChatGPT-4 and an education-focused tool (Teacher's Buddy) against sixth-grade Engineering Design-Based Learning stages: the education-focused platform outscored the general-purpose one across all eight quality criteria, yet 7 of 11 experts still rated the plans only "applicable by correction." Both platforms generated pedagogically richer output from English than Turkish prompts even when asked to localize for Mersin, Turkey — a [[digital-divide|digital-equity]] concern where prompt language shapes instructional quality. - **Tutoring and dialogue:** [[intelligent-tutoring|AI tutoring systems]] use generative AI for conversational instruction. [[socratic-method|Socratic dialogue]] and [[golrang-propact-pair-programming-2026|collaborative tutoring]] exploit generative capabilities for [[pedagogy|pedagogical]] interaction. - **Simulated patients and case consistency.** A multi-expert annotated corpus of 4,815 student-AI messages from the MeduAI-SP platform ([[ai-standardized-patient-scaffolding-medical-2026|Yang et al., 2026]]) found that only about 0.68% of LLM-generated standardized-patient responses contained clear fidelity problems, and progressive disclosure was rated clinically appropriate in roughly 99.3% of patient messages. This supports the claim that generative-AI simulated patients can sustain case consistency and inquiry-dependent, non-premature disclosure under structured YAML scripting (qwen-max), making them a stable-enough environment for outcome research rather than only for plausibility demonstrations — while the system deliberately withheld diagnoses and [[summative-assessment|summative]] scores during learning. - **Assessment:** [[automated-essay-scoring|Essay scoring]], [[automated-assessment|automated grading]], and [[formative-assessment]] increasingly rely on generative models. [[benchmark|Benchmarks]] substantiate this shift for open-ended work: [[pecuchova-automated-grading-open-ended-genai-2026|Pecuchova, Benko & Drlik (2025)]] found that context-sensitive GenAI models (GPTo1 reaching almost-perfect agreement with human graders) sharply outperformed earlier sentence-embedding approaches on grading open-ended student responses, which relied on rigid reference matching and misclassified valid but differently-worded answers. [[olvet-genai-scoring-open-ended-medical-2026|Olvet et al. (2026)]] extend this to pre-clerkship [[medical-education|medical]] education, where GPT-4's scoring of open-ended questions reached substantial-to-almost-perfect inter-rater agreement with faculty (weighted kappa up to 0.94) — but only after humans iteratively refined the rubric across three rounds and remained in the loop to arbitrate discrepancies — while the most synthetic, holistic-rubric question stalled at moderate (κw = 0.54). This is evidence that generative assessment reliability is shaped as much by human rubric engineering and error-pattern analysis as by the raw model. Yet the same fluency does not generalize across item types: [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat et al. (2026)]] found ChatGPT-5 matched human faculty on objective pharmacy-exam items (CCC 0.935–1.000) but not on short-answer (≈0) or essay (0.341–0.854) items, and a structured rubric did not reliably close the gap. - **Risks:** [[hallucination-risk|Hallucination]], [[cognitive-offloading|Over-Reliance]], [[cognitive-offloading]], and [[academic-integrity]] concerns arise specifically from generative AI's fluency and [[accessibility]]. - **Learning environment generation:** Specialized generative models now turn a course brief directly into finished learning artifacts. [[cogevol-learning-environment-generation-2026|CogEvol (Tu et al. 2026)]], a family of models trained for single-pass generation of structured slides and self-contained interactive HTML pages, completes a slide in a median of 17 seconds and an interactive page in 59 — replacing minutes-long multi-turn [[agentic-ai|agent]] [[scaffolding]]. Reliability is enforced via a production pipeline that converts real failures into 53,687 verified SFT samples plus a hybrid rule-plus-VLM reward for GRPO-based RL. This positions generative AI as a content authoring engine with implications for [[teacher-role|teacher]] and [[curriculum-design|curriculum]] production workflows, and for evaluating whether AI-generated learning environments are functionally and pedagogically sound rather than merely visually polished. ### The knowledge base's generative AI coverage With 80+ articles, generative AI is the knowledge base's largest technology thread. Research spans effectiveness studies ([[genai-meta-analysis-programming-learning|meta-analyses]]), safety concerns ([[hazra-safetutors-pedagogical-safety-2026|tutor harms]], [[eduguard-safe-rag-llm-tutor|guardrailing]]), and design principles ([[instructional-guidance-genai-learning|instructional guidance]]). Generative UI is the newest capability in this thread: models that emit a working interactive artifact — sliders, manipulable simulations — rather than prose. [[generative-ui-education-learning-interactives-2026|Kovshov et al. (2026)]], a Google Research team, report that off-the-shelf generative UI is not yet pedagogically precise enough for complex constructs, but that decomposing a learning objective into progressive leveled goals and wrapping generation in critique and self-improvement loops yields interactives expert teachers rate as acceptable. Theirs is an orchestration design: teachers state objectives, approve them and select among candidate simulations, so the binding constraint on bespoke [[simulation|interactive learning material]] shifts from production to specification, and [[guardrails|pedagogical guardrails]] are embedded in the generation pipeline rather than left to teacher vigilance afterwards. Beyond these core strands, recent work extends the evidence base across [[governance|institutional]], interactional, and domain contexts. Qin (2026) documents how Lingnan University institutionalized GenAI literacy for all undergraduates as part of a digital liberal-arts transformation. Chang and Li (2026) show that student-AI conversations encode discipline-associated cognitive [[student-engagement|engagement]], with ~62% of prompts reflecting higher-order cognitive demand. Neto and colleagues (2026) [[meta-analysis-systematic-review|systematically review]] GenAI in scenario-based healthcare education, finding [[prompt-engineering|prompt design]] functions as instructional specification but is rarely aligned with instructional frameworks (34.8%) or reported in reproducible detail (34.8%). GenAI also powers role-play simulations of learners for practice-based [[teacher-education|teacher training]]: [[zhuang-zhang-chatgpt-math-teacher-education-2026|Zhuang and Zhang (2025)]] built *Student GPT*, a custom ChatGPT chatbot that simulated a [[k-12|middle school]] student holding common ratio-reasoning [[misconceptions]], giving preservice mathematics teachers affordable, content-specific practice at diagnosing student thinking — evidence that prompt design (a literature-grounded prompt reliably elicited target conceptual errors, 0.98 vs. 0.40) can steer an off-the-shelf generative model into a useful pedagogical persona. A [[li-language-educators-genai-review-2026|systematic review of language educators]] (Li et al. 2026) finds educators value GenAI most for preparatory content work — lesson planning, materials creation, and writing support — while hesitating on live classroom use, with concerns centering on [[academic-integrity|academic integrity]] (plagiarism and [[assessment-validity|assessment validity]]), professional displacement, and technostress; adoption is shaped by professional-identity, pedagogical, technical, institutional, and integrity factors, and competency gaps map to episteme, techne, and phronesis. Content generation likewise reaches beyond [[math-education|mathematics]] into co-designing learning resources with teachers — for example, teacher-AI co-designed [[simulation]] scaffolds for [[stem-education|drone STEM]] learning that preserve pedagogical validity and contextual relevance. In children's STEAM [[arts-design-and-media-education|arts education]], [[luo-tahir-chatgpt-steam-lesson-planning-2026|Luo and Tahir (2025)]] experimentally quantified the gains of ChatGPT-assisted over teacher-generated lesson plans (expert-rated median 20.5 vs. 17.6, p = .002, large effect) — yet the same study documents that fluent output carries real failure modes for classroom generation: plans that are idealized or impractical for daily teaching, missed child-safety constraints (e.g., suggesting carving knives for [[early-childhood-elementary-ai-education|young children]]), Western-centric cultural bias, and logically flawed or irrelevant image/resource generation. The contribution is a prompt framework (Role–Instructions–End Goal plus a "four points and one line" quality rubric) that turns the reliability question from whether the model can generate into how prompts and evaluation criteria must constrain it for pedagogical use. [[equity-in-ai-education|Equity]]-oriented uses remain underexplored; an all-girls GenAI makerspace initiative in Europe combined two GenAI tools with feminist pedagogy to address persistent gender inequities in computing participation, analyzing girls' GenAI-generated images and stakeholder reflections. Assistive and inclusive applications are a growing strand: [[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]] — a [[qualitative-research|qualitative]] case study of 21 visually impaired undergraduates in Palestine — found GenAI tailors pace, content, and delivery to individual learning profiles, simplifies complex academic texts, and converts content across modalities, with learners viewing it as complementing rather than replacing teachers. - **Generative AI as a [[pedagogical-agent|pedagogical agent]] in elementary critical media literacy.** Demir and Akar (2026) operationalize the 5E instructional model with generative AI tools (ChatGPT for reflective questions and Q&A, Grammarly and Canva AI for content refinement, Padlet for [[peer-assessment|peer feedback]]) embedded phase-by-phase rather than as isolated add-ons, in an 18-hour critical media literacy program for fourth-grade Turkish students aligned to the Turkish Language and Social Studies curricula. The AI-supported group showed large gains in media reading (+3.50), writing (+1.67), and total media literacy (+5.17, all p < .01) with between-group effect sizes of Cohen's *d* = 1.12 (reading), 1.18 (writing), and 1.31 (total literacy), while the control group advanced only modestly. Qualitative analysis surfaced six domains of critical media literacy growth — digital self-protection and [[privacy|data privacy]], purposeful and responsible media use, safe communication and boundary awareness, [[critical-thinking|critical evaluation]] and misinformation awareness, online risk awareness, and media [[ethics]]/digital citizenship — illustrating how generative AI can be designed into a curriculum as a scaffolded pedagogical agent that cultivates critical evaluation rather than short-circuiting it. ### Generative AI in specialized domains: dyslexia support A 2026 interdisciplinary systematic review (Dabaghi, D'Urso & Sciarrone, PRISMA-guided, 2018–2024, n=72) finds that **generative AI is under-utilized in the dyslexia-support domain**. GAI research (all from 2024) clusters into intelligent [[conversational-ai|chatbots]], [[teacher-role|teacher training]] support, and exploratory studies, and is rapidly overtaking classical [[machine-learning|ML]] as the tool of choice — yet rigorous experimentation and real-world validation remain largely absent. The review's future-trends analysis points to GAI-powered personalized materials and real-time adaptive feedback, [[multimodal|multi-modal]] diagnostic models integrating eye-tracking, EEG, and behavioral [[learning-analytics|analytics]], NLP-driven [[intelligent-tutoring|intelligent tutoring systems]] and conversational agents, and educator-facing support tools. This illustrates both the promise of generative AI for content generation and interactive support in a specialized, high-need domain and the risk that its adoption outpaces the evidence base. ## Connected Concepts - [[llm]] — the model class underlying generative AI - [[prompt-engineering]] — how outputs are shaped - [[rag]] — retrieval-augmented grounding - [[ai-literacy]] — the competency needed to use it effectively - [[ai-education]] — the broader field - [[intelligent-tutoring]] — conversational and generative tutoring systems - [[cognitive-offloading]] — the over-reliance risk generative AI amplifies - [[hallucination-risk]] — a core reliability risk of generated content - [[academic-integrity]] — integrity concerns from fluent generation - [[automated-assessment]] — generative models in grading and feedback - [[ai-technologies]] — the umbrella of AI techniques and models - [[higher-ed]] — a primary deployment context - [[k-12]] — a primary deployment context ## Connected Articles - [[typology-generative-ai-tools-education-2026]] — Typology of Generative AI Tools for Education - [[generative-ui-education-learning-interactives-2026]] — Harnessing Generative UI for Education: Tailored Learning Interactives - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[ssail-safe-sound-ai-learning-2026]] — SSAIL: A Design Framework for Safe and Sound AI for Learning - [[generative-ai-k12-teaching-learning-systematic-review-2026]] — Systematic review of generative AI in K-12 teaching and learning (Marzano 2026) - [[genai-higher-education-systematic-review-2026]] — Systematic review of GenAI in higher education - [[conversational-ai-agents-umbrella-review-2026]] — Umbrella review of conversational AI agents in education - [[genai-educational-outcomes-meta-analysis]] — Meta-analysis of GenAI learning outcomes - [[genai-meta-analysis-programming-learning]] — Meta-analysis of GenAI in programming learning - [[zhao-genai-higher-order-thinking-meta-2026]] — GenAI and higher-order thinking meta-analysis - [[genai-performance-vs-learning]] — Performance vs. learning with GenAI - [[generative-ai-reduced-study-time-math]] — Cognitive surrender: study-time decline with GenAI - [[metacognitively-discordant-completion-genai-2026]] — Metacognitive discordance in GenAI completion - [[hazra-safetutors-pedagogical-safety-2026]] — Harms of AI tutoring agents - [[eduguard-safe-rag-llm-tutor]] — Guardrailing a safe RAG LLM tutor - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From substitution to scaffolding: breaking the harm cycle - [[beyond-detection-authentic-assessment-ai-2025]] — Redesigning authentic assessment for an AI-mediated world - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans - [[cogevol-learning-environment-generation-2026]] — CogEvol: Learning Environment Generation - [[ai-digital-transformation-liberal-arts-lingnan-2026]] — Digital transformation of a liberal arts university toward a research-intensive model in the GenAI era (Qin 2026) - [[student-ai-conversations-cognitive-engagement-2026]] — Discipline-associated Bloom-level cognitive engagement in student-AI conversations (Chang & Li 2026) - [[demir-akar-ai-media-literacy-children-2026]] — AI-based critical media literacy program for children - [[khlaif-assistive-genai-visually-impaired-2026]] — Assistive GenAI for visually impaired learners - [[li-language-educators-genai-review-2026]] — Language educators' practices and development with GenAI - [[dabaghi-ai-dyslexia-education-review-2026]] — AI to help people with dyslexia in education - [[luo-tahir-chatgpt-steam-lesson-planning-2026]] - [[zhuang-zhang-chatgpt-math-teacher-education-2026]] - [[pecuchova-automated-grading-open-ended-genai-2026]] - [[karaismailoglu-ai-lesson-plans-science-experts-2026]] - [[falahat-chatgpt-grading-pharmacy-exams-2026]] - [[olvet-genai-scoring-open-ended-medical-2026]] --- ## [Large Language Models (LLMs)](https://edtechdev.github.io/aied/concepts/llm/) > **Large Language Models (LLMs)** — [[machine-learning|neural network]] models trained on vast text corpora that generate human-like text, powering most modern [[ai-education|AI in education]] applications. LLMs are the computational backbone of generative AI tutoring, assessment, and content generation in education. ## Questions to Consider - What do you believe an AI [[conversational-ai|chatbot]] 'knows' when it answers you? The page frames LLMs as generating probable text rather than retrieving verified facts — how does that distinction change how much you would trust a model's explanations? - LLMs are described as the engine behind most modern AI education tools — tutoring, grading, content generation, and even diagnosing what students know. Of these uses, which do you think is most and least appropriate for a probabilistic text generator, and why? - The page reports that three different LLMs produced sharply divergent support plans for the same learning-analytics input, each with different demographic assumptions. If models aren't interchangeable as advisors, what does that mean for an institution that adopts one? - Because LLM output is sensitive to prompts and settings, two people can get very different results from the same model. How should this influence how you — as a learner or designer — phrase requests, and how much you trust a single output? - A key limitation is hallucination — plausible-sounding but ungrounded content. In a tutoring or grading context, what would it take for you to feel confident the model wasn't inventing something, and what safeguards would you demand before letting it assess a real student? ## Introduction ### LLMs as the engine of AIED LLMs are the most-referenced concept in the knowledge base (60+ articles) because they underpin nearly every AI education application: - **Tutoring:** [[intelligent-tutoring|AI tutors]] use LLMs for dialogue, explanation, and [[problem-solving]] guidance. [[pedagogical-llm-training|Pedagogical training]] adapts general LLMs for educational use. - **Assessment:** [[automated-assessment|Grading systems]], [[automated-essay-scoring|essay scoring]], and [[llm-item-difficulty-prediction|item difficulty prediction]] leverage LLM capabilities. [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] show GPT-4o can estimate the difficulty of K-5 math and reading items (N = 5170) calibrated under the Rasch IRT model: zero-shot ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but varied by grade, while a feature-based strategy in which the LLM extracts cognitive and linguistic features for tree-based models reached correlations up to r = 0.87 — evidence that structured feature extraction can outperform a single holistic LLM judgment. Across the aggregate grading literature, a PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 empirical studies (2023–2025) concludes that LLMs match human raters on short, well-structured tasks with detailed rubrics yet cannot fully replace human judgment on complex, open-ended, or subjective work, and that model version is a dominant determinant of grading quality ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). Reliability also varies sharply by item type: [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat et al. (2026)]] found ChatGPT-5 matched faculty closely on objective pharmacy-exam items (CCC 0.935–1.000) but was unreliable on short-answer (CCC ≈0) and essay (0.341–0.854) items, and that providing a rubric did not consistently improve agreement. - **[[multimodal]] reasoning LLMs as graders:** when a multimodal, reasoning-capable LLM (GPT-o4-mini) graded a 296-student handwritten general-[[chemistry-education|chemistry]] exam page-by-page against rubric images, single-run total scores were highly reproducible (ICC(A,1) = 0.967; averaging five runs reached 0.993) and agreed strongly with TA totals (R² = 0.91), yet item-level reliability was sharply format-dependent — textual and reaction-equation answers graded well while drawing and graphing were worse than random (background grids distract AI vision). This shows an LLM grader's [[trust|trustworthiness]] is a function of response format and task, not just raw model capability, and that [[human-in-the-loop-ai|selective deferral]] via confidence filters is needed for high-stakes use ([[cvengros-grading-handwritten-chemistry-ai-2026]]). - **Content:** [[generative-ai|Generative AI]] content creation relies on LLMs. [[automated-question-generation|Question generation]] and [[ai-generated-instructional-videos-computing-ed|video generation]] are LLM-driven. - **Safety:** [[pedagogical-safety]], [[hallucination-risk]], and [[hazra-safetutors-pedagogical-safety-2026]] [[research-methods-aied|research]] examine LLM-specific risks. - **Diagnosis:** [[knowledge-tracing]] and [[cognitive-diagnosis]] increasingly incorporate LLMs for richer [[student-modeling|student modeling]]. Grounding matters enormously for error diagnosis: [[reddig-maclellan-personalized-feedback-llm-2026|Reddig, Arora & MacLellan (2025)]] showed that supplying GPT-4 the tutor interface structure plus Bayesian [[knowledge-tracing]] skill estimates raised logical-error identification from 40% to 81% on factoring (overall error diagnosis ~87.8%), while multi-step problems and responses with several errors remained weak cases and hallucinated "common-[[misconceptions|misconception]]" diagnoses persisted — evidence that an LLM's diagnostic value is as much a function of the structured context and [[student-modeling|learner-model]] signals it receives as of the model itself. - **Assessment model shift (2017–2024):** Morley et al.'s scoping review of auto-marking short-answer [[science-education|science]] questions traces the field's move from fine-tuning smaller [[educational-nlp|BERT]] models (dominant through 2021) toward prompting larger LLMs (GPT-1/2/3.5/4) from roughly 2022 — adopted via [[prompt-engineering]] rather than fine-tuning — with domain-augmented models, rubric-aware prompting, and chain-of-thought lifting accuracy. Yet GPT models were rarely benchmarked against BERT on standard corpora, few auto-markers could explain their marks, and [[bias-mitigation|bias]] was seldom examined, cautions that apply to LLM assessment generally ([[auto-marking-short-answer-science-2026]]). ### Model-specific research The knowledge base covers both general-purpose LLMs (GPT-4, Claude) and education-specific adaptations. [[cstutorbench-slm-tutors|Small language model benchmarks]] compare SLM performance for tutoring. [[educational-llm-alignment|Educational alignment]] research addresses how to make LLMs pedagogically appropriate. A classroom study across three frontier families — [[oppenheimer-llms-collaborative-learning-partners-2026|Oppenheimer, Cash & Connell Pensky (2025)]] — found that ChatGPT, Gemini, or Claude could act as collaborative critique partners for argumentative writing: over a semester of iterative essays, students improved on argument quality, [[prompt-engineering|prompt engineering]], and response-to-[[ai-feedback-quality|AI feedback]] by roughly a full standard deviation each (all p < .001) and engaged deeply (87.8% rebutting LLM claims), positioning general-purpose LLMs as viable [[collaborative-learning|collaborative learning]] partners rather than mere answer generators. A complementary line of work reframes LLMs from static graders into emulators of [[pedagogy|pedagogical]] reasoning. [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] showed that GPT-4, scaffolded with a semantically precise, iteratively co-refined rubric, could approximate human [[evaluative-judgment|evaluative judgment]] in [[design-based-research|design-based learning]]: initial LLM–human agreement was poor (Cronbach's Alpha = 0.393; Kappa −0.06 to 0.18), but iterative rubric refinement raised mean agreement from 54.75% to 81.25% (final Alpha = 0.798, Kappa 0.40–0.55), and K-means clustering of human and LLM score matrices showed highly correlated centroids (r = 0.89). The study positions the rubric as a mediating interface between human pedagogical intent and machine inference — evidence that off-the-shelf LLMs are not interchangeable as evaluators either, and that their assessment behavior is a design outcome shaped by the rubric and prompts they are given. Raw model capability differentiates grading too: [[benchmark|benchmarking]] eleven GenAI and sentence-embedding models on 1,885 open-ended [[automated-assessment|responses]], [[pecuchova-automated-grading-open-ended-genai-2026|Pecuchova, Benko & Drlik (2025)]] found only GPTo1 reached almost-perfect agreement with expert human graders (Fleiss' Kappa 0.82), with Claude3 and PaLM2 slightly behind, while reference-aligned models such as BERT fell far short — showing that frontier-model context-sensitivity matters for reliable open-ended assessment. Model differences also matter for high-stakes downstream uses. [[lopez-pernas-llm-appropriate-student-support-2026|López-Pernas et al. (2026)]] showed that three LLMs produced sharply divergent student-support prescriptions for the same [[learning-analytics|learning-analytic]] input, and each imposed distinct demographic priors on the learner profiles they generated — evidence that off-the-shelf LLMs are not interchangeable as prescriptive advisors. Similarly, [[olvet-genai-scoring-open-ended-medical-2026|Olvet et al. (2026)]] found that GPT-4's scoring of pre-clerkship [[medical-education|medical]] open-ended questions climbed to substantial-to-almost-perfect inter-rater agreement with faculty (weighted kappa up to 0.94) only after three rounds of iterative rubric refinement, and fell to moderate (κw = 0.54) on a holistic-rubric item — reinforcing that rubric design, not raw capability alone, is the decisive lever for LLM scoring reliability. Model-specific behavior also shows up in how LLMs respond to skeptical users: an algorithmic audit queried ten frontier LLMs 500 times each with a rural-Montana [[k-12]] AI-skeptic persona to test whether [[ai-technologies|AI systems]] consulted by skeptical users are predisposed to encourage adoption. Eight of ten acknowledged user concerns then redirected to AI-[[student-engagement|engagement]] framings; composite scores spanned from 3.85 (Claude Sonnet) to 7.52 (Gemini 3.1 Pro Preview), with a cross-family AI scorer panel clearing Cohen's kappa >= 0.70. The pattern was a model-dependent design outcome. Model capability also depends on how models are combined: Bird (2026) fine-tuned eight state-of-the-art transformers (BERT, ELECTRA, RoBERTa, XLNet, ERNIE, ALBERT, DistilBERT, Longformer) to classify English literature by UK Key Stage, finding the best unimodal transformer (BERT) reached only an F1 of 0.75 — while fusing a fine-tuned ELECTRA with a computational-linguistics neural network lifted F1 to 0.996, showing that transformer text classification alone is limited and that fusion with complementary features is where the gains lie. ## Connected Concepts - [[generative-ai]] - [[prompt-engineering]] - [[rag]] - [[hallucination-risk]] - [[pedagogical-safety]] - [[intelligent-tutoring]] - [[automated-assessment]] - [[ai-literacy]] - [[knowledge-tracing]] - [[higher-ed]] - [[scaffolding]] - [[pedagogical-llm-training]] - [[learning-by-teaching]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) ## Connected Articles - [[llm-interaction-depth-task-quality-recall-2026]] — What students ask matters: LLM interaction depth, task quality, and immediate recall (Tsiligkiris 2026) - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[assessing-quality-ai-generated-exams-field-2025]] — Assessing the quality of AI-generated exams: a large-scale field study - [[nspa-neuro-symbolic-pedagogical-alignment-2026]] — Neuro-symbolic pedagogical alignment (NSPA) - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans (Mathew et al. 2026) - [[educational-llm-alignment]] - [[cstutorbench-slm-tutors]] - [[hazra-safetutors-pedagogical-safety-2026]] - [[llm-item-difficulty-prediction]] - [[eduguard-safe-rag-llm-tutor]] - [[llm-difficulty-calibration-programming-exams-2026]] - [[elbench-education-llm-benchmark-2026]] - [[student-llm-interaction-taxonomy-review-2026]] - [[learnlm-improving-gemini-learning]] — LearnLM: pedagogical instruction following - [[teachlm-post-training-llms-education]] — TeachLM: post-training with authentic learning data - [[conversational-ai-agents-umbrella-review-2026]] — Umbrella review of conversational AI agents in education - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[frontier-ai-redirect-skeptical-rural-staff-2026]] — Algorithmic audit: how frontier LLMs redirect skeptical rural K-12 staff - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[auto-marking-short-answer-science-2026]] - [[reddig-maclellan-personalized-feedback-llm-2026]] - [[oppenheimer-llms-collaborative-learning-partners-2026]] - [[pecuchova-automated-grading-open-ended-genai-2026]] - [[cvengros-grading-handwritten-chemistry-ai-2026]] - [[falahat-chatgpt-grading-pharmacy-exams-2026]] - [[olvet-genai-scoring-open-ended-medical-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] --- ## [RAG (Retrieval-Augmented Generation)](https://edtechdev.github.io/aied/concepts/rag/) > **RAG (Retrieval-Augmented Generation)** — an AI architecture that combines information retrieval with text generation, allowing [[llm|LLMs]] to ground responses in external knowledge sources rather than relying solely on training data. In education, RAG addresses hallucination, enables [[curriculum-design|curriculum]]-grounded tutoring, and powers domain-specific [[intelligent-tutoring|AI tutors]]. ## Questions to Consider - You've probably seen an AI [[conversational-ai|chatbot]] confidently state something false. What does 'grounding' a model's response in external documents change about that failure mode, and what new failure modes might it introduce? - RAG retrieves relevant materials and feeds them to the generator. Before you read, what assumptions does this make about the quality of the retrieved content — and about whether the retrieved text is actually the right thing to teach? - The page contrasts RAG with fine-tuning: retrieval grounds responses in up-to-date sources without retraining, while fine-tuning embeds behaviors. If you were building a curriculum-aligned tutor, which approach would you trust for accuracy, and which for teaching style? - RAG is presented as the main answer to hallucination in education. But consider: if the retrieval source itself contains errors, or is outdated, can RAG still hallucinate? Where might the guarantee of 'grounded in verified content' break down in practice? - For a developer or instructor: what does a tutor need to 'know' beyond the textbook content — pedagogy, when to withhold answers, how to probe understanding? Where would RAG alone fail to provide that, and what would you combine it with? ## Introduction ### How RAG is used in education - **Domain-specific retrieval with notation awareness:** [[algorag-rag-theoretical-cs-education-2026|AlgoRAG]] indexes textbooks, 847 lecture slides, 312 solved practice problems, 156 worked proof templates and 89 complexity worksheets for theoretical [[cs-education|computer science]] courses, adding mathematical entity recognition and notation-aware re-ranking; it answered all 179 instructor-authored exam questions within a 240-second timeout (mean 38.0 seconds) but produced BLEU-4 = 0.0000 and a 0.7620 rubric score, which illustrates both the value of the architecture and the limits of the metrics used to judge it. - **Hallucination reduction:** [[eduguard-safe-rag-llm-tutor|EduGuard]] and [[eduzone-llm-safety-k12|EduZone]] use RAG to keep AI tutor responses grounded in verified educational content, reducing [[hallucination-risk]]. - **Curriculum-grounded tutoring:** [[retrieval-augmented-tutoring-algorithm-kite|KITE]] retrieves relevant curriculum materials to inform tutoring responses, ensuring alignment with course content. - **Textbook and materials indexing:** [[book-level-synthetic-textbook-organization|Synthetic textbook organization]] indexes educational content for retrieval. [[structrag-diagram-reasoning-ai-tutoring|StructRAG]] extends retrieval to structured diagrams. - **Training pipeline integration:** [[pedagogical-llm-training|Pedagogical LLM training]] uses RAG to ground tutor training in educational best practices. - **Course-specific academic support:** [[course-specific-rag-help-seeking-higher-ed-2026|Beacon]] retrieves from a single programming module's approved teaching materials to serve students who hesitate to approach a lecturer, and 89% of the 15 evaluating students rated its responses highly aligned with course materials; the design point is that grounding is an institutional answer to the mismatch between general-purpose [[llm|LLMs]] and module-level expectations. - **Ingest-time structure versus query-time retrieval:** [[wiki-llm-indexing-ml-classes-2026|Wright (2026)]] compiled the same DS3001 machine-learning course corpus into seven cross-referenced wiki concept pages carrying source citations, and set it against a tuned vector-RAG baseline of chunked-embedding retrieval. Over 59 human-written questions the compiled wiki out-answered the tuned index (9.95 vs. 9.05 of 10, with a bootstrap CI on the difference excluding zero) and was more often grounded in the material the answerer actually saw (98% vs. 81%), with both gaps roughly tripling on questions that needed material from more than one page (cross-page scores 9.93 vs. 8.14, where RAG's grounded rate fell from 87% to 64%). The grounding gap was not a retrieval failure: only 2 of vector RAG's 11 ungrounded answers were retrieval misses, while the other 9 had the relevant excerpts in context and still added unsupported detail — evidence that structure at ingest constrains elaboration, not just access. ### RAG vs fine-tuning RAG serves a complementary role to [[llm]] fine-tuning — retrieval provides up-to-date, domain-specific grounding without retraining, while fine-tuning embeds [[pedagogy|pedagogical]] behaviors. The knowledge base's research explores both approaches and their combination. ## Connected Concepts - [[llm]] - [[generative-ai]] - [[hallucination-risk]] - [[knowledge-graph]] - [[edtech-platform]] - [[intelligent-tutoring]] - [[pedagogical-llm-training]] - [[pedagogical-safety]] - [[k-12]] - [[higher-ed]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) ## Connected Articles - [[eduguard-safe-rag-llm-tutor]] - [[eduzone-llm-safety-k12]] - [[retrieval-augmented-tutoring-algorithm-kite]] - [[structrag-diagram-reasoning-ai-tutoring]] - [[book-level-synthetic-textbook-organization]] - [[veriforge-narrative-drafting-scaffolding-2026]] - [[pchl-he-framework-genai-content-creation-2026]] - [[conversational-agents-novice-programmers-scoping-2025]] — Scoping review of conversational agents for novice programmers - [[algorag-rag-theoretical-cs-education-2026]] — AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory - [[personalized-educational-video-generation-2026]] — Dynamic Learning Solutions: A System for Personalized Educational Video Generation - [[course-specific-rag-help-seeking-higher-ed-2026]] — Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education - [[wiki-llm-indexing-ml-classes-2026]] — Potential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing --- ## [Prompt Engineering](https://edtechdev.github.io/aied/concepts/prompt-engineering/) > **Prompt engineering** — the practice of designing and refining inputs to large language models to achieve desired outputs. In education, prompt engineering serves dual roles: as a learner skill (students must learn to prompt effectively) and as a system design lever (developers craft prompts that shape [[intelligent-tutoring|AI tutoring]] behavior). ## Questions to Consider - You've likely typed a prompt into an AI tool recently. Now consider this: the way you phrased it isn't neutral — it may reveal how you planned, thought, and allocated your effort. What might your own prompting habits say about how you approach problems? - A study found that users who phrase requests skillfully systematically get better output than those expressing the same intent less adroitly. If you accept that 'prompt privilege' is real, is fair access to AI best fixed by [[teacher-role|teaching]] everyone to prompt better, or by redesigning the system to not demand that skill — and what are the trade-offs of each? - Is prompting a 'trick' to be memorized, or a genuine intellectual skill? One line of [[research-methods-aied|research]] treats it as professional judgment within a discipline (journalism, law, [[medical-education|medicine]]); another treats it as a core of AI literacy. Which view aligns with your own experience of what actually separates good prompts from bad ones? - Well-designed prompts can scaffold student thinking, while poorly used ones can encourage cognitive offloading. Can you recall a moment when an AI answer did the thinking for you? What about the prompt — or your intent — made that happen, and could it have been designed to do the opposite? - Prompting is both a learner skill and a system-design lever: some tutors now automatically route and select prompts for the user. As prompting moves from the user to the system, what do students lose — and what do they gain? - Set a small goal before you read: after learning about prompt engineering, decide on one concrete way you'll change how you write prompts in your own work, and what result you'll check to know it worked. ## Introduction Prompt engineering is central to effective [[generative-ai]] use in education. Unlike traditional programming interfaces, LLMs respond to natural language — but the quality, accuracy, and [[pedagogy|pedagogical]] value of those responses depend heavily on prompt design. Research in this knowledge base reveals that prompting is not a neutral act: it reflects how students think, plan, and allocate cognitive effort. [[miles-prompt-literacy-human-centered-genai-framework-2026|Miles, Haber-Curran and Arar (2026)]] sharpen what the term covers by distinguishing prompt engineering, the technical optimization of inputs for performance, from prompt literacy, the rhetorical, ethical and reflective work of clarifying purpose, reading output critically and revising with stated reasons. Their Prompt Literacy Cycle (Clarify Purpose, Craft the Prompt, Engage with Output, Refine the Prompt, Reflect) and a sample process rubric make the distinction teachable, and they argue that instruction which optimizes outputs alone leaves the ethical and epistemological dimensions of LLM use untouched. ### How prompt engineering appears in the research - **Prompting as cognitive trace:** [[misiejuk-cognitive-offloading-prompting-2026|Misiejuk et al.]] show that prompt patterns reveal [[cognitive-offloading|cognitive offloading]] — high-quality work uses context-rich, polite, and instructional prompts; low-quality work shows reactive disagreement without domain grounding - **Prompting as literacy:** [[tracing-genai-literacy-interaction-patterns|Tracing GenAI literacy]] and [[aaai2026-prompting-literacy-k12|K-12 prompting literacy]] research frame prompting as a core [[ai-literacy]] component - **Prompting as system design:** [[cotal-formative-assessment-scoring-2026|CoTAL]] uses [[human-in-the-loop-ai|human-in-the-loop]] prompt engineering for [[formative-assessment|formative assessment]] scoring; [[choi-anchor-aes-prompting-2025|anchor-based prompting]] improves [[automated-essay-scoring|automated essay scoring]] - **Adaptive prompt routing:** [[learning-to-prompt-adaptive-tutoring|Learning to Prompt]] treats prompt selection as part of the tutoring system itself — subject-aware prompt routing over 14 pedagogical features, where a stochastic router selects the best prompt per conversation. This shifts prompting from a learner skill into an adaptive system-design lever, improving [[student-engagement|engagement]] and efficiency (28.1% vs 19.6% exercise conversion in a real-world A/B test). - **Prompt modalities:** [[voice-text-prompt-problems-computing-education|Voice vs. text input research]] examines whether prompting modality affects [[learning-gains|learning outcomes]] - **Scaffolded prompting:** [[guided-llm-scaffolding-independent-learning|Guided LLM scaffolding]] and [[scaffolding-critical-engagement-genai-minority-students|critical engagement scaffolding]] teach structured prompting as a learning intervention - **Prompt privilege and equity:** [[prompt-privilege-equitable-ai-access-2026|Jin et al.]] show prompting expertise is unevenly distributed — users who phrase requests skillfully systematically get better output than those expressing the same intent less adroitly. Their Prompt Equity Transformer shifts prompt optimization from the user to the AI system, arguing that [[equity-in-ai-education|equitable]] output should be engineered into the model rather than demanded of novices. - **Prompting as [[situated-learning|situated]] professional judgment.** Beyond literacy and system design, prompting can be framed as a *disciplinary practice*. The [[dierickx-taxonomy-llm-tasks-critical-ai-literacy-journalism-2026|Dierickx et al. taxonomy]] for journalism treats task definition and prompting as a form of professional judgment exercised within a domain's epistemic and ethical norms — translating journalistic work into explicit tasks (newsgathering → sensemaking → editing → publication/distribution) makes assumptions, priorities, and [[ethics|ethical considerations]] visible, and turns prompting into a pedagogical tool for critical AI literacy. Its logic transfers to other knowledge-intensive professions (law, medicine, public policy). - **Prompt design as instructional specification.** Neto and colleagues (2026) find in their [[meta-analysis-systematic-review|systematic review]] of GenAI in healthcare education that prompt design functions as a form of instructional specification, encoding the cognitive targets and quality criteria implicit in expert authoring — yet only 34.8% of studies aligned generated content with instructional frameworks and only 34.8% reported prompting in enough detail to reproduce. Looi, Liu, and Sun (2026) further show how prompt architecture can embed pedagogical rules (correctness gates, anti-spoiler boundaries, goodbye gates) to constrain [[llm]] tutoring behavior in procedural domains. - **Rubric-guided and role-aware prompting.** [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] showed that rubric-guided prompting — treating the rubric as a semantic interface between human pedagogical intent and machine inference — drove LLM–human agreement on student design work from 54.75% to 81.25% (Cronbach's Alpha 0.393 → 0.798). Rubrics engineered for LLMs must balance precision and flexibility: too vague invites free interpretation, too rigid reduces the model to pattern-matching. Role-aware prompting — evaluating the same artifact under instructor, peer-reviewer, and grant-reviewer prompts — produced qualitatively distinct, epistemically different feedback, showing that prompt design shapes not just accuracy but the evaluative stance of the output. - **Context-aware prompting for assessment.** Context-aware prompting of pre-trained language models automates the coding of [[collaborative-learning|collaborative problem-solving]] skills from process data, modeling dependencies between behavior codes and fusing cognitive and social abilities. This enables structured CPS analysis at scale and in real time, overcoming the labor intensity of manual coding schemes. - **Role-based templates and quality rubrics for teacher planning.** [[luo-tahir-chatgpt-steam-lesson-planning-2026|Luo and Tahir (2025)]] empirically develop a prompt framework for children's STEAM arts [[curriculum-design|lesson planning]] that pairs a Role (R) – Instructions (I) – End Goal (E) template (adapted from RISEN) with a "four points and one line" optimization rubric — standardized, practical, engaging, and complete, plus an extension dimension. Applying the rubric to critique and refine prompts kept generated plans acceptable to practicing art teachers (mean ratings above 4/5) while exposing recurring gaps ([[personalized-learning|personalization]], [[pedagogical-safety|child-safety]] constraints, cultural bias) that plain one-shot prompting left unaddressed — showing prompt templates plus explicit evaluation criteria function as a quality-control scaffold for classroom generation. - **Role and constraint design as the independent variable.** [[wang-teacher-student-centered-agents-physics-2026|Wang et al. (2026)]] compare two agents built on the same model and platform at temperature 0.3 whose only difference is how the prompt specifies role, skills, and constraints: an expert teacher agent answering from a bounded textbook knowledge source versus an empathic student-centered agent scripted to diagnose [[misconceptions]] and check comprehension. The role difference alone shifted learning performance, cognitive load, flow experience, and perceived empathy, showing that role specification is an instructional-design decision with measurable effects rather than a stylistic flourish ([[pedagogical-agent]]). ### Connections to broader concepts Prompt engineering connects to [[scaffolding]] — well-designed prompts can scaffold student thinking rather than bypass it. It intersects with [[metacognition]] and [[ai-literacy]], as effective prompting requires understanding both the AI's capabilities and one's own learning goals. The [[cognitive-offloading]] research directly links prompt quality to whether AI use supports or undermines learning. - **Writing skill drives prompting, and both predict [[vibe-coding]] success.** In a preregistered CHI 2026 study (N=100), [[vibe-coding-writing-cs-achievement-2026|Thorgeirsson, Weidmann & Su]] found that written-communication proficiency predicted GUI-oriented vibe-coding performance (r = .29), with human-graded prompt quality *mediating* the link — response-process evidence that clear, structured prose translates into better natural-language programming prompts. Both writing skill and [[cs-education|CS achievement]] were independent predictors, and CS achievement (r = .39) carried roughly twice the unique variance, so improving prompting alone is unlikely to fully substitute for programming fundamentals in LLM-native development. - **Prompting strategy predicts performance.** An [[isaza-chatgpt-engineering-prompting-2026|empirical study of 128 engineering students]] found that AI Query Efficiency (clear, well-structured prompts) and AI-Driven [[problem-solving]] (strategic integration of AI output into reasoning) were the strongest predictors of academic success — even after controlling for GPA — indicating prompting is a teachable skill that shapes how effectively students learn with AI. - **A usable taxonomy, and which prompt categories actually pay off.** [[teacher-ai-literacy-prompt-feedback-quality-2026|Jacobsen et al. (2026)]] translate technical strategies into the 3K model (*Kontext, Kernauftrag, Klarheit* — context, core task, clarity): eleven practice-oriented categories, each with a good/average/suboptimal rubric, and each tested as an experimental variation on feedback generated for pre-service teachers' learning goals. Domain-specific technical language was the decisive category — replacing subject terminology with everyday paraphrases significantly reduced feedback quality across three models (β = −0.412) — while adding concrete examples and removing the chain-of-thought instruction produced no significant difference from the baseline in the first study; examples did help once the analysis was rerun with the best-performing model-prompt combinations (β = 0.52). Prompt quality and model choice together explained 42.8% of the variance in rated feedback quality, which is the paper's case that prompt engineering is a measurable and teachable competency rather than a stylistic preference — and that its categories are not interchangeable in effect size. ## Connected Concepts - [[vibe-coding]] - [[guardrails]] - [[scaffolding]] - [[ai-literacy]] - [[agentic-ai]] - [[metacognition]] - [[curriculum-design]] - [[cognitive-offloading]] - [[writing-education]] - [[k-12]] - [[generative-ai]] - [[learning-design]] - [[cs-education]] - [[higher-ed]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) ## Connected Articles - [[wang-teacher-student-centered-agents-physics-2026]] — Agent role and constraint prompts as the design variable in physics learning (Wang et al. 2026) - [[gpt-item-generation-l2-listening-2026]] — Prompting vs. fine-tuning for GPT-based L2 listening item generation (Aryadoust & Wong 2026) - [[llm-interaction-depth-task-quality-recall-2026]] — What students ask matters: LLM interaction depth, task quality, and immediate recall (Tsiligkiris 2026) - [[ye-arpg-real-time-coaching-llm-prompting-2026]] — ARPG+: real-time coaching for educational LLM prompting - [[dierickx-taxonomy-llm-tasks-critical-ai-literacy-journalism-2026]] — Task-based taxonomy of LLM tasks for critical AI literacy in journalism - [[benali-genai-academic-writing-2026]] - [[ying-genai-journalism-assessment-2026]] - [[enright-staff-perspectives-genai-2026]] - [[prompt-privilege-equitable-ai-access-2026]] — Prompt Privilege: measuring & mitigating accessibility disparities in LLM access - [[principal-trait-analysis-human-ai-skills-2026]] — Principal Trait Analysis: data-driven traits of human-AI collaboration - [[llms-text-linguistics-teaching-2026]] — LLMs in text linguistics teaching - [[idea-framework-metacognitive-genai-2026]] — The IDEA framework for metacognitively regulated GenAI use - [[lin-llm-interactive-lesson-generation]] — LLM generation of interactive tutor-training lessons (Lin et al. 2025) - [[aaai2026-prompting-literacy-k12]] - [[ai-adoption-training-public-sector]] - [[ase-26-agentic-software-engineering-curriculum]] - [[choi-anchor-aes-prompting-2025]] - [[guided-llm-scaffolding-independent-learning]] - [[learning-to-prompt-adaptive-tutoring]] - [[llm-intervention-design-cs-review]] - [[misiejuk-cognitive-offloading-prompting-2026]] - [[tracing-genai-literacy-interaction-patterns]] - [[pchl-he-framework-genai-content-creation-2026]] - [[probing-ai-generated-physics-solutions-2026]] - [[genai-assisted-problem-posing-physics-2026]] - [[unesco-ai-guidelines-chemical-education-2026]] — UNESCO AI guidelines translated to chemical education; epistemic drift - [[learnai-just-in-time-ai-cocreation-university-2026]] — LearnAI: Just-in-Time AI Co-Creation Across Disciplines - [[student-ai-inquiry-types-cs2-2026]] — Analysis of Types of Inquiries in Student-AI Interaction - [[learnlm-improving-gemini-learning]] — LearnLM: pedagogical instruction following vs prompt engineering - [[teachlm-post-training-llms-education]] — TeachLM: prompt engineering as a stopgap - [[li-dbagent-llm-educational-agent-cs-2026]] — LLM-based educational agent (DBagent) in CS education - [[pedagogy-ai-mistakes]] — The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking (Hosseini 2026) - [[chatgpt-qiskit-homework-autogradable-2026]] — ChatGPT solves Qiskit homework; autogradable design - [[isaza-chatgpt-engineering-prompting-2026]] — Prompting behaviors predict engineering student performance - [[rule-integrated-llm-tutoring-primary-math-2026]] — Rule-guided vs ad-hoc scaffolding in an LLM tutoring system for primary mathematics (Looi et al. 2026) - [[genai-scenario-based-healthcare-education-2026]] — Systematic review of GenAI in scenario-based healthcare education (Neto et al. 2026) - [[student-ai-conversations-cognitive-engagement-2026]] — Discipline-associated Bloom-level cognitive engagement in student-AI conversations (Chang & Li 2026) - [[context-aware-prompting-cps-skill-identification-2026]] — Context-aware prompting for automated collaborative problem-solving skill coding - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design - [[luo-tahir-chatgpt-steam-lesson-planning-2026]] - [[miles-prompt-literacy-human-centered-genai-framework-2026]] — Prompt engineering vs prompt literacy: a five-phase human-centered GenAI engagement framework with a five-step Prompt Literacy Cycle (Miles, Haber-Curran & Arar 2026) - [[teacher-ai-literacy-prompt-feedback-quality-2026]] — Prompt engineering and model selection as predictors of AI-feedback quality (Jacobsen et al. 2026) - [[context-prompts-physics-assignments-2026]] — Artificial Intelligence Driven Physics Assignments using Context Prompts --- ## [Vibe Coding](https://edtechdev.github.io/aied/concepts/vibe-coding/) > **Vibe coding** — building software by iteratively prompting a large language model and judging the resulting behavior, without directly reading or editing the underlying source code. Popularized by Andrej Karpathy in 2025 as the workflow where one "forgets the code even exists," vibe coding is the LLM-native realization of natural-language programming and end-user development — framings now treated as synonyms in this knowledge base — in which prose becomes the primary programming interface. ## Questions to Consider - Karpathy's original framing said you should "forget the code even exists." Before you read further, ask yourself: is not seeing the code a feature (it lowers barriers) or a risk (you cannot verify or fix what you cannot see)? What does the answer imply for who should be allowed to vibe-code? - Research on who succeeds at vibe coding found that traditional [[cs-education|computer-science achievement]] still predicts success even when the user never touches code. If that surprises you, what hidden skill might CS training be building that prose alone does not capture? - The same study found writing skill predicts vibe-coding performance largely *because* it produces higher-quality prompts. If prompting is really the bottleneck, is the right fix to [[prompt-engineering|teach people to prompt better]] — or to redesign tools so they demand less prose skill? - Vibe coding is often celebrated as making "anyone" a developer. But if writing skill and CS knowledge both shape outcomes, does vibe coding widen access to building software or merely relocate the skill barrier from code to prose? - Some developers distinguish "pure" vibe coding (never reading code) from AI-assisted coding where you review and edit what the model wrote. Where do you think genuine learning — versus [[cognitive-offloading|over-reliance]] — is more likely to happen, and why? ## Introduction Vibe coding describes an interaction style enabled by LLM-integrated development platforms (Replit, Lovable, Cursor, and others): the user specifies a program in natural language, the model generates a working system, and the user iterates based on observed behavior rather than by editing source. The term was coined by OpenAI co-founder Andrej Karpathy in February 2025 to capture the experience of relying on the model to such an extent that "the code" fades from awareness. Vibe coding sits at the convergence of several strands this knowledge base already tracks — it is [[generative-ai|generative AI]] applied to [[cs-education|programming]], an extreme form of [[prompt-engineering|prompt-driven]] work, a concrete instance of [[human-ai-collaboration]], and the clearest route yet to [[teacher-role|non-programmers]] and end users building their own software (end-user development). It also has deep roots. The idea of programming in ordinary language long predates LLMs — from COBOL's aspiration to be "an English-language programming system for non-professional programmers," through Donald Knuth's literate programming, to research on natural-language programming with constrained subsets of English. Only with LLMs did it become feasible to map genuinely conversational, underspecified instructions to runnable code. Vibe coding is the particular variant in which the user deliberately does not inspect or edit the generated source, relying entirely on iterative prompting and behavioral evaluation. ### Defining the construct: "pure" vs. code-visible vibe coding The definition of vibe coding is still in flux. Some use the term broadly to mean any AI-guided programming; others insist it refers strictly to "building software with an LLM without reviewing the code it writes." Google Cloud distinguishes a "pure" no-code version (consistent with Karpathy's definition) from a version where the user understands and refines the generated code. This distinction matters for [[research-methods-aied|research]]: a controlled study of vibe-coding proficiency requires a well-defined construct. The CHI 2026 study of predictors of vibe-coding proficiency deliberately targeted the "pure," no-code variant — participants could not view or edit the generated source, so measured performance reflected the ability to specify, refine, and debug behavior through prose and observed output alone (see [[vibe-coding-writing-cs-achievement-2026|Thorgeirsson et al.]]). ### Who succeeds at vibe coding: evidence A preregistered cross-sectional study (N = 100 tertiary students) provides the first controlled, participant-level evidence on which skills predict vibe-coding success. Both [[writing-education|written-communication proficiency]] (r = .29) and computer-science achievement (r = .39) significantly predicted performance on expert-vetted, GUI-oriented vibe-coding tasks, with CS achievement remaining significant after controlling for domain-general cognitive skills (partial r = .281). In a joint model CS achievement contributed roughly twice the unique variance of writing skill, but both added independent predictive value. Critically, human-graded prompt quality mediated the writing→performance link, giving response-process evidence that clear prose operates by producing better prompts. Because the environment hid the source code, CS knowledge could only help indirectly (through problem decomposition, algorithmic thinking, and mental models of control flow) — so the authors argue their CS estimate is a *lower bound* for AI-assisted programming in which users may also edit code directly ([[vibe-coding-writing-cs-achievement-2026|Thorgeirsson et al., 2026]]). ### Vibe coding as end-user development and teacher tooling A major promise of vibe coding is that it lets non-programmers — including [[teacher-role|teachers]] and domain experts — build their own software, an LLM-era form of end-user development. A [[gaide-vibe-coding-k12-teachers|GAIDE framework study]] showed K-12 teachers (non-programmers) using vibe coding in an eight-week workshop to create AI-powered learning tools, raising their [[ai-literacy|AI literacy]] and demonstrating "learning-by-creating" as a professional-development model. In higher education, an instructor rapidly built a [[vibe-coding-programming-process-visualizer|programming-process visualizer from IDE activity logs]] via vibe coding in a matter of days, making students' programming processes visible for teaching and [[academic-integrity]] review. These cases position vibe coding not merely as a learner skill but as an authoring capability that [[educational-development|reshapes who can create educational technology]]. ### Learning, agency, and the risk of over-reliance Vibe coding reopens core questions about what is learned when AI automates implementation. Because the user does not read code, they must trust the model's behavior — which makes vibe coding a high-stakes case of the tension between [[agency]] and [[cognitive-offloading|over-reliance]] that runs through AI-assisted programming. Curricula are responding by shifting from teaching implementation toward teaching how to direct, verify, and audit AI-generated artifacts (see [[reshaping-cs-education-genai|reshaping undergraduate CS]] and [[agentic-ai|agentic software engineering]]). Vibe coding also changes the learner's epistemic position: success depends less on writing code than on expressing intent precisely and evaluating behavior against goals, competencies closer to [[computational-thinking|computational thinking]] and structured writing than to traditional syntax mastery. ### Connections to related concepts Vibe coding connects naturally to [[prompt-engineering]] (prompt quality is the mechanism of prose-driven development), [[cs-education]] (as the domain where the technique is most used and most contested), [[computational-thinking]] (the mental modeling that predicts success even without code access), [[writing-education]] (writing becoming a programming skill), and [[agentic-ai]] (directing a model toward an artifact rather than hand-building it). It also intersects with [[ai-literacy]] and [[teacher-role]], since the ability to build one's own tools changes what teachers and learners can do. Finally, it raises [[academic-integrity]] and assessment questions identical to those AI code generation raises across computing education. A faculty-level case study in this knowledge base supplies the organizational layer. [[zimmer-ai-intrapreneurship-faculty-innovation-2026|Zimmer (2026)]] describes *AI intrapreneurship* — educators building their own tools instead of waiting for institutional procurement — including an author who does not code using Claude Code to build a checker for 321 course links. The decisive enablers were organizational rather than technical: work discretion, rewards, and time availability, the last described as most obviously in deficit in academic settings and undercut by promotion and tenure. The security picture stayed sober, since Veracode's 2025 analysis found only 55% of AI-generated code secure, so vibe-coded classroom tools still need a review pass before handling student data or connecting to an LMS. ## Connected Concepts - [[generative-ai]] - [[llm]] - [[prompt-engineering]] - [[cs-education]] - [[computational-thinking]] - [[writing-education]] - [[agentic-ai]] - [[human-ai-collaboration]] - [[ai-literacy]] - [[teacher-role]] - [[cognitive-offloading]] ## Connected Articles - [[vibe-coding-writing-cs-achievement-2026]] — Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency (CHI 2026 empirical study) - [[gaide-vibe-coding-k12-teachers]] — A Guiding Framework for K-12 Teachers in Creating AI-powered Learning Technologies through Vibe Coding - [[vibe-coding-programming-process-visualizer]] — From Idea to Classroom in Days: Using "Vibe Coding" to Create a Programming Process Visualizer from IDE Activity Logs - [[prompt-problems-nl-programming-mistakes]] — Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks - [[code-to-learn-genai-artifact-construction-2026]] — Code to Learn with Generative AI: A Theoretically Grounded Framework for Artifact Construction in Upper-Secondary Education - [[reshaping-cs-education-genai]] — Reshaping Undergraduate CS Education for Generative AI - [[flowcode-ai-creative-coding]] — Flowcode: An AI-Powered Programming Environment for Scaffolding Iteration in Creative Computing Education - [[zimmer-ai-intrapreneurship-faculty-innovation-2026]] — AI intrapreneurship: faculty building their own tools, and the organizational enablers that decide whether the impulse survives (Zimmer 2026) --- ## [Multimodal AI](https://edtechdev.github.io/aied/concepts/multimodal/) > **Multimodal AI** — [[ai-technologies|AI systems]] that process, understand, or generate content across multiple modalities — text, images, audio, video, and structured data — and the educational questions these systems raise. In [[ai-education|AI in education]], multimodal AI appears in three distinct roles: as the *learning content* learners create and engage with ([[multimodal-learning-genai|multimodal learning]]), as the *capability boundary* of tutoring systems that must interpret diagrams and graphs ([[syal-multimodal-dialogue-stem-2026|multimodal tutoring]]), and as the *assessment signal* used to evaluate understanding ([[multimodal-item-parameter-estimation-2026|multimodal measurement]]). ## Questions to Consider - Think of a graph, force diagram, or schematic you've ever struggled to explain in words. What does that experience suggest about the limits of a text-only [[intelligent-tutoring|AI tutor]] trying to help with image-rich problems? - A [[physics-education|physics]] tutor answers text-based problems ~96% of the time but drops to ~74% on problems that embed meaning in diagrams. Before reading, what do you think causes this 'multimodal interference'—and can you think of a fix that doesn't involve retraining the model? - You've likely generated both text and images with AI tools. Have you found that '[[prompt-engineering|prompting]] for pictures' differs from prompting for text? What skills might students need to translate an abstract idea into a precise visual prompt? - Multimodal AI can grade essays, generate feedback with audio narration, and even reconstruct exam item statistics from image-and-text items. What does the shift from text-only to multimodal assessment signal (or risk) for [[bias-mitigation|fairness]] and validity? - How could the fact that AI support is less reliable on the very diagram-heavy problems that build deep [[stem-education|STEM]] understanding create an [[equity-in-ai-education|equity]] gap between learners? Who is most affected? - Multimodal systems can translate text to audio or visuals to support inclusive learning, but they also enable fine-grained classroom sensing. Where is the line between helpful multimodal access and surveillance? ## Introduction Multimodality in AI refers to the capacity to work across different representational forms rather than text alone. Modern [[generative-ai]] and [[llm]] systems increasingly accept and produce images, audio, and video in addition to text, opening new possibilities and new risks for education. Grounded in social semiotic theory, which holds that meaning is made across modes — not just words — multimodal AI changes how [[teacher-role|teaching]], learning, and assessment are designed and evaluated.([[multimodal-learning-genai]]) ## Three faces of multimodal AI in education ### 1. Multimodal learning and content creation Multimodal AI enables learners to produce and engage with content across text, image, audio, and video. An educator's guide to multimodal learning with generative AI positions these tools as a "cyber-social" partner: they complement — but cannot replace — human meaning-making.([[multimodal-learning-genai]]) - **[[ai-literacy|AI literacy]] in multimodal contexts** is layered: basic awareness of multimodal platforms, intermediate co-creation and [[critical-thinking|critical evaluation]] of outputs, and advanced design of multimodal activities and assessments.([[multimodal-learning-genai]]) - **Multimodal prompting** is itself a demanding epistemic practice. Students who prompt for images as well as text discover that "prompt literacy is different between prompting for text than it is for pictures" — translating abstract meaning into machine-readable multimodal prompts requires a precise visual vocabulary and exposes system limitations and bias.([[multimodal-prompting-ai-literacy]]) - **Multimodal assessment** shifts from essays to artifacts combining text, image, audio, and video, with educators using AI to [[scaffolding|scaffold]] creation and feedback rather than replace the learner's own production.([[multimodal-learning-genai]]) - **Learner multimodal composing as a critical-thinking scaffold carries a trade-off.** [[lu-ai-multimodal-writing-critical-thinking-2026|Lu et al. (2027)]] show that having upper-primary students turn written narratives into AI-generated images and short videos supported sustained gains in interpretation, analysis, evaluation, and explanation — but not inference. Because the visuals made story meaning explicit, students reported less need to infer implicit meaning from text alone; peer collaboration, not the multimodal tool, restored occasions for inference. Multimodal AI's value as a meaning-making partner is thus dimension-specific and depends on [[learning-design|instructional design]] that deliberately re-introduces the inferential and [[self-regulated-learning|self-regulatory]] work the externalization can short-circuit. - **Learner multimodal [[writing-education|composition]] as critical AI literacy.** [[burriss-multimodal-composition-critical-ai-literacy-2026|Burriss et al. (2026)]] analyze 22 eleventh graders' 90-second to 3-minute video public service announcements on self-chosen AI [[ethics]] issues — surveillance through school-regulated laptops and electronic "hall passes," [[privacy|informed consent]], and punitive algorithmic accusation — as [[ai-literacy|critical AI literacy]] enacted through composing across moving image, sound, text, and students' own bodies. Across all seven films harm was portrayed as emerging from human–machine entanglement rather than from the tool alone (an anthropomorphized "AI stalker" was played by a human actor in three of seven), and 15 of 18 end-of-unit responses said composing changed their understanding of AI ethics. The authors argue multimodal products both *demonstrate* and *communicate* critical competence — productive artifacts, reflections, and civic discourse can serve as [[assessment]] evidence that text-only literacy scales structurally miss. ### 2. Multimodal tutoring and the capability boundary When LLM-based tutors must solve problems that embed meaning in graphs, force diagrams, schematics, or tables, their accuracy degrades sharply — the **Multimodal Interference Effect**.([[syal-multimodal-dialogue-stem-2026]])([[syal-multimodal-dialogue-stem-2026]]) - On OpenStax physics problems, text-only accuracy of ~96% drops to **~74%** on image-rich problems, consistently across model families.([[syal-multimodal-dialogue-stem-2026]]) - **Visual Processing Errors** — failures to extract information from graphs or diagrams — dominate the error taxonomy and are the most correctable failure mode. - A simple structured-dialogue intervention (have the model describe what it sees, correct only *observable* misreadings without giving away physics, then re-prompt) restores accuracy to **~95%** with zero retraining.([[syal-multimodal-dialogue-stem-2026]]) - This is an **equity concern**: students working on image-rich problems — precisely the problems that build deep conceptual understanding in STEM — currently receive less reliable AI support than those on text-only exercises. - **The boundary is a profile, not a level — and artistic imagery sits outside the region models handle well.** [[muse-vlm-artistic-image-benchmark-2026|MUSE (Zhu et al., 2026)]] evaluates 30 open- and proprietary VLMs on 12 tasks over 1,174 commissioned artworks, and the capability spread across dimensions is wider than any aggregate score suggests: scene classification is near-mature (23 of 30 models above 75.0, median 81.0) while emotion detection tops out at 39.5 and the open-ended tasks that require models to *articulate* their evidence score at 50.90 (visual clue identification) and 49.18 (emotion cause inference) on semantic similarity. Compositional and viewpoint-dependent reasoning fail hardest — where the ground truth specifies no definite lateral or vertical relation, 90.0% and 73.3% of models assert one anyway, only 43.3% place the girl correctly in depth, and no model resolves all three dimensions of a single item. Failures also cascade: a mis-grounded character is then justified with a fluent rationale built from nearby visual semantics (butterflies, birds), which is the outcome most dangerous in tutoring because the explanation reads as competent. For image-based [[language-learning|language learning]] this argues for dimension-level validation on the imagery a course actually uses, rather than importing a general multimodal score, and for extending the grounding checkpoint described below — describe what is seen, and where, before reasoning from it — to [[situated-learning|situated]] artistic content ([[muse-vlm-artistic-image-benchmark-2026]]). The practical design implication is a **visual grounding checkpoint** in multimodal tutoring: a deliberate step where the system describes what it sees before attempting a solution, giving the student or a human supervisor a chance to correct perceptual errors.([[syal-multimodal-dialogue-stem-2026]]) ### 3. Multimodal assessment and measurement Multimodal AI broadens both the *content* of assessment and the *signal* used to score it. - **Multimodal feedback systems** integrate structured text, slide references, and streaming audio narration. In one study, [[ai-feedback-quality|AI multimodal feedback]] matched educator feedback for learning while *significantly outperforming* it on student perceptions.([[multimodal-ai-feedback-learning]]) - **Multimodal item response estimation** uses fine-tuned multimodal LLMs to reconstruct item characteristic curves (IRT / 3PL) directly from predicted option probabilities on image-and-text items, connecting multimodal AI to [[educational-measurement]] and [[item-response-theory]].([[multimodal-item-parameter-estimation-2026]]) - **Educational vision-language model evaluation** and [[mllm-scientific-visualization-literacy|multimodal LLM literacy]] extend the field's evaluation toolkit to multimodal reasoning and [[visualization]].([[drawedumath-vlm-struggling-students-2026]])([[mllm-scientific-visualization-literacy]]) - **Multimodal grading of handwritten [[chemistry-education|chemistry]] exposes a format-dependent capability boundary:** [[cvengros-grading-handwritten-chemistry-ai-2026|Cvengros & Kortemeyer]] graded a 296-student handwritten general-chemistry final page-by-page against rubric images with a multimodal, reasoning LLM, scoring textual answers and chemical-reaction equations reliably (normed F1 highest) but drawing and graphing *worse than random* — background grids visually distract AI vision and scientific diagrams/chemical structures remain hard to interpret — reinforcing that multimodal AI's vision is not robust to representation-heavy work and is best deployed with [[human-in-the-loop-ai|human deferral]] of graphical items ([[cvengros-grading-handwritten-chemistry-ai-2026]]). ## Multimodal AI for language and accessible learning Multimodal systems also expand access and [[personalized-learning|personalization]]. AI-guided audio-[[video-education|video learning]] tools adapt playback speed, produce multimodal video summaries, and support pronunciation practice.([[ai-guided-learning-audiovideo-2026]]) Multimodal knowledge graphs reason across images and text for educational tasks,([[multimodal-knowledge-graph-educational-reasoning]]) and multimodal representations improve [[inclusive-learning]] by translating information across modes (e.g., text to audio or visual). Domain applications include handwritten-math grading and diagnosis,([[llm-cognitive-diagnosis-handwritten-math]]) [[affective-tutoring|affective tutoring]] with multimodal signals,([[multimodal-affective-its-presentation]])([[kar-mathbuddy-affective-math-tutoring-2025]]) text-to-image learning in specialized fields,([[nuclear-diffusion-text-to-image-learning-2026]]) and privacy-aware multimodal classroom sensing.([[privacy-aware-classroom-incident-recognition-2026]]) Bird (2026) demonstrates a text-internal form of multimodality: fusing a fine-tuned ELECTRA transformer with computational-linguistics feature analysis to classify English literature by UK Key Stage, where the fused model (F1 0.996) far surpassed every unimodal baseline — evidence that combining representational forms, even within text, can outperform single-model approaches. ## Challenges and design implications 1. **Close the multimodal gap.** Multimodal tutoring systems should include visual grounding and structured-dialogue scaffolds rather than assuming vision capabilities are robust.([[syal-multimodal-dialogue-stem-2026]]) 2. **Treat multimodal prompting as a teachable skill.** AI literacy curricula must address modality-specific prompting, coherence across modes, and critical evaluation of multimodal outputs.([[multimodal-prompting-ai-literacy]]) 3. **Preserve human meaning-making.** Multimodal AI should augment, not replace, the learner's own construction and evaluation of meaning across modes.([[multimodal-learning-genai]]) 4. **Extend evaluation to multimodal validity.** [[assessment-validity|Assessment validity]], bias, and reliability must be examined when AI scores or generates multimodal artifacts.([[multimodal-item-parameter-estimation-2026]])([[ai-ed-evaluation]]) 5. **Watch equity and privacy.** Unreliable support on image-rich problems and the data demands of multimodal sensing both carry equity and privacy implications.([[syal-multimodal-dialogue-stem-2026]])([[privacy-aware-classroom-incident-recognition-2026]]) ## Connected Concepts - [[generative-ai]] - [[llm]] - [[knowledge-graph]] - [[intelligent-tutoring]] - [[ai-literacy]] - [[prompt-engineering]] - [[feedback]] - [[assessment]] - [[educational-measurement]] - [[item-response-theory]] - [[student-modeling]] - [[socratic-method]] - [[scaffolding]] - [[ai-ed-evaluation]] - [[benchmark]] - [[higher-ed]] - [[equity-in-ai-education]] - [[privacy]] - [[stem-education]] - [[inclusive-learning]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) - [[virtual-and-augmented-reality]] — gesture, voice and spatial input as learning channels - [[speech-and-voice-technologies]] - [[arts-design-and-media-education]] ## Connected Articles - [[burriss-multimodal-composition-critical-ai-literacy-2026]] — Video PSA composition on AI ethics as critical AI literacy pedagogy (Burriss et al. 2026) - [[student-attention-estimation-fairness-2026]] — Fairness-Aware Multimodal Transformer Modeling for Real-Time Student Attention Estimation - [[omniphys-multimodal-physics-benchmark-2026]] - [[ni-lam-multiliteracies-ai-portfolio-2026]] - [[drawedumath-vlm-struggling-students-2026]] — VLM performance on handwritten student math work (DrawEduMath, Lucy et al. 2026) - [[multimodal-learning-genai]] — Educator's guide to multimodal learning with generative AI (MMLD-AI model) - [[syal-multimodal-dialogue-stem-2026]] — The Multimodal Interference Effect and structured-dialogue recovery in STEM - [[multimodal-ai-feedback-learning]] — Multimodal AI feedback matches educators on learning, exceeds on perceptions - [[multimodal-prompting-ai-literacy]] — Students' multimodal prompting as epistemic work in AI literacy - [[multimodal-item-parameter-estimation-2026]] — Estimating IRT item parameters with multimodal LLMs - [[ai-guided-learning-audiovideo-2026]] — AI-guided audio-video learning support - [[multimodal-knowledge-graph-educational-reasoning]] — Multimodal knowledge graphs for educational reasoning - [[mllm-scientific-visualization-literacy]] — Multimodal LLM literacy for scientific visualization - [[multimodal-affective-its-presentation]] — Multimodal signals in affective intelligent tutoring - [[kar-mathbuddy-affective-math-tutoring-2025]] — Affective multimodal math tutoring - [[llm-cognitive-diagnosis-handwritten-math]] — LLM cognitive diagnosis of handwritten math - [[nuclear-diffusion-text-to-image-learning-2026]] — Text-to-image learning in nuclear engineering education - [[privacy-aware-classroom-incident-recognition-2026]] — Privacy-aware multimodal classroom sensing - [[genai-cybersecurity-ocr-multimodal-instruction-2025]] — Multimodal OCR instruction in cybersecurity education - [[cfes-p24-multimodal-slide-auditing-2026]] — CFES-P24: Benchmarking Multimodal LLMs for Slide Auditing - [[diagramir-educational-math-diagram-evaluation]] — DiagramIR: evaluating visual math diagrams from LLM-generated code - [[ai-grading-handwritten-physics-2026]] — AI grading of handwritten physics assessments (Olympiad) - [[gemini-lualatex-physics-video-transcription-2026]] — Gemini+LuaLaTeX math-accessible physics video transcription - [[bird-multimodal-educational-literature-2026]] — Multimodal fusion for classifying educational literature - [[lu-ai-multimodal-writing-critical-thinking-2026]] — Multimodal AI composing and critical thinking in primary writing (Lu et al. 2027) - [[cvengros-grading-handwritten-chemistry-ai-2026]] - [[geovad-bench-visual-chain-of-thought-geometry-2026]] — Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving - [[muse-vlm-artistic-image-benchmark-2026]] — MUSE: 12 tasks over 1,174 artworks show VLM capability as a dimension-specific profile, weakest in affective interpretation and viewpoint-dependent spatial reasoning (Zhu et al. 2026) - [[ai-assisted-physics-lab-report-assessment-2026]] — AI-Assisted Assessment of Experimental Physics Laboratory Reports: Potential, Limitations, and Support for Teaching Practice --- ## [Speech and Voice Technologies](https://edtechdev.github.io/aied/concepts/speech-and-voice-technologies/) > **Speech and voice [[ai-technologies|technologies]]** — the family of AI systems in which the spoken channel carries the interaction: automatic speech recognition (ASR) that transcribes, analyses, and grades what a learner says; text-to-speech (TTS) that generates narrated or dialogue-form instruction; voice-first agents that hold real-time spoken conversation with learners; and automated capture and scoring of oral performance. Where [[conversational-ai]] research is mostly about text chatbots and [[language-learning]] covers second-language pedagogy broadly, this page follows the audio itself — what a speech interface changes about learning, what it costs, and who it includes or excludes when the interface is a voice rather than a screen. ## Questions to Consider - ASR feedback helps stronger learners more than weaker ones: in one survey of 325 Chinese [[higher-ed|undergraduates]], the payoff of reflective behavior and motivation was weak and non-significant at low proficiency and grew markedly stronger at average and high levels. If a tool raises the average while widening the gap, is it a success? - TTS narration matched an instructor's own voice on comprehension, concentration, and overall evaluation — but dialogue-format TTS beat single-speaker TTS on comprehension and cognitive engagement while sounding *less* natural. Which would you choose for a first introduction to a concept, and which for a dense procedural walkthrough? - Two findings run against intuition: students who read the AI's output most heavily during interpreting tasks had the weakest delivery fluency, and a learner who answered in two words was marked correct while a fuller spoken answer was flagged. How much does a voice interface measure verbal fluency rather than understanding? - A voice agent with a British accent was treated as a tool, while agents with Indian and African American accents were anthropomorphized and treated as peers in [[k-12]] [[group-work|group work]]. If accent changes how a learner relates to an agent, who should decide which voice an educational product ships with? - Voice-first and offline systems exist mainly to reach learners that screen-based [[edtech-platform|edtech]] excluded — blind children, deaf and hard-of-hearing learners, learners in workshops with no connectivity. What would it take for accessibility to be the starting assumption of a speech product rather than a later fix? - Nearly every claim on this page comes from a short, single-site, self-report study with a fixed presentation order. Which of these findings would you trust enough to change your assessment or lesson design before a replication exists? ## Introduction Speech is the oldest educational interface and the newest one for AI. Systems that listen and speak now sit in language classrooms, vocational workshops, and high-school inquiry lessons. They are attractive for the usual reason — speech is fast, works hands-free and eyes-free, and is the medium in which pronunciation, fluency, and oral reasoning live — and for a newer one: a time-limited spoken answer is hard to outsource to a text generator, which makes the spoken channel a candidate answer to [[academic-integrity]] pressure on written assessment. The research here splits into four clusters: ASR pronunciation feedback, TTS-generated instruction, voice-first and spoken-dialogue partners, and AI-supported oral [[assessment]]. Running through all four is a question about [[equity-in-ai-education|equity]] — speech interfaces remove one barrier (sight, reading speed, typed literacy) while introducing others (accent, recognition accuracy, hardware, connectivity). ## ASR and Pronunciation Feedback The most consistent message from the ASR work is that the tool is not the treatment — the feedback design is. [[asr-english-speaking-feedback-metacognition-2026|Chen et al. (2026)]] surveyed 325 undergraduates at a Chinese [[teacher-role|teacher]]-training university and modeled accuracy, usage frequency, [[ai-feedback-quality|feedback quality]], and reflection-task design against reflective behavior and intrinsic [[motivation]], then against speaking improvement. Accurate error correction and structured reflection tasks drove both [[feedback]] internalization and reflection. More frequent use boosted reflection but had no independent effect on motivation. Recognition accuracy raised motivation, plausibly by building trust in the tool, but did not by itself trigger deeper processing — a learner can register a flagged error without analyzing its cause. Reflection was the stronger predictor of speaking gains, and proficiency moderated both pathways: weak and non-significant at low proficiency, markedly stronger at average and high levels. The advice that follows is to explain errors articulatorially rather than flag them, ramp complexity with readiness, and frame correction supportively. A second strand asks whether AI pronunciation feedback changes learners' disposition to speak. [[genai-pronunciation-feedback-wtc-2026|Lu et al. (2026)]] surveyed 1,701 Chinese university EFL learners and used covariance-based structural equation modeling with bias-corrected bootstrapping to test whether perceptions of [[generative-ai]] pronunciation feedback related to willingness to communicate, with pronunciation [[self-efficacy]] as mediator. The association was positive, and self-efficacy partially mediated it: the indirect path accounted for 69.9% of the total effect while a direct effect remained. All constructs were [[self-report-measures|self-reported]] at one time point, so the study describes a mechanism learners perceive rather than one that was manipulated. Automated pronunciation feedback can also be built without error labels. [[ai-guided-learning-audiovideo-2026|Kawamura (2026)]] describes Profy, which learns what good and poor pronunciation look like from largely unannotated speech via self-supervised learning, then shows learners which waveform regions drove the judgment and how far their acoustics sit from native-speaker distributions. With 10 Japanese learners of English rated by five American listeners, intelligibility improved, and unlike an elicited-imitation baseline the pre- and post-practice confidence intervals did not overlap — a small sample, but one that gives *where* and *how much* rather than a binary verdict. ASR also serves the classroom in a second role — converting spoken input into text that learners read. [[ai-mediated-input-medical-english-asr-2026|Stanchev (2026)]] rendered a 2,692-token Medical English coursebook passage on proteins with one text-to-speech voice (15:27, female Canadian) and submitted the identical audio to four online ASR services, then measured both word error rate (WER) and **coverage** (returned words as a share of the reference). The two free services returned near-complete transcripts (87.96% and 88.34% coverage, WER 15.23% and 15.00%), while the two paid services exported only previews (32.76% and 55.35% coverage, WER 73.21% and 48.77%). The measurement point is that WER alone conflates unavailable text with inaccurate text and would rank an export restriction as a transcription failure, so coverage belongs beside the error rate. Deletions dominated every service and clustered on structural markers such as figure references and page numbers, but a handful of meaning-altering substitutions — imino → amino, cystine → cysteine, protonated → protonatid, pH → phase — each inverted a biochemical fact, which is why the authors propose a three-stage workflow (generate the input, export the full transcript, verify domain terminology) before a transcript reaches learners, and why a terminology pass is the cheapest guardrail for medical English. ## TTS and Voice-Generated Instruction Generated narration turns out to be a viable substitute for a recorded teacher, and the interesting variation lies inside the format rather than between human and machine. [[llm-tts-dialogue-lesson-generation|Kumoi et al. (2026)]] built a three-stage, [[human-in-the-loop-ai|human-in-the-loop]] pipeline in which an [[llm]] drafts slides and TTS-optimized narration while the educator fact-checks at each stage, then ran a within-subject study with 245 first-year students at a Japanese prefectural high school across instructor-voiced video, single-speaker TTS, and Expert-×-Novice dialogue TTS. Comprehension, concentration, and overall evaluation did not differ significantly across the three (p > .14, r < 0.11), and equivalence testing kept all three within a ±0.5 margin: synthetic narration did not degrade the experience. Dialogue narration was better on comprehension (p = .006, r = .130), on reported deepened thinking (p = .019, r = .121), and on being able to explain the content to a friend (p = .004, r = .152), and 66.9% preferred it as most enjoyable. Two caveats: dialogue audio was rated significantly less natural (r = −.238), consistent with extra load from multiple speakers, and the dialogue group arrived with less [[prior-knowledge|prior knowledge]] (66.0% reporting knowing "nothing at all" versus 34.1%), which makes its advantage a conservative estimate. The follow-up argues that format should be matched to the learner, not the content. [[tts-dialogue-lessons-learner-characteristics-2026|Watanabe et al. (2026)]] gave 222 first-year high-school students teacher–student (TS), student–student (SS), and teacher–teacher (TT) dialogue lessons generated by an LLM and voiced through a TTS API, fitting linear mixed-effects models through an aptitude-treatment interaction lens. TT raised [[motivation]] more than TS, but the effect depended on learning style: the interaction with the Concrete Experience factor was positive (b = 0.162, p < .001) while the interaction with the reflection-and-conceptualization factor was negative (b = −0.238, p = .002). TT also drew a significantly *lower* overall evaluation than TS (b = −0.126, p = .005), which the authors attribute to intrinsic [[cognitive-offloading|cognitive load]] from dense expert-to-expert exchange; free responses name "stiffness of speech" in TT and "unnatural manner of speaking" in SS. Both TTS studies confound format with content and fixed viewing order, so their effect sizes are preliminary. ## Voice-First Companions and Spoken Dialogue Partners What changes when the interlocutor is a machine? [[ai-interlocutor-l2-spoken-dialogue|Scheinberg et al. (2026)]] had 78 university learners of German across four sites complete a counterbalanced spot-the-difference task with both a human peer and a real-time AI partner, then analyzed diarized ASR transcripts. Human dialogue was faster and better balanced, with many short turns, while the AI condition looked like supported monologue — fewer, longer turns, a smaller learner share of the floor, and higher within-turn fluency. The AI's verbose, syntactically regular input was associated with greater short-term uptake and stronger syntactic priming after controlling for input volume, and satisfaction tracked learners' own production fluency rather than how much they picked up. Voice-first design becomes a necessity rather than a convenience when the learner cannot use a screen. [[kutti-ai-voice-first-learning-companion|Kutti AI (Fadurudeen 2026)]] inverts edtech's visual assumption to reach an estimated 1.4 million blind children worldwide: children hear [[curriculum-design|curriculum]] content, answer aloud, and receive spoken feedback with no visual dependency. Three choices carry the design — a lightweight struggle-detection engine fusing response latency, wrong-attempt counts, and keyword hesitation cues to decide when to hint or simplify; a cross-language answer-matching pipeline so that code-switching and pronunciation variation are not penalized; and an offline on-device ASR path that removes the connectivity requirement. It is a systems contribution with no learning-outcomes evaluation, so its [[pedagogy|pedagogical]] claims remain hypotheses. Accent, meanwhile, is a design variable with social consequences. [[agent-voice-accents-k12-group-learning|Ravi et al. (2026)]] had 33 teachers interact with a GenAI voice agent in British, Indian, and African American accent conditions in K-12 [[collaborative-learning|group learning]]. The British-accented agent was treated largely as a tool and engaged with in detached, utility-based ways, while the Indian- and African American-accented agents were more readily anthropomorphized and integrated as peers, with stronger trust and reliance over time. Turn-taking, questioning patterns, and perceived [[community-of-inquiry|social presence]] all shifted with accent — the voice that makes an agent easy to treat as a neutral resource can make it easier to ignore as a collaborator. Heavier use of a spoken support tool is not automatically better. [[student-ai-interaction-consecutive-interpreting-2026|Kuang, Li and Weng (2026)]] had 22 interpreting trainees complete bidirectional computer-assisted consecutive interpreting tasks in systems built on ASR and machine translation, capturing eye movements, pen notes, and voice output. Four interaction profiles emerged — Intensive Engagers, Fast Scanners, Traditionalists, and Frequent Switchers — and 58.3% of stage-level observations changed profile between comprehending the source and producing the target. Only comprehension-stage patterns predicted product quality, and the AI-heaviest cluster scored lowest on fluency of delivery (5.46 against 6.17–6.46) and target language quality (5.70 against 6.35–6.67). Attention spent reading an AI transcript is attention not spent building one's own representation; the authors recommend teaching learners to reflect on their own strategy rather than prescribing one, since the pattern is invisible unless surfaced. A preregistered randomized field experiment tests the spoken channel against typed text within the same tutor, and finds the learning is in the tutoring rather than in the voice. [[ai-tutor-modality-randomized-field-experiment-2026|Yang, Van Alstyne and Dellarocas (2026)]] randomized 86 students in an online-MBA corporate-finance module between a structured tutor grounded in course materials and a holdout, then alternated each tutored student's channel weekly between voice and text so the comparison is identified inside each student: weekly mastery differed by −0.01 points of 11 (p = .98), and two one-sided tests rejected any true difference larger than ±0.3 SD. Voice transformed the process instead — 1.34 dialogue turns per minute against 0.75 in text weeks, and 11.3 questions against 4.6 — while the thinking pause collapsed (median tutor-to-student reply gap of 27 seconds in voice against 54 in text, where typed turns spent 25.9 seconds before the first keystroke plus 13.9 seconds composing) and the speech channel cost 2.8× more to deliver, at \$0.22 per voice minute against \$0.0165 per text message (\$413.46 voice and \$149.13 text over five weeks for 52 seats). On this evidence voice is an adoption and engagement lever, not an encoding one — which is exactly the trade-off any institution adding a spoken channel has to price. ## AI in Oral and Spoken Assessment Oral assessment has always been pedagogically strong and logistically expensive, and speech technology now attacks the cost side. [[asynchronous-oral-assessment-2026|Pentland, Lowenthal and Krier (2026)]] describe asynchronous oral assessments (AOAs): prompts delivered just-in-time, brief time-limited webcam responses that cannot be revisited, and grading against embedded rubrics with auto-generated transcripts. Across two courses taught by a single instructor without scheduling constraints, students scored higher on AOAs than on in-person multiple-choice exams — significantly in Study 2 (midterm median 92.5 vs 70, p < .001; final 94.2 vs 86.4, p = .002) and directionally in Study 1, with moderate cross-format correlations (τ = .44; τ = .25, non-significant). Students shifted to more active preparation, 90.91% saw the format as closer to workplace communication than written exams, and 81.82% reported engaging more actively with content. The authors are careful that these are format score differences, not evidence of [[learning-gains|learning gains]], and that the integrity advantage is inferred rather than measured; a re-scoring check with an LLM found instructor scores systematically higher with moderate-to-good agreement (ICC 0.73 and 0.60), offered as a reliability check rather than an endorsement of [[automated-assessment|automated grading]]. [[fenton-oral-exams-ai-authentic-assessment-2025|Fenton (2025)]] supplies the rationale in review form: because orals are real-time and interactive, students cannot generate answers in advance and memorize them, and the format probes reasoning at higher levels of Bloom's taxonomy rather than recall. Its benefits — [[personalized-learning|personalization]], authenticity, work-readiness, deeper knowledge — come with named costs: scheduling and logistics, learner anxiety, and bias risk around gender, ethnicity, language, response speed, and non-anonymous marking, though the evidence suggests orals can be as inclusive as written exams for some learners, including those with dyslexia. Two deployments show what the technology adds. [[ai-supported-oral-assessment-tvet-2026|AkoVoice (Adams 2026)]] assessed 33 learners across four Level 3 Automotive classes and one Level 3 Engineering class in New Zealand, on learner-owned phones against one mid-range laptop running open-weight models — Mistral 7B via Ollama, faster-whisper for speech-to-text, Chatterbox for TTS — with up to 12 learners assessed simultaneously and no internet connection at any point. AI was cast as evidence-surfacer rather than judge, leaving the assessor's judgment intact. Learner reaction was positive: 21 of 33 (64%) called the task realistic, and none of the 33 disagreed that speaking in real time felt right compared with a written [[eportfolio|portfolio]]. The performance data is more interesting. Word counts for identical questions varied five- to eight-fold between learners in every cohort (93–523, 92–796, 309–811), yet verbosity did not predict accuracy on closed questions; nine learners answered in 2 to 13 words and were all marked correct, the shortest being "3500 kgs", while the fuller "the safe working load is three tons" was flagged against a 3.5-ton marking guide. Adjacent Horticulture (14 learners) and Dairy (11 learners) trials found a 95% match between the agent's preliminary grade and the human tutor's grade. On the item side, [[gpt-item-generation-l2-listening-2026|Aryadoust and Wong (2026)]] compared iterative [[prompt-engineering|prompt engineering]] against fine-tuning for L2 listening items, producing 40 tests and 240 multiple-choice items: prompt refinement improved quality but plateaued, while fine-tuning GPT-4.1 on the optimized prompt with prompt design held constant improved generation further. Spoken assessment also risks measuring the wrong thing — [[multimodal-embodied-cognition-oral-explanations-2026|Morphew et al. (2026)]] track gesture alongside transcribed speech and argue that assessing only speech misses [[embodied-learning|embodied]] evidence of understanding, so that fluency can be mistaken for conceptual knowledge. ## Accessibility and Equity in the Spoken Channel Speech technology's accessibility record is mixed: the same channel that removes one barrier can install another. Deaf and hard-of-hearing learners are the clearest mismatch. [[llm-question-generation-deaf-hard-of-hearing-2026|Chen et al. (2026)]] designed an LLM question-generation system for DHH learners watching video, drawing on Language Deprivation Theory to explain why text-based prompting fits poorly with sign-based first languages. It added Visual Questions (timestamps where visual information is likely misread — rapid movement, misaligned captions, dense on-screen text) and Emotion Questions (timestamps where prior DHH learners reported frustration or confusion), refining a final bank of 30 questions with learners and instructors. With 16 users the prototype improved [[self-efficacy]] (M = 5.70, SD = 1.12 on a 7-point scale), and Deaf participants selected visual questions more often than hard-of-hearing participants, who reported reading captions fast enough not to need them. Accessibility here had to be built into generation, not appended to it. AkoVoice adds two equity considerations that rarely surface together. First, data: recordings were encrypted at rest, hosted locally, and deleted after 90 days, and the authors treat a participant's voice as a personal and cultural expression rather than a resource for training models or building biometric profiles — a stance that matters because the EU AI Act classifies AI used to evaluate learning outcomes in [[professional-training|vocational training]] as high-risk. Second, accent and register: speech-to-text accuracy was not fully tested across learner accents, and rubric criteria such as "engaging" were flagged as culturally specific and potentially penalizing to second-language speakers, prompting a review notes field. Set beside the accent findings in K-12 group work and the near-total visual assumption of mainstream edtech, the pattern is that the voice interface removes sight, reading-speed, and connectivity barriers while leaving accent, recognition accuracy, hardware cost, and cultural register open. Scope check. This page is about the spoken channel and the technology that carries it. For second-language pedagogy at large — writing development, grammatical accuracy, motivation, teacher practice — go to [[language-learning]], which spans L2 instruction with AI and treats speaking as one strand among several. For text-based chatbots and dialogue design independent of modality, go to [[conversational-ai]]; several articles here involve dialogue, but their question is about speech, not about prompting a chat model. For screen readers, captioning, tactile graphics, and assistive tools that are not primarily speech-based, go to [[assistive-technology]], which sits under the broader [[accessibility]] and [[inclusive-learning]] strands of this wiki. ## What Remains Uncertain Almost every strong claim here rests on one short study. The two TTS lesson studies ran in single Japanese high schools with fixed presentation order that confounded format with content, and their learning-outcome measures were self-reported on in-house scales. The ASR feedback model came from one Chinese university and was cross-sectional, so its proficiency-moderation finding describes covariance rather than causation. The willingness-to-communicate study measured perceptions of feedback, not feedback quality. The interpreting profile study is a 22-participant exploratory analysis in which output-stage effects did not reach significance. The AOA and AkoVoice deployments are single-institution, and AkoVoice reports no [[summative-assessment|summative]] learning evidence because its voice assessments ran in parallel to actual assessments. Kutti AI is a systems contribution with no outcome evaluation at all. What the collection does support is a set of design claims: feedback quality beats practice volume in ASR use, dialogue narration can aid comprehension at the cost of naturalness, format should be matched to learner characteristics rather than fixed, AI should surface evidence rather than replace the human assessor, and speaking in real time measures something writing does not — including an accuracy surprisingly indifferent to how many words the learner uses. Whether any of that holds at another site, another language, or another year is untested. ## Connected Concepts - [[language-learning]] - [[conversational-ai]] - [[accessibility]] - [[assistive-technology]] - [[assessment]] - [[authentic-assessment]] - [[human-in-the-loop-ai]] - [[multimodal]] - [[inclusive-learning]] - [[equity-in-ai-education]] ## Connected Articles - [[asr-english-speaking-feedback-metacognition-2026]] — ASR speaking feedback: quality beats quantity, and proficiency moderates who benefits (Chen et al. 2026) - [[genai-pronunciation-feedback-wtc-2026]] — Pronunciation self-efficacy mediates GenAI pronunciation feedback and willingness to communicate (1,701 EFL learners) - [[llm-tts-dialogue-lesson-generation]] — LLM + TTS lesson pipeline: synthetic narration matched instructor voice; dialogue aided comprehension - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learning-style × dialogue-format interactions in TTS-generated lessons (222 students) - [[ai-interlocutor-l2-spoken-dialogue]] — Human versus AI interlocutors: supported monologue, uptake, and syntactic priming in L2 German - [[kutti-ai-voice-first-learning-companion]] — Voice-first offline companion with real-time struggle detection for visually impaired children - [[agent-voice-accents-k12-group-learning]] — Agent accent shapes anthropomorphization and collaboration in K-12 group work (33 teachers) - [[asynchronous-oral-assessment-2026]] — Asynchronous oral assessments: scalable spoken assessment, with a learning-gains caveat - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Review making the case for oral exams as authentic, integrity-protective assessment - [[ai-supported-oral-assessment-tvet-2026]] — AkoVoice: offline voice assessment in vocational workshops where AI surfaces evidence, not grades - [[student-ai-interaction-consecutive-interpreting-2026]] — Four interaction profiles among interpreting trainees using ASR/MT support - [[gpt-item-generation-l2-listening-2026]] — Prompting versus fine-tuning GPT for L2 listening item generation - [[multimodal-embodied-cognition-oral-explanations-2026]] — Gesture plus speech: why speech-only oral assessment misses evidence of understanding - [[ai-guided-learning-audiovideo-2026]] — Adaptive audio speed, voice-preserving video summaries, and unlabeled pronunciation feedback - [[llm-question-generation-deaf-hard-of-hearing-2026]] — LLM question generation for deaf and hard-of-hearing learners - [[ai-tutor-modality-randomized-field-experiment-2026]] — When AI Tutors Speak: Evidence from a Randomized Field Experiment - [[ai-mediated-input-medical-english-asr-2026]] — AI-Mediated Input Transformation in Medical English: Speech Recognition and Transcript Reliability --- ## [Visualization](https://edtechdev.github.io/aied/concepts/visualization/) > **Visualization** — the use of data visualizations, infographics, dashboards, charts, diagrams, and other graphical representations to make information comprehensible for learning and analysis. Across education, visualization is increasingly both generated by AI (text-to-image, [[multimodal]] slide and chart analysis) and used as the interface through which learners, teachers, and analytics systems reason over shared data. ## Questions to Consider - A chart or dashboard can make data clear — but is simply *seeing* a visualization the same as understanding it? The page's [[research-methods-aied|research]] suggests how you interact with a visualization matters more than the chart itself. Recall a dashboard or graph you looked at but didn't really learn from. What was missing from the mere display? - Conventional learning dashboards follow a 'show data, hope for insight' model. The finding here is that learners who answer questions about their data *before* seeing the metrics reflect and calibrate better than those who passively view charts. Why might being forced to predict first change what you get out of seeing the actual data? - AI can now generate accurate visualizations of specialized content — one study raised domain accuracy from 12% to 78% by fine-tuning a text-to-image model on nuclear concepts. But the page also finds no model is uniformly competent. Where would you trust an AI-generated visual, and where would you insist on checking it against a human expert? - Participants in one study found AI-generated data comics more engaging and comprehensible, yet many also flagged misinformation risk and information overload. How do you weigh the appeal of a compelling AI visual against its potential to mislead — and what would you verify before trusting or using it? - The page warns that a model's *confidence* is not its *reliability*: systems can sharply diverge on severity judgments while getting basic constructs right. If you relied on an AI tool to assess slides, essays, or data, how would you discover where its confidence hid a serious error? - One study found students spent most of their gaze on code despite elaborate visual scaffolds — visual aids simply didn't capture attention for everyone. What does that suggest about assuming a nice diagram or dashboard will automatically help all learners engage? What else besides visuals shapes how people actually use a tool? ## Introduction Visualization is the use of graphical representation — dashboards, charts, diagrams, infographics and multimodal displays — to make learning and learning data intelligible. Its most established educational role is the [[learning-analytics|learning analytics]] dashboard, where the dominant *show data and hope for insight* model has given way to interactive designs: the evidence indicates that how learners interact with a representation matters more than whether they see it, and that self-elicit prompts and [[pedagogical-agent|pedagogical agents]] improve calibration more than passive metrics do. The same question — does the display provoke thinking or substitute for it? — links visualization to [[desirable-difficulties]] and [[metacognition]]. ## Visualization as a Learning Interface The most established role of visualization in education is the learning analytics dashboard. Conventional Learning Analytics Dashboards (LADs) operate on a "show data → hope for insight" model, presenting behavioral metrics in charts that learners passively view. Research on [[interactive-learning-dashboards-engagement]] challenges this paradigm: when a dashboard adds an [[llm]]-powered [[pedagogical-agent|pedagogical agent]] and an interactive Judgment of Learning self-assessment, the "elicit" condition — where learners answer questions about their data before seeing metrics — produced more reflection and more accurate mastery calibration than either a passive dashboard or a "telling" agent. The lesson is that how learners interact with visualizations matters more than merely seeing them. This connects to [[learning-analytics]] and [[self-regulated-learning]], where visual feedback supports [[metacognition|metacognitive]] judgment calibration rather than simple information display. [[teacher-role|Teacher]]-facing dashboards add a distinct set of design lessons. [[mejia-domenzain-ml-findings-teachers-blended-2026|Mejia-Domenzain et al. (2026)]] found that teachers systematically preferred simpler, more traditional visualizations (bar plots, pie charts, legends) even when more complex designs (e.g., heatmaps) yielded more detailed insights — visual preference did not always align with informativeness, echoing debates about pie charts' comparative readability. Visualization literacy (VL) did not drive design preferences, but higher-VL teachers produced deeper, more detailed interpretations (e.g., more of them identified trends in time-series data), confirming VL as a [[research-methods-aied|confounder]] for gauging how teachers read analytics designs. For group comparison, teachers strongly favored superposition over juxtaposition, and preferred plots that displayed full information (e.g., including a "students who did not watch" group) rather than explicit difference encoding, although younger teachers ranked difference plots higher. These findings argue that dashboard design must balance teachers' stated preferences against the interpretative depth that more complex encodings afford. Dashboards also function as shared representations that bridge human and AI reasoning. The CLARA system uses LLM-generated artifacts — concept maps and seven-dimension collaboration assessments — as common ground between dashboard users and [[agentic-ai|AI agents]], indexing them into separate vector collections so both parties reason over the same visible, queryable material. Similarly, the Expert Cognition Dashboard reframes analytics as "cognition intelligence," turning raw learner behaviors into interpretable cognition structures across individual, class, and AI Twin expert levels. These systems position visualization not as an output but as embedded reasoning infrastructure within [[ai-technologies]]-native education. ## AI-Generated and Multimodal Visual Content A second major strand concerns AI producing visualizations directly. [[nuclear-diffusion-text-to-image-learning-2026]] shows that domain-adapted text-to-image models can generate accurate illustrations of specialized STEM concepts: fine-tuning Stable Diffusion on nuclear domain images raised domain accuracy from 12% to 78%, enabling instructors to produce correct reactor-component and safety-system visualizations on demand. This generative capacity is powerful but uneven. [[mllm-scientific-visualization-literacy]] [[benchmark|benchmarks]] six multimodal large language models against 485 human participants on scientific visualization literacy, finding no uniform competence: closed-source Gemini exceeded the human mean on several subsets while all [[open-source]] models fell below it, with particular weaknesses in fine-grained [[quantitative-research|quantitative]] estimation and texture-based or integration-based visualizations. AI should therefore support — not substitute for — human visualization literacy, a finding with direct [[ai-literacy]] and [[formative-assessment]] implications. Ethics and reliability temper enthusiasm for AI-generated visual content. [[data-comics-for-education-evaluating-effectiveness-benefits-ethics]] found [[generative-ai|GenAI]]-assisted data comics improved [[student-engagement|engagement]] and comprehension over conventional visualizations regardless of prior visualization literacy, yet participants raised concerns about misinformation risk and [[academic-integrity|authorship]] attribution, and two-thirds flagged downsides such as information overload from overly busy layouts. The counterfactual CFES-P24 benchmark extends this scrutiny to [[cfes-p24-multimodal-slide-auditing-2026|slide auditing]], showing that multimodal LLMs can reliably recognize [[learning-design]] constructs (operations, principles, evidence localization) while sharply diverging on comparative judgment and severity calibration — evidence that composite scores conceal which capability fails and that confidence is not reliability. Together these works argue for layered [[ai-ed-evaluation|evaluation of AI]]-generated visuals rather than holistic ratings. ## Slide, Comic, and Multi-View Tools in Practice Practical systems apply these principles at scale. AISSA combines LLM-based rubric scoring with learning analytics dashboards to deliver automated, iterative feedback on student presentation slides, processing 90 presentations in 1–3 minutes each at cents-per-evaluation cost with high perceived [[usability-research|usability]] — while students selectively applied feedback, sometimes disregarding recommendations that conflicted with their visual design. In [[cs-education|coding education]], Flowcode pairs a code-structure flowchart with a learning-oriented chat to help novice creative coders understand and extend found examples, where visualization and [[desirable-difficulties|productive friction]] steer AI use toward learning rather than bypass. Yet visual [[scaffolding|scaffolds]] are not universally effective: [[code-anchor-multi-view-visualization]] found students spent ~47% of gaze time on code despite visual scaffolds, driven by agency, representational fit, and the perceived legitimacy of metaphorical views. These [[student-experience|learner-experience]] findings caution that visualization design must attend to [[affective-computing|affective]] and social factors, not just cognitive affordances. ## Implications Across these twelve works, visualization emerges as a dual-use medium: AI increasingly generates and interprets visualizations, while dashboards and interactive visuals serve as the shared surface for human-AI sensemaking. Generative text-to-image and multimodal analysis extend the reach of visualization into specialized [[stem-education|STEM]] content and [[ai-feedback-quality|automated feedback]], but uneven model competence, severity-calibration failures, and [[ethics|ethical]] concerns over misinformation and authorship demand careful, layered verification. For designers and educators, the strongest conclusion is that interactivity and engagement — eliciting learner reasoning over visual data, letting users control cognitive effort, and treating AI-produced visuals as shared infrastructure rather than endpoints — matter more than the fidelity of the chart itself. ## Connected Concepts - [[learning-analytics]] - [[multimodal]] - [[generative-ai]] - [[ai-technologies]] - [[ai-literacy]] - [[storytelling-in-education]] - [[learning-design]] - [[assessment-validity]] - [[virtual-and-augmented-reality]] — spatial and three-dimensional representation ## Connected Articles - [[interactive-learning-dashboards-engagement]] — Rethinking learning visualizations as engagement tools via pedagogical agents - [[clara-collaboration-literacy-dashboard]] — AI-augmented analytics dashboard with concept maps and 7C assessments - [[wordstream-glass-learning-analytics]] — Quantitative encoding of qualitative learning analytics - [[mllm-scientific-visualization-literacy]] — Benchmarking multimodal LLMs for scientific visualization literacy - [[nuclear-diffusion-text-to-image-learning-2026]] — Domain-adapted text-to-image models for nuclear concept visualization - [[data-comics-for-education-evaluating-effectiveness-benefits-ethics]] — Effectiveness, benefits, and ethics of AI-assisted data comics - [[cfes-p24-multimodal-slide-auditing-2026]] — Counterfactual benchmark for multimodal slide auditing - [[aissa-slides-analysis]] — AI-based student slides analysis tool for academic presentations - [[mejia-domenzain-ml-findings-teachers-blended-2026]] — Making ML findings accessible to teachers in blended classrooms --- ## [Educational NLP](https://edtechdev.github.io/aied/concepts/educational-nlp/) > **Educational NLP** applies language [[ai-technologies|technologies]] to learning: [[llm-item-difficulty-prediction]], [[teaching-feedback-classification-benchmark]], [[llm-sentiment-analysis-education-research]], and [[vocabulary-difficulty-prediction]] show LLMs advancing analysis of student language at scale ([[educational-measurement]], educational-nlp). ## Questions to Consider - When an LLM analyzes thousands of student essays or discussion posts for sentiment, what might it be getting right, and what about the language of learning do you suspect it's missing? - Natural language processing can now estimate the difficulty of vocabulary and test items, and classify [[teacher-role|teaching]] feedback at scale. If those predictions feed adaptive systems, who checks whether the machine's judgments about language are actually right for the learners using them? - How is analyzing student language different from understanding it? Where might the line between correlation and genuine insight blur when NLP scales up sentiment and feedback analysis? - This concept connects NLP to tutoring, student modeling, and measurement. Before reading, how much of 'understanding a student' do you think can be captured from their written or spoken language alone — and what gets left out? ## Introduction ### What educational NLP does Natural language processing in education applies computational methods to the language of teaching and learning — student essays, responses, discussion posts, feedback, and instructional text. [[llm|LLMs]] have dramatically expanded what can be analyzed automatically, enabling fine-grained understanding of student language that was previously impractical at scale. ### Applications documented in the knowledge base - **Analysis of student language.** [[llm-sentiment-analysis-education-research]] applies LLM-based sentiment analysis to educational research, extracting emotional and evaluative signals from student text at scale, feeding [[learning-analytics]] and [[affective-computing]]. - **Prediction and measurement.** [[llm-item-difficulty-prediction]] and [[vocabulary-difficulty-prediction]] use language models to estimate item and text difficulty — core inputs to [[educational-measurement]], [[adaptive-learning]], and [[item-response-theory]] models. - **Readability and [[curriculum-design|curriculum]] alignment.** Bird (2026) fuses transformer text classification with computational-linguistics features to classify English literature by UK Key Stage, reaching an F1 of 0.996 — a data-driven complement to [[vocabulary-difficulty-prediction]] and [[llm-item-difficulty-prediction]] for [[educational-measurement]] and reading-level alignment. - **Feedback and classification.** [[teaching-feedback-classification-benchmark]] provides a [[benchmark]] for classifying teaching feedback, advancing [[feedback|Feedback Loop]] research and [[pedagogical-llm-training]]. - **Short-answer assessment in science.** Morley et al.'s [[meta-analysis-systematic-review|scoping review]] of transformer-based auto-marking of short-answer science questions (2017–early 2024) shows BERT-family models became the field's dominant workhorse for [[automated-assessment|free-text marking]] before larger [[llm|LLMs]] were adopted via [[prompt-engineering|prompting]], and that models augmented with domain knowledge — extra pre-training, rubric or textbook data, meta-learning — consistently outperformed those without ([[auto-marking-short-answer-science-2026]]). - **Context-sensitivity vs. reference matching in open-response grading.** Benchmarking eleven [[generative-ai|GenAI]] and sentence-embedding models on 1,885 software-engineering open-ended answers, [[pecuchova-automated-grading-open-ended-genai-2026|Pecuchova, Benko & Drlik (2025)]] show that context-sensitive [[llm|LLMs]] (GPTo1 best, almost-perfect human agreement) beat cosine-similarity reference-based models (BERT, RoBERTa, T5, USE), which systematically misclassified valid but differently-phrased responses. Their NLI analysis revealed that many semantically correct answers fell into the *contradiction* category relative to reference answers — evidence that educational NLP grading must accommodate students' short, diverse, own-word phrasing rather than rigid reference alignment. - **Learner-generated question classification.** [[lee-learner-question-types-ai-education-2026|Lee, Atif & Kang (2026)]] classify 434 authentic learner questions from 11 IT students across 12 courses into three [[constructivist]] instructional roles — knowledge transmitter, facilitator, and co-learner — and benchmark four transformers on the task. DeBERTa led at 86.36% accuracy (F1 86.52%) with 96.67% precision on factual knowledge-transmitter questions, yet only 78.79% precision on facilitator queries; fine-tuned BERT reached the best co-learner recall (92.00%) at lower precision. The result mirrors the field's recurring pattern that strong aggregate scores mask weak discrimination on higher-order categories: conceptual overlap between roles, ambiguous learner intent, and domain-specific technical phrasing misread as cognitive depth all defeat surface lexical features, arguing for context-aware embeddings, multi-turn dialogue signals, and intent-sensitive features ([[cross-dataset-bloom-question-classification]], [[llm-educational-question-cognitive-depth]]). ### Connection to tutoring and measurement Educational NLP underpins both the analysis of learner language ([[student-modeling]], [[knowledge-tracing]]) and the generation of adaptive instructional content ([[intelligent-tutoring]], [[scaffolding]]). [[ai-generated-interactive-fiction-education-2026]] demonstrates NLP-driven content generation for learning, while [[zerkouk-comprehensive-review-its-2025]] situates NLP within the broader [[intelligent-tutoring]] landscape. As LLM-based analysis grows, [[rct]] and [[research-methods-aied]] frameworks matter for validating that NLP-derived insights genuinely improve learning. ## Connected Concepts - [[intelligent-tutoring]] - [[student-modeling]] - [[knowledge-tracing]] - [[socratic-method]] - [[scaffolding]] - [[adaptive-learning]] - [[pedagogical-llm-training]] - [[metacognition]] - [[rct]] - [[learning-analytics]] - [[educational-policy-ai]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) ## Connected Articles - [[lee-learner-question-types-ai-education-2026]] — Transformer classification of learner questions into constructivist roles (Lee, Atif & Kang 2026) - [[bert-discourse-english-teaching-2026]] — Automatic discourse relation classification with BERT for English teaching - [[studychat-student-dialogues-chatgpt-ai-course-2026]] — The StudyChat dataset of student–LLM dialogues in an AI course - [[nspa-neuro-symbolic-pedagogical-alignment-2026]] — Neuro-symbolic pedagogical alignment (NSPA) - [[ai-generated-interactive-fiction-education-2026]] - [[zerkouk-comprehensive-review-its-2025]] - [[diagramir-educational-math-diagram-evaluation]] — DiagramIR: IR-based evaluation of math diagrams - [[shap-llm-rationales-teaching-quality-assessment]] — SHAP and LLM rationales for rubric-based teaching quality - [[distilling-self-explaining-lm-learning-analytics-2026]] — Distilling self-explaining LM for learning analytics - [[bird-multimodal-educational-literature-2026]] — Multimodal fusion for classifying educational literature - [[auto-marking-short-answer-science-2026]] - [[pecuchova-automated-grading-open-ended-genai-2026]] --- ## [Reinforcement Learning](https://edtechdev.github.io/aied/concepts/reinforcement-learning/) > **Reinforcement learning** trains AI tutors and agents through reward signals: [[special-r1-rl-special-education]], [[singh-eduqwen-pedagogical-rl-2026]], [[pedagogical-safety-rl]], and [[ai-coaching-rl-skill-development]] align RL with pedagogical objectives, including safety and skill transfer ([[intelligent-tutoring]], [[agentic-ai]]). ## Questions to Consider - An RL tutor 'learns' what to do by maximizing a reward signal. Before you read, what could be wrong with an AI that optimizes for a reward — specifically if the reward is something like 'student clicks continue' or 'correct answer now'? - The page notes that reward design encodes educational values. If you had to specify the reward an AI tutor should maximize, what would you put in it — and what would your reward accidentally ignore or reward incorrectly? - RL trains agents to make long-horizon sequences of decisions (what hint, when to advance difficulty, how to pace) rather than single answers. How is that different from the moment-to-moment correctness you might naively reward — and why does the difference matter for learning? - Safety constraints can be integrated into RL so that reward optimization doesn't come at the cost of learner well-being. Think of a 'helpful' behavior a reward-optimizing tutor might exhibit that would actually be pedagogically harmful (e.g., giving away answers to inflate completion). Where would your safety line go? - Reward optimization can preserve or destroy productive struggle, depending on design. From your experience, is 'student completes task' the same as 'student learns'? Where have you seen an AI optimized for the former while undermining the latter? ## Introduction ### How reinforcement learning works in AIED Reinforcement learning (RL) trains an agent by rewarding desired behavior — the agent learns a policy that maximizes cumulative reward through trial and error. In AI in education, RL is used to train tutoring agents and learning companions that must make sequences of decisions (what hint to give, when to advance difficulty, how to pace practice) rather than single answers. This makes RL well suited to [[adaptive-learning]] and [[intelligent-tutoring]] where long-horizon pedagogical decisions matter. ### Applications documented in the knowledge base - **Pedagogically aligned RL.** [[singh-eduqwen-pedagogical-rl-2026|EduQwen]] uses an RL-SFT-RL pipeline to train a model that *guides* rather than answers, aligning reward with pedagogical goals; [[special-r1-rl-special-education]] applies RL to tutor design for [[special-education]]. - **Safety and skill transfer.** [[pedagogical-safety-rl]] integrates safety constraints into RL-based tutoring so that reward optimization does not come at the cost of learner well-being; [[ai-coaching-rl-skill-development]] shows RL-driven coaching that supports genuine skill development and transfer. - **Simulation and practice.** [[history-aware-student-simulation]] and [[q-learning-lab-rl-teaching]] use RL and simulated learners to train and evaluate [[pedagogical-agent|pedagogical agents]], connecting RL to [[student-modeling]] and [[learning-analytics]]. ### Evidence across the field A PRISMA-standard [[riedmann-reinforcement-learning-education-review-2026|systematic review of RL in education (Riedmann, Schaper & Lugrin, 2025)]] synthesized 89 studies (2000–2024), finding a sharp post-2016 growth in [[adaptive-learning]] and [[intelligent-tutoring|tutoring]] applications concentrated in STEM (especially [[math-education]]). It reports that model-free RL dominated (n = 72) with Q-learning the most common algorithm, yet classical RL was more consistently effective than Deep RL (61% vs 36% of papers showing significant superiority); that adaptation split into content-scheduling (n = 53) and guidance-related (n = 36) mechanisms, with RL beating baselines more often on guidance; and that learning gain — especially normalized learning gain — was the most effective reward source. The review also warns that over half of studies (n = 54) skipped statistical testing, so the field's growth has outpaced its methodological rigor. ### Connection to the knowledge base RL underpins much modern [[agentic-ai]] and [[intelligent-tutoring]] design, where the agent must optimize long-term learning rather than a single correct response. It connects to [[pedagogical-llm-training]] (RL as a training method), [[scaffolding]] (reward design that preserves productive struggle), and [[self-regulated-learning]] (agents that help learners regulate their own strategy). Because reward design encodes educational values, RL research in AIED is tightly tied to [[pedagogical-safety]] and to the equity considerations of [[equity-in-ai-education|equitable]] tutor behavior. ## Connected Concepts - [[intelligent-tutoring]] - [[student-experience]] - [[stem-education]] - [[self-regulated-learning]] - [[scaffolding]] - [[active-learning]] - [[edtech-platform]] - [[higher-ed]] - [[learning-analytics]] - [[open-source]] - [[pedagogical-safety]] - [[pedagogical-llm-training]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) ## Connected Articles - [[history-aware-student-simulation]] - [[q-learning-lab-rl-teaching]] - [[singh-eduqwen-pedagogical-rl-2026]] - [[residencyrl-clinical-rl-training-2026]] - [[learnlm-improving-gemini-learning]] — LearnLM: RLHF for pedagogical instruction following - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[riedmann-reinforcement-learning-education-review-2026]] --- ## [Knowledge Graph](https://edtechdev.github.io/aied/concepts/knowledge-graph/) > **Knowledge graph** — a structured representation of concepts and their relationships used to model domain knowledge, student understanding, and learning dependencies in [[ai-education|AI in education]] systems. Knowledge graphs enable AI systems to reason about what students know, what they need to learn next, and how concepts relate to each other. ## Questions to Consider - A knowledge graph captures not just concepts but their relationships — prerequisites, similarity, hierarchy. Why might knowing how concepts relate be more useful to an adaptive system than a flat list of skills? - How are prerequisite relationships in a domain like yours? Can you think of a topic where students routinely struggle because they're missing a foundational concept the graph would reveal? - The page describes using knowledge graphs to detect knowledge gaps — where learners are missing foundational concepts. How might surfacing that gap change what an AI tutor decides to teach next? - Knowledge graphs can be built manually or automatically by LLMs from educational text. What are the risks of letting an AI construct the concept structure that a tutor will then reason over? - If knowledge graphs provide the domain structure that AI agents reason over, what happens to trust and accuracy when the graph itself contains an error or a biased relationship? - A knowledge graph is described as the structural backbone enabling fine-grained diagnosis and personalized paths. In your own [[teacher-role|teaching]] or design, what would you need a knowledge graph of your subject to capture — and what would it leave out? ## Introduction Knowledge graphs provide the structural backbone for many intelligent education systems. Unlike flat lists of skills or concepts, knowledge graphs capture prerequisite relationships, similarity, and hierarchical organization — essential for [[adaptive-learning]], [[knowledge-tracing]], and [[student-modeling]]. ## How knowledge graphs are used in AIED Knowledge graphs are a recurring structural mechanism across the knowledge base's AIED [[research-methods-aied|research]], serving several distinct roles: - **[[knowledge-tracing]] models** use concept graphs to propagate student proficiency estimates across related skills, improving prediction accuracy when data is sparse. - **[[student-modeling]] systems** leverage knowledge graphs to represent what learners know in a semantically meaningful way, enabling fine-grained diagnosis. - **[[adaptive-learning]] platforms** use prerequisite graphs to sequence content and recommend [[personalized-learning|personalized learning]] paths. - **[[cognitive-diagnosis]] frameworks** like [[xie-hillm-cd-2026|HiLLM-CD]] construct concept trees from educational text using LLMs, eliminating manual annotation. - **Knowledge-graph-augmented tutoring:** [[quantum-education-its|ITAS]] uses a knowledge graph of quantum concepts (with explicit prerequisite relationships) to drive a multi-agent tutoring system, traversing the graph to select next topics for counterintuitive material. - **Curriculum and course modeling:** [[coursegraph-cs-course-comparison-2026|CourseGraph]] compares CS course structures across institutions using graph representations; [[learnity-graphs-lifelong-learning-framework-2026|Learnity graphs]] model [[lifelong-learning|lifelong learning]] pathways. - **Prerequisite-relation learning:** [[proprl-prerequisite-relation-learning|ProPrL]] learns prerequisite relations among concepts, formalizing the edges that knowledge graphs encode. - **Knowledge-gap detection:** [[knowledge-gap-detection-ai-tas|Knowledge gap detection]] uses graph-based reasoning in AI teaching assistants to identify where learners are missing foundational concepts. - **[[multimodal]] and explainable reasoning:** [[multimodal-knowledge-graph-educational-reasoning|multimodal knowledge graphs]] extend graph structure across content modalities; [[fair-explainable-edu-recommendations|fair and explainable recommendations]] combine knowledge-graph embeddings with sequential modeling (a hybrid HKG-GRU framework). - **Instructionally structured graphs for resource recommendation:** [[hybrid-cf-kg-recommendation-multimodal-teaching-2026|Liu, Sun & Song (2026)]] decompose each teaching-resource entity into four instructional dimensions (teaching context, cognitive level, technological feature, cultural adaptability), compute user-dependent semantic similarity over those dimensions, and fuse it with collaborative filtering via an ability- and progress-aware coefficient — encoding pedagogical structure directly into the recommendation signal rather than treating resources as consumption items. - **Ontology-based knowledge bases:** [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026|Ivanova (2026)]] proposes a layered, hybrid knowledge-base architecture grounded in description logic that replaces the classic ITS single-ontology models with **systems of mapped ontologies** — adding procedural (rule-based), probabilistic/fuzzy, and ML-extracted implicit knowledge — plus a metadata framework for describing, discovering, and reusing educational ontologies. - **[[scaffolding|Scaffolding]] and writing:** [[veriforge-narrative-drafting-scaffolding-2026|Veriforge]] and [[visual-query-tracer-declarative-logic-learning|visual query tracing]] apply graph-based structure to narrative drafting and declarative-logic learning. - **Human-curated literary graphs, and what an audit exposes:** [[incipit-axiom-grounded-scaffolding-literary-creation-2026|Incipit]] graphs literary premises — 1,455 axiom records, 1,464 mappings to 149 works, and 472 typed relationships — with [[llm|language models]] proposing candidate formulations that human curators selected and grounded. Its recomputed audit is as instructive as its structure: every endpoint resolves and no duplicate or self-link remains, yet 1,448 of the 1,455 axioms map to exactly one work (so cross-work reuse is sparse), the context taxonomy cannot separate its two context types, and no provenance record survives, leaving the snapshot unable to reconstruct its own pipeline. Structural validity is not interpretive quality, and a curated graph without provenance can be neither audited nor refreshed. ## LLM-driven knowledge graph construction Recent research explores using [[llm|LLMs]] to automatically construct knowledge graphs from educational content. The [[xie-hillm-cd-2026|HiLLM-CD]] framework uses multi-agent LLM pipelines to generate exercise-concept links and hierarchical concept trees, reducing reliance on expert annotation. This connects to broader [[generative-ai]] applications in curriculum design and automated content organization, and to [[rag]] (retrieval-augmented generation), where graph-structured knowledge can improve retrieval quality over flat similarity search. ## Relationship to other concepts Knowledge graphs connect to [[learning-design]] (defining what to teach), [[curriculum-design]] (how to sequence it), and [[learning-analytics]] (extracting insights from student interaction data). They are foundational to [[intelligent-tutoring]] systems that need structured representations of educational domains. As AI agents become more common in education, knowledge graphs provide the domain structure that [[agentic-ai|agentic systems]] reason over — a pattern seen in [[quantum-education-its|ITAS]] and knowledge-gap-detection teaching assistants. ## Connected Concepts - [[adaptive-learning]] - [[knowledge-tracing]] - [[intelligent-tutoring]] - [[cognitive-diagnosis]] - [[student-modeling]] - [[learning-analytics]] - [[curriculum-design]] - [[learning-design]] - [[generative-ai]] - [[llm]] - [[rag]] - [[agentic-ai]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[incipit-axiom-grounded-scaffolding-literary-creation-2026]] — A curator-built graph of 1,455 literary axioms with a structural audit and no provenance record (Liu & Zhao 2026) - [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] — Ontology-based layered hybrid knowledge model for personalized e-learning - [[learnity-graphs-lifelong-learning-framework-2026]] — Learnity graphs for lifelong learning - [[veriforge-narrative-drafting-scaffolding-2026]] — Veriforge: narrative-drafting scaffolds - [[quantum-education-its]] — Quantum education intelligent tutoring (ITAS) - [[multimodal-knowledge-graph-educational-reasoning]] — Multimodal knowledge graphs for educational reasoning - [[coursegraph-cs-course-comparison-2026]] — CourseGraph: CS course comparison - [[proprl-prerequisite-relation-learning]] — ProPrL: prerequisite-relation learning - [[knowledge-gap-detection-ai-tas]] — Knowledge-gap detection in AI teaching assistants - [[visual-query-tracer-declarative-logic-learning]] — Visual query tracer for declarative logic learning - [[learnopt-exam-cognitive-structure]] — LearnOpt: exam cognitive structure - [[fair-explainable-edu-recommendations]] — Fair and explainable educational recommendations - [[hybrid-cf-kg-recommendation-multimodal-teaching-2026]] — Hybrid CF–KG cross-domain recommendation for multimodal teaching resources - [[concept-catalyst-engineering-scaffolds]] — Concept Catalyst engineering scaffolds - [[xie-hillm-cd-2026]] — HiLLM-CD: LLM-driven cognitive diagnosis - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[cogevol-learning-environment-generation-2026]] — CogEvol: Learning Environment Generation - [[ai-information-extraction-undergraduate-thesis-2026]] — AI-powered information extraction supporting undergraduate thesis and research-based learning (An et al. 2026) --- ## [Robots in Education](https://edtechdev.github.io/aied/concepts/educational-robotics/) > **Robots in education (educational robotics)** — the use of physical or simulated robots as tools for [[teacher-role|teaching]] and learning. Educational robotics spans a wide spectrum: from programmable kits that teach computational thinking and programming, to socially assistive and humanoid robots that tutor, tell stories, model sign language, or rehearse social skills. It is valued for fostering [[problem-solving|problem solving]], [[critical-thinking|critical thinking]], [[creativity]], and STEAM engagement, and for making abstract computing concepts tangible through embodied interaction. The knowledge base's robotics corpus spans [[curriculum-design|curriculum]]-integrated programming, LLM-powered conversational tutors, socially assistive [[storytelling-in-education|storytelling]] robots, and role-play for social-emotional learning. It is underpinned by two closely related areas absorbed here: **social robots** (robots designed for social interaction and relationship-building) and **human–robot interaction (HRI)** (the study of how people perceive, trust, and learn with robots). ## Questions to Consider - A robot in the classroom adds an embodied, [[community-of-inquiry|social presence]] that a chatbot on a screen can't. What do you think the physical body and social cues of a robot change about how students learn, trust, and engage — and what might they distract from? - Social robots use human-like speech, gestures, and personality to teach, tell stories, or rehearse social skills. Is a robot that looks and acts human inherently better for learning, or could that social presence bring risks (misinformation, over-reliance, privacy) that software-only tools don't? - LLMs now let social robots converse fluently. If a robot can talk like a tutor, what still depends on its physical embodiment — and where does adding a 'body' really matter for learning versus just being a novelty? - Think of a time you learned something by physically manipulating an object or watching your actions produce a visible result. How might programming a physical robot ground abstract ideas (like program logic) more effectively than writing code on a screen? ## Introduction Educational robotics is a distinct but closely related application of [[ai-education|AI in education]]. Unlike software-only [[intelligent-tutoring|intelligent tutoring]] or [[llm]] chatbots, robots add an **embodied** and often **social** presence — a physical agent that learners can see, manipulate, and (increasingly) converse with. This embodiment is central to their [[pedagogy|pedagogical]] value: it grounds abstract program logic in observable behavior, and it can support relationship-building and emotional engagement that disembodied systems cannot. ### Social robots and human–robot interaction Two strands shape the social side of robotics in education. **Social robots** are robots designed to engage people through social interaction, using human-like cues such as speech, gesture, facial expression, and personality to communicate, teach, assist, or accompany. In education, social robots (humanoids like iCub, Pepper, Reachy, and companion robots) are used for tutoring, storytelling, role-play, language support, and as study companions. Their social presence is the key differentiator from software-based [[agentic-ai|AI agents]], enabling relationship-building and emotional engagement. Advances in [[llm|large language models]] have dramatically expanded what social robots can say and do, enabling fluent, adaptive conversational tutoring — while also introducing risks such as misinformation, [[cognitive-offloading|over-reliance]], and [[privacy]] violations, motivating knowledge-based design approaches. **Human–robot interaction (HRI)** is the interdisciplinary study of how people and robots interact, encompassing perception, communication, collaboration, and the social, cognitive, and [[ethics|ethical]] dynamics of that interaction. In education, HRI underlies how learners perceive, trust, and learn with robots — whether programming a robot, conversing with a tutoring robot, or rehearsing social scenarios. HRI [[research-methods-aied|research]] examines how robot appearance, behavior, task context, and embodiment shape [[usability-research|user experience]], trust, agency, and learning. Key concerns in educational HRI include preserving human [[agency]], building [[trust]], supporting [[self-efficacy]], and ensuring that interaction with robots supports rather than undermines autonomy and social learning. It connects robotics to [[human-ai-collaboration]] and [[social-emotional-learning]]. ### How robots are used in education [[teaching-with-robots-five-types-perspective-2026|Christ et al. (2026)]] add a role typology rather than a technology list: their five workshop-derived types of classroom robot differ by *pedagogical function and abstraction level*, not by hardware. Type a runs a non-interactive demonstration of generic social patterns (a scripted emotion theater followed by discussion of dynamics such as escalation or misunderstanding); type b is a touch-reactive interactive robot supporting participative physical theater, [[embodied-learning|embodied learning]], boundary awareness and emotion regulation; type c is a spoken-language partner that shows empathy and remembers interactions with a single pupil, creating a protected one-to-one setting for self-disclosure; type d is externally guided by a hidden specialist like a puppet with extra degrees of freedom, aimed at flattening social hierarchy; and type e is a non-interactive robot replaying actions recently observed in the school so pupils can reflect on situated behavior — the contrast with type a being exactly its context-specific rather than generalized abstraction. The typology is explicit about being an unvalidated design space grounded in one national mental-health program, so it is a menu for designing and evaluating robot roles, not evidence that any of them works. - **Computational thinking and programming:** Programmable robots (e.g., LEGO, block-based platforms) help learners connect code to real outcomes. [[computational-thinking-educational-robotics-secondary-2026|Valls i Pou]] links computational thinking to secondary STEAM curricula, and [[roboblockly-conversational-block-robotics-ct-2026|RoboBlockly Studio]] combines block programming with a [[conversational-ai|conversational AI]] agent and embodied robot feedback. [[edusim-llm-robotic-simulation-education-2026|EduSim-LLM]] lets beginners control simulated robots with natural language. - **Tutoring and knowledge delivery:** [[knowledge-based-design-generative-social-robots-2026|Knowledge-based design research]] and [[teachy-mini-generative-social-robot-higher-ed-2026|Teachy Mini]] develop LLM-powered generative social robots that tutor higher-education students, addressing risks like misinformation and [[cognitive-offloading|overreliance]]. [[task-context-trust-educational-hri-2026|Research on trust]] shows that what a robot does (task context) shapes learner trust more than its appearance, with the highest trust during instructional tasks. - **Storytelling and engagement:** [[motibo-digital-storytelling-robots-motivation-2026|MotiBo]] and [[robobuddy-llm-social-robots-classroom-2025|RoboBuddy]] use interactive, LLM-powered social robots for storytelling to boost motivation and engagement, while [[icub-humanoid-storytelling-llm-hri-2025|the iCub narrative HRI study]] explores co-creative storytelling between humans and humanoids. - **Social-emotional learning and [[inclusive-learning|inclusion]]:** [[remind-robot-mediated-roleplay-antibullying-2026|REMind]] uses robot-mediated role-play to rehearse anti-bullying bystander intervention, and [[pepper-robot-sign-language-lis-2025|work with the Pepper robot]] explores robot sign-language communication to support Deaf learners. [[pepper-social-robot-formal-education-scoping-review-2026|A scoping review]] maps Pepper's use in formal education. - **Autonomy and agency:** [[human-autonomy-agency-hri-review-2025|A systematic review]] synthesizes how HRI affects human autonomy and sense of agency, bridging design frameworks with [[regulation|regulatory]] demands (EU AI Act, IEEE Ethically Aligned Design). [[social-robot-study-companions|Social robots as study companions]] and [[enhancing-creative-writing-with-robot-llm-integration-the-interplay-of-embodimen|robot–LLM integration in creative writing]] further explore robot roles. - **Project-based and game-based approaches:** [[bots-blocks-project-based-robotics-education-2026|Bots and Blocks]] presents a project-based robotics course, and [[game-based-gamified-robotics-education-review-2026|a systematic review]] compares game-based learning and gamification in robotics education. - **Reinforcement learning and sim-to-real in a full robotics workflow.** [[teaching-rl-humanoid-robotics-high-school-2026|Dong, Cao and Wang (2026)]] convert an end-to-end research robotics workflow — assembly, electrical checks, simulation-based policy training, and physical deployment — into a [[k-12|high-school]] course built on one open humanoid (a ToddlerBot, reported parts cost under USD 6,000) across eight three-hour sessions. Pairs share one robot and train a [[reinforcement-learning|walking policy]] in [[simulation]] before deploying it to hardware, with safety gates (a passed standing test before walking) making the dependency order visible. Because a shared artifact rewards the team rather than the individual, the framework separates robot performance from individual understanding: students rotate roles, each submits a separate prediction and explanation at every checkpoint, and supported stepping is explicitly not treated as evidence of conceptual mastery — the authors' warning that passing a robot milestone is not understanding it. - **Competition robotics as an ecosystem problem, not a kit problem.** [[arc-hubs-k12-ai-robotics-rural-2026|Jacobson et al. (2026)]] locate the binding constraint on K–12 robotics less in curriculum or hardware than in sustained local technical mentorship, and show it is distributed geographically: in Indiana, FIRST LEGO League participation collapsed in the 2020 remote season, urban participation gradually recovered, and rural participation did not, remaining near its post-2020 level through 2025–2026. Their ARC framework makes mentorship the engineered object — colleges run a credit-bearing course that prepares undergraduates as workshop mentors for nearby teams, and mature school programs become secondary hubs whose experienced students become peer mentors for further schools, so reach propagates beyond any university's catchment through a self-reinforcing loop. A one-university trial created three rural FLL teams and moved undergraduate community connection from 1.86 to 4.00 on a five-point scale (the largest of any measured shift, ahead of confidence teaching technical concepts at +1.29), while a spatially explicit Markov simulation of Indiana's 1,925 public schools projected 992 school programs after 40 years under moderate assumptions against 161 with no ARC. The evidence is feasibility-level — seven mentors and four parents, retrospective self-reports, no control group — but the framing is portable to any robotics program: what scales or fails to scale is mentorship capacity and hub geography, not the robot ([[arc-hubs-k12-ai-robotics-rural-2026]]). - **Child development and young learners:** [[ai-toys-child-development-2026|AI-enabled toys and child development]] shifts the lens to commercial AI toys in early childhood, examining how AI-enabled playthings affect child development and play. This extends educational robotics beyond classroom robots to the consumer toys children encounter at home, raising questions about [[pedagogical-agent|agents]] in play, [[trust-calibration|trust calibration]], [[agency]], and [[well-being]] for the youngest learners — an area where design guidance is thinner than for school-age robotics curricula. - **Tangible coding and social robots in pre-K [[ai-literacy|AI literacy]].** Lee (2026) integrates unplugged play, tangible coding (Bee-Bot, Ozobot), and guided dialogue with a social AI robot in the Play With AI (PL-AI) curriculum for pre-K and kindergarten. The [[design-based-research|design-based research]] documents how these embodied, tangible robotics activities support children's emerging reasoning about AI concepts, with four design principles — embodied play, tangible coding, guided dialogue, and teacher co-design — offering a developmentally appropriate model for [[early-childhood-elementary-ai-education|early childhood]] robotics and AI education. - **Two paradigms for young learners: coding robots and generative social robots.** [[creative-project-approach-ai-early-childhood-2025|Yang, Li and Lee (2025)]] frame early-childhood robotics as the pairing of two [[pedagogy|pedagogical]] paradigms, each with a distinct theoretical base. **Coding robots** (Bee-Bot, KIBO, Matatalab) descend from Papert's LOGO and embody [[constructivist|constructionism]] — children learn by making and build [[computational-thinking|computational thinking]] through tangible programming. **Generative social robots**, powered by [[generative-ai|generative AI]], are grounded in [[sociocultural-learning|social constructivism]], acting as conversational peers or tutors who [[scaffolding|scaffold]] learning within the child's Zone of Proximal Development and support social-emotional development. Their five-step **Creative Project Approach** for integrating both robot types into the Project Approach keeps teachers as facilitators who guide child–robot interaction, balance automation with [[creativity]], and preserve child [[agency]]. ### Embodiment and pedagogy A defining theme is that robots are effective when they support genuine learning goals — not as isolated technical exercises. The value of a robot depends on the pedagogical context: teaching computational thinking ([[computational-thinking]]), supporting [[stem-education|STEAM]], building [[cs-education|programming]] skills, motivating learners ([[motivation]], [[student-engagement|engagement]]), or supporting [[social-emotional-learning]] and [[equity-in-ai-education|inclusion]]. Robotics also connects to [[project-based-learning]], [[game-based-learning]], and [[experiential-learning]]. Key design considerations include preserving learner [[agency]], building [[trust]], supporting [[self-efficacy]], and grounding learning in [[embodied-learning|embodied interaction]]. In [[language-learning]], [[robot-assisted-language-learning-meta-analysis-2026|meta-analytic evidence]] points to the effectiveness of embodied robot-assisted language learning. - **Pathways to learning AI-powered robotics.** [[educational-robotics-pathways-2026|A qualitative study]] of high school students in a robotics+AI curriculum found learning through real-world practice, designing, and playful creative expression (constructionist, epistemological-pluralist lens). ## Connected Concepts - [[early-childhood-elementary-ai-education]] — Early childhood and elementary AI education - [[computational-thinking]] - [[cs-education]] - [[stem-education]] - [[embodied-learning]] - [[human-ai-collaboration]] - [[project-based-learning]] - [[game-based-learning]] - [[llm]] - [[motivation]] - [[student-engagement]] - [[social-emotional-learning]] - [[agency]] - [[trust]] - [[well-being]] - [[ethics]] - [[privacy]] - [[language-learning]] - [[k-12]] - [[higher-ed]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) ## Connected Articles - [[pepper-social-robot-formal-education-scoping-review-2026]] — Scoping Review of the Pepper Robot in Formal Education - [[robot-assisted-language-learning-meta-analysis-2026]] — Meta-analysis of AI-enhanced embodied robot-assisted language learning - [[white-wu-robotics-ai-education-2026]] — Robotics and AI in Education - [[computational-thinking-educational-robotics-secondary-2026]] — Computational Thinking and Educational Robotics - [[roboblockly-conversational-block-robotics-ct-2026]] — RoboBlockly Studio - [[edusim-llm-robotic-simulation-education-2026]] — EduSim-LLM - [[knowledge-based-design-generative-social-robots-2026]] — Knowledge-Based Design for Generative Social Robots - [[teachy-mini-generative-social-robot-higher-ed-2026]] — Teachy Mini - [[motibo-digital-storytelling-robots-motivation-2026]] — MotiBo - [[robobuddy-llm-social-robots-classroom-2025]] — RoboBuddy - [[remind-robot-mediated-roleplay-antibullying-2026]] — REMind - [[task-context-trust-educational-hri-2026]] — Task Context and Trust in Educational HRI - [[human-autonomy-agency-hri-review-2025]] — Human Autonomy and Agency in HRI - [[icub-humanoid-storytelling-llm-hri-2025]] — iCub Narrative HRI - [[pepper-robot-sign-language-lis-2025]] — Pepper and Sign Language - [[social-robot-study-companions]] — Social Robots as Study Companions - [[enhancing-creative-writing-with-robot-llm-integration-the-interplay-of-embodimen]] — Robot-LLM Integration in Creative Writing - [[game-based-gamified-robotics-education-review-2026]] — Game-Based and Gamified Robotics Education - [[bots-blocks-project-based-robotics-education-2026]] — Bots and Blocks - [[educational-robotics-pathways-2026]] — Pathways to Learning AI-Powered Educational Robotics (2026) - [[tsingidou-ct-robotics-kindergarten-2026]] — Robot-mediated CT in kindergarten - [[ai-toys-child-development-2026]] — AI-enabled toys and child development - [[play-ai-pre-k-kindergarten-ai-literacy-2026]] — Play With AI (PL-AI): play-centered AI literacy curriculum for pre-K and kindergarten (Lee 2026) - [[creative-project-approach-ai-early-childhood-2025]] — The Creative Project Approach: integrating coding and generative social robots into early-childhood projects (Yang, Li & Lee 2025) - [[teaching-with-robots-five-types-perspective-2026]] — Five functionally distinct types of classroom robot, from scripted demonstration to one-to-one empathic dialogue (Christ et al. 2026) - [[arc-hubs-k12-ai-robotics-rural-2026]] — ARC: a hubs-based framework that treats technical mentorship capacity and hub geography, not hardware, as the constraint on rural K–12 robotics programs (Jacobson et al. 2026) - [[teaching-rl-humanoid-robotics-high-school-2026]] — Teaching Reinforcement Learning and Humanoid Robotics to High-School Students: An Expert-Validated Curriculum Design on a Low-Cost Open Platform --- ## [Conversational AI](https://edtechdev.github.io/aied/concepts/conversational-ai/) > **Conversational AI (CAI) agents** — AI-driven speech- or text-based agents that simulate and automate conversations, from rule-based chatbots to NLP/ML and [[multimodal]] LLM-based assistants — are among the most widely used AI interfaces in education, valued for [[teacher-role|teaching]], psychological, and metacognitive support even as technical, cognitive, and [[ethics|ethical]] concerns persist. ## Questions to Consider - When you've used a chatbot or AI assistant, did you think of it as a teacher, a search engine, or something else? How did that framing shape how much you actually learned from it? - Conversational AI describes HOW an agent talks, not what it's built to do. Could a chatbot be conversational yet completely unpedagogical — and what would tell you the difference? - One study found that AI literacy — not general tech savvy — predicted whether students were willing and able to use a chatbot. Why might knowing how AI works matter more than being 'good with computers'? - A raw general chatbot can short-circuit reasoning by answering immediately, while a structured tutor preserves productive struggle by withholding answers. What design choices decide which kind of agent a student meets? - Students using a law-course chatbot did a third of their interactions after hours — evidence that 24/7 availability is a real benefit. But does always-on access also carry risks you'd want to design against? - The biggest barrier to chatbot adoption in one study was a badly designed pop-up, not distrust or academic-integrity fears. What does that suggest about where AI-in-education investments actually fail? ## Introduction Conversational AI (CAI) is the umbrella term for AI-driven agents that carry on spoken or written dialogue, most commonly realized as chatbots and, more recently, [[generative-ai|generative]] [[llm]]-based assistants such as ChatGPT, Claude, and multimodal educational avatars. Modern CAI agents fall into [[machine-learning]]-based, NLP-based, and hybrid categories, with text-based agents the most prevalent in education. As learning tools they function as [[intelligent-tutoring|intelligent tutors]], [[feedback]] providers, [[student-ai-interaction|interaction partners]], and administrative assistants — overlapping with [[pedagogical-agent|pedagogical agents]] while spanning a broader set of applications. ## How conversational AI appears in the knowledge base **An umbrella-review synthesis.** The [[conversational-ai-agents-umbrella-review-2026|umbrella review of CAI agents]] (34 review articles) shows CAI utilization is concentrated in teaching and learning support (97.1% of reviews), psychological and motivational support (91.2%), and [[metacognition|metacognitive]] and personal development (88.2%), while administrative support, [[research-methods-aied|research]] management, and healthcare education lag. The review documents that human–AI relationship concerns persist across all CAI generations, with [[academic-integrity]] and data [[privacy]] emerging as newer ethical issues, and calls for HCI-grounded, evidence-based design and stronger [[ai-literacy]] support. **From chatbots to tutoring agents.** The knowledge base traces CAI's evolution from rule-based FAQ chatbots toward [[intelligent-tutoring|tutoring-focused]] [[pedagogical-agent|agents]]. The [[conversational-ai-tutors-framework|conversational AI tutors framework]] argues proven ITS [[ai-technologies|technologies]] ([[knowledge-tracing]], affect detection, [[student-modeling|student modeling]]) should anchor generative tutors while [[generative-ai]] supplies flexible dialogue. Research on [[measuring-llm-tutors-teach-vs-solve|whether LLM tutors teach or solve]] and [[stanford-evidence-base-ai-k12-2026|tutoring-specific vs general AI]] shows pedagogically designed [[guardrails]] matter: raw general chatbots can short-circuit reasoning while structured tutors preserve [[desirable-difficulties|productive struggle]]. **Interaction and collaboration.** Conversational agents are increasingly framed as interaction partners rather than answer-givers. [[student-ai-interaction]] captures how learners prompt, question, and verify with CAI in practice. In [[collaborative-learning]], agents mediate participation and shared [[regulation]], and in [[language-learning]] they provide real-time conversational practice. The [[human-ai-collaboration]] thread examines when this partnership preserves versus substitutes for the learner's cognitive work. LLMs as critique partners illustrate the preservation side concretely: [[oppenheimer-llms-collaborative-learning-partners-2026|Oppenheimer, Cash & Connell Pensky (2025)]] had ChatGPT, Gemini, or Claude critique students' argumentative essays across a semester, and learners improved on writing, [[prompt-engineering|prompt engineering]], and response-to-feedback while rating the exchanges useful, engaging, and enjoyable — with active rebuttal of model claims (87.8%) showing they treated the conversational partner critically rather than passively. **The role the agent plays shapes the interaction.** CAI design is not neutral about its persona: [[liao-role-adaptive-ai-companion-book-talk-2026|Liao (2026)]] found a fixed "student peer" companion in elementary book talk sustained longer interactions yet dominated the conversation (lower student word/sentence share) and hit an "affective ceiling" — matching a human teacher on factual recall but falling short on emotional and future-oriented reflection — arguing CAI should adapt its role (peer, teacher assistant, parent advisor) rather than stay monolithic. [[xu-genai-collaborative-space-2026|Xu et al.]] extend this to small groups, showing generative AI acts as both an *agent* and a *collaborative space* in synchronous and asynchronous collaborative dynamics, where the interaction design decides whether it scaffolds or supplants group cognition. Tone can also be set by an external classifier rather than by the conversation alone: [[culturally-aware-student-stress-chatbot-2026|Sukoon (Bashir & Afzal, 2026)]] trains a Random Forest on 20 survey features to assign a low, moderate, or high stress level, and that output selects one of three response tiers inspired by the Stepped Care Model — warm and encouraging at low stress, grounding and non-judgmental at high stress — before an [[open-source]] LLM (GLM-4.5-Air via OpenRouter) takes over the dialogue with the full conversation history and a culturally adapted system prompt re-sent on every turn. The authors chose a free-access [[multilingual-learning|multilingual]] model deliberately to keep deployment feasible in regional universities, and note that Urdu appears through prompt engineering rather than a genuinely bilingual pipeline — a limitation that matters whenever cultural appropriateness is claimed as a system property rather than an evaluated outcome. **Student perspectives.** Real-world usage shows adoption hinges on [[ai-literacy]] and [[usability-research|user experience]] more than on technical capability. A human-centered [[mixed-methods-research|mixed-methods]] study of the "Jordan Chatbot," a GPT-4o-based [[pedagogy|pedagogical]] agent in an Australian law course, found students hold positive attitudes and perceive gains in knowledge while strongly supporting [[academic-integrity]] requirements; over a third of interactions occurred after hours, confirming the value of 24/7 availability ([[colbran-student-perspectives-genai-chatbots-2026|Colbran, Jha & Schiavone 2026]]). Notably, AI literacy — not general technology proficiency — predicted willingness and confidence to use the chatbot, and usability (an intrusive pop-up design) was the largest barrier among non-users, ahead of trust, preference for staff, and academic-integrity fears.([[colbran-student-perspectives-genai-chatbots-2026]]) The study recommends human-centered design, explicit AI policies and assessment labels, staff and student training, and continuous error monitoring — evidence that effective CAI deployment is as much a design and literacy problem as a technical one. At the other end of the age spectrum, [[vahedian-children-attitudes-ai-chatbot-2026|Vahedian Movahed & Martin (2025)]] studied children aged 6–14 interacting with AMA, a topic-bounded, age-tailored chatbot (astronomy, sneakers and shoes, dinosaurs), finding broad openness to and high trust in the AI as an information source — children even tested its credibility with known-answer questions — alongside gaps in critical engagement and digital-safety awareness that argue for age-sensitive, trust-aware conversational-ai design and explicit [[privacy]] instruction. **Risks and ethics.** CAI agents carry persistent risks of [[cognitive-offloading|over-reliance]] and [[cognitive-offloading|cognitive offloading]] (the leading ethical concern in the umbrella review), plus technical limitations, [[hallucination-risk|hallucination]], bias, [[ai-detection|plagiarism]], and [[equity-in-ai-education|equity]] barriers. These concerns animate [[ai-literacy]] and [[reducing-ai-misuse]] and require [[educational-policy-ai|policy]] and ethical-[[governance]] responses. A further trust concern arises around the adoption advice staff increasingly receive: staff are urged to consult conversational AI about whether to adopt AI, yet such systems are built by organizations with a commercial stake in adoption. An audit of ten frontier LLMs found most acknowledge skeptical users' concerns before redirecting to engagement framings, raising questions about the neutrality of AI adoption advice. **Game-based conversational agents.** Beyond tutoring, conversational agents are being embedded in digital game-based learning. Wenzel, Geiger, and Liening (2026) use action design research to derive the **CAIS-GBL** framework — four design principles and fifteen design features for AI conversational agents in digital game-based learning — grounded in theory-driven meta-requirements spanning cognitive, motivational, [[affective-computing|affective]], and [[sociocultural-learning|socio-cultural]] [[student-engagement|engagement]] and an equity-by-design stance. Their instantiated agent (Lara) in a business [[simulation]] game was positively received for cognitive and [[community-of-inquiry|social presence]] and support for [[self-regulated-learning|self-regulated learning]], evaluated with student teachers and in a field study — a practical blueprint for [[adaptive-learning|adaptive instructional support]] via [[game-based-learning|conversational agents in serious games]]. ## Relationship to pedagogical agents and intelligent tutoring Conversational AI is best understood as an **interaction modality** that overlaps — but does not coincide with — two more established constructs in the knowledge base: [[pedagogical-agent|pedagogical agents]] and [[intelligent-tutoring|intelligent tutoring systems (ITS)]]. **Conversational AI as the medium, not the pedagogy.** CAI names *how* the agent communicates (natural-language dialogue, spoken or text). It says little on its own about *what* the agent is built to do. Pedagogical agents, by contrast, are defined by their **instructional role** — an AI component that engages learners through dialogue, questions, or prompts to support [[metacognition|metacognitive processes]], [[feedback]], and [[scaffolding]]. [[intelligent-tutoring|Intelligent tutoring systems]] are defined by their **architecture and modeling** — a diagnostic backbone of [[knowledge-tracing]], [[student-modeling|student modeling]], and pedagogical decision logic that tracks what the learner knows and adapts instruction. A single agent can be all three at once: e.g. a [[conversational-ai-tutors-framework|conversational AI tutor]] is a CAI agent (dialogue interface) that functions as a pedagogical agent (tutoring strategies) built on an ITS foundation (student modeling). The distinction matters because a CAI agent need not be pedagogically grounded at all — a plain FAQ chatbot is conversational AI without being a pedagogical agent or a tutor. **The pedagogical-agent lens.** Pedagogical agents use the conversational medium to enact teaching strategies — eliciting self-assessments, [[socratic-method|Socratic questioning]], role-specialized facilitation in [[agentic-ai|multi-agent]] designs (Teacher, Assistant, Classmate, Analyzer). Not every CAI agent is a pedagogical agent, but the two heavily overlap: the umbrella review of CAI agents found teaching and learning support (97.1%) and metacognitive development (88.2%) dominate CAI applications, meaning most education-focused CAI agents function pedagogically. The [[conversational-agents-novice-programmers-scoping-2025|novice-programmer scoping review]] sharpens this: only 4 of 23 conversational agents explicitly grounded design in [[learning-theories|learning theory]] — most were pedagogical in intent but not in foundation. **The ITS lens.** Intelligent tutoring contributes the *[[cognitive-diagnosis|cognitive diagnostic]] machinery* that raw conversational models lack. The [[conversational-ai-tutors-framework|conversational AI tutors framework]] argues proven ITS technologies should anchor generative tutors: knowledge tracing, affect detection, and student modeling supply the structure, while [[generative-ai]] and [[llm|LLMs]] supply flexible dialogue. This is the key design tension — conversational AI provides natural, scalable interaction, but without ITS-style structure it risks [[cognitive-offloading|over-scaffolding]], hallucination, or bypassing the learner's productive struggle. Research such as [[measuring-llm-tutors-teach-vs-solve]] and [[stanford-evidence-base-ai-k12-2026]] shows that pedagogy-oriented criteria (guiding questions, calibrated hints) must be designed in explicitly. **In short:** conversational AI is the **interface/medium**, pedagogical agents are the **role**, and intelligent tutoring is the **underlying modeling and instructional logic**. Educationally valuable CAI agents sit at the intersection of all three — conversational in interface, pedagogical in intent, and tutor-like in their modeling of the learner. ## Practical guidance Choose conversational agents to support teaching, [[motivation]], and [[metacognition]] rather than merely to answer questions, and design for HCI-grounded, participatory, user-centered interaction. Guard against [[cognitive-offloading|over-reliance]] by pairing CAI with [[ai-literacy]] instruction and [[feedback]] that keeps the learner cognitively productive. Attend to AI literacy and usability explicitly — since these — not general digital skill — drive adoption and non-use ([[colbran-student-perspectives-genai-chatbots-2026|Colbran, Jha & Schiavone 2026]]) — and pair deployment with clear AI-use policies, assessment labels, and training. Evaluate CAI on pedagogical outcomes — not just task completion — and plan for equity and [[accessibility]] from the start rather than as an afterthought. ## Connected Concepts - [[intelligent-tutoring]] - [[pedagogical-agent]] - [[generative-ai]] - [[llm]] - [[student-ai-interaction]] - [[ai-literacy]] - [[human-ai-collaboration]] - [[cognitive-offloading]] - [[feedback]] - [[metacognition]] - [[self-regulated-learning]] - [[language-learning]] - [[academic-integrity]] - [[hallucination-risk]] - [[equity-in-ai-education]] - [[reducing-ai-misuse]] - [[speech-and-voice-technologies]] - [[parents-and-families]] ## Connected Articles - [[usher-faraon-who-grades-best-2026]] — Comparing ChatGPT, peer, and instructor grading across project quality levels (Usher & Faraon 2026) - [[llm-agents-5e-esl-grammar-2026]] — LLM agents with 5E framework for ESL grammar acquisition (Yang et al. 2026) - [[llm-interaction-depth-task-quality-recall-2026]] — What students ask matters: LLM interaction depth, task quality, and immediate recall (Tsiligkiris 2026) - [[semantic-variability-llm-conversation-assessment-2026]] - [[colbran-student-perspectives-genai-chatbots-2026]] — Student perspectives on GenAI chatbots (mixed methods) - [[saihi-ahmed-genai-adoption-personas-higher-ed-2026]] — Adoption personas for AI chatbots - [[conversational-ai-agents-umbrella-review-2026]] — Umbrella review of conversational AI agents in education - [[conversational-ai-tutors-framework]] — Conversational AI tutors framework - [[measuring-llm-tutors-teach-vs-solve]] — Measuring whether LLM tutors teach or solve - [[stanford-evidence-base-ai-k12-2026]] — Tutoring-specific vs general AI - [[rethinking-scaffolding-llm-tutors]] — Rethinking scaffolding in LLM tutors - [[genai-higher-education-systematic-review-2026]] — GenAI in higher education systematic review - [[conversational-agents-novice-programmers-scoping-2025]] — Scoping review of conversational agents for novice programmers - [[dai-chatbots-problem-posing-primary-2026]] — GenAI chatbots and problem posing in primary science - [[ba-ai-agents-cscl-review-2026]] — AI agents in computer-supported collaborative learning review - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[aivaluate-anxiety-assessment-2026]] — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026) - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — The substitution-to-scaffolding AI harm cycle - [[lee-wu-gender-motivation-genai-achievement-2026]] — Gender and motivation in GenAI achievement - [[conversational-agents-business-simulation-gaming-2026]] — CAIS-GBL framework for AI conversational agents in business simulation games (Wenzel et al. 2026) - [[frontier-ai-redirect-skeptical-rural-staff-2026]] — Algorithmic audit: how frontier LLMs redirect skeptical rural K-12 staff - [[liao-role-adaptive-ai-companion-book-talk-2026]] — Role-adaptive AI companion for elementary book talk; affective ceiling of fixed-role agents (Liao 2026) - [[xu-genai-collaborative-space-2026]] — GenAI as agent and collaborative space in small-group dynamics (Xu et al. 2026) - [[oppenheimer-llms-collaborative-learning-partners-2026]] - [[vahedian-children-attitudes-ai-chatbot-2026]] - [[culturally-aware-student-stress-chatbot-2026]] — An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning - [[sophie-clinical-communication-ai-assessment-2026]] — Scalable AI-based clinical communication training and automated assessment ## Citation Ganguly, A., Mehjabin, N., Malik, A., & Johri, A. (2025). [*Conversational AI agents in education: an umbrella review*](https://doi.org/10.1007/s43681-025-00916-0). *AI and Ethics*, 6, 72. --- ## [Simulation](https://edtechdev.github.io/aied/concepts/simulation/) > **Simulation** — the use of modeled environments, agents, or scenarios to support learning through practice and feedback in contexts that are safe, repeatable, and often otherwise inaccessible. Simulations let learners act, make errors, and see consequences without real-world cost, and are increasingly powered by AI and agent-based modeling. ## Questions to Consider - Recall a time you learned something by doing it in a safe, low-stakes environment — a lab, a mock exercise, a flight or game simulator. What made that practice effective, and what might be lost if the simulation were too realistic or not realistic enough? - The page argues simulations let learners make errors and see consequences 'without real-world cost.' What do you think is gained, and what might be lost, when the cost of a mistake drops to nearly zero? - If an AI can simulate patients, students, or conversation partners for practice, where would you draw the line between valuable rehearsal and practice that fails to transfer to real human interaction? - Why might a learner's awareness of a simulation's limits — its [[trust|trustworthiness]] — matter as much as how faithfully it models reality? - How could the same simulation technology that helps someone learn also mislead them, and what would you need to know to tell those two outcomes apart? ## Introduction Simulation sits at the core of [[experiential-learning|experiential]] and [[active-learning]] [[pedagogy|pedagogies]]. It provides the deliberate practice, [[productive-failure|productive failure]], and [[feedback|feedback loops]] that build skill and judgment. AI has transformed simulation in two ways: it powers more realistic and adaptive simulated environments, and it generates [[simulating-students|simulated learners]], patients, or interlocutors that make practice scalable. Behavioral evidence shows that *how* learners engage with a simulation varies systematically rather than uniformly: tracing online learners building ecological models in VERA, [[an-goel-self-directed-modeling-2026|An, Hammock & Goel (2025)]] classified [[student-engagement|engagement]] into Observation (frequent runs and parameter adjustment with little model building), Construction (hands-on building with little simulation), and Exploration (full construct–parameterize–simulate cycles), with Explorers producing the most complex and diverse models and observation-heavy learners largely copying existing ones — an argument for designing simulation environments that push learners toward full-cycle activity. ### AI and simulation - **AI-powered environments:** adaptive simulations adjust difficulty and scenarios to a learner's state, linking to [[adaptive-learning]] and [[reinforcement-learning]]-based coaching. - **Simulated agents:** AI can simulate patients (for medical training), students (for [[teacher-role|teacher]] practice), or conversation partners, making high-stakes interpersonal practice accessible and repeatable. In [[teacher-education|teacher education]], [[zhuang-zhang-chatgpt-math-teacher-education-2026|Zhuang and Zhang (2025)]] built *Student GPT*, a custom ChatGPT [[conversational-ai|chatbot]] that role-played a [[k-12|middle school]] student holding common ratio-reasoning [[misconceptions]], giving preservice [[math-education|mathematics]] teachers affordable, content-specific practice at diagnosing student thinking — and used an [[affective-computing|Affective]], Communicative, Technical (ACT) coding framework to systematically assess the simulated student's role-play strengths (clarity, relevance, error consistency) and authenticity weaknesses (teacher-like tone, role confusion). - **Role-play puts the learner in the part.** Where simulated agents supply the counterpart, role-play gives the learner that part instead. [[remind-robot-mediated-roleplay-antibullying-2026|Sanoubari and colleagues (2026)]] had 18 children aged 9-10 watch a bullying scene enacted by social robots, reason about each character's position, then rehearse defending by puppeteering a robotic avatar, and reported gains in perceived [[self-efficacy]] for defending plus better-calibrated beliefs about whether confronting a bully actually stops it. Their framing, robot-mediated applied drama, keeps a human facilitator in the Forum Theatre role and confines automation to narrative control, which is a useful reminder that the demanding part of role-play is the reflection rather than the machinery. [[lock-integrating-ai-online-learning-higher-ed-2025|Lock, Arteaga and Johnson (2025)]] place role-play alongside simulation among the strategies that AI-supported online learning draws on. - **Simulated learners:** models of student behavior let [[research-methods-aied|researchers]] and designers test tutoring systems and [[curriculum-design|curriculum]] before live deployment, grounding [[student-modeling]] and [[knowledge-tracing]]. - **Trust and fidelity:** the value of a simulation depends on how faithfully it models the real context — and on the learner's awareness of its limits, connecting to [[trust-calibration]]. - **[[generative-ai|GenAI]] in simulation-based learning.** [[genai-scenario-based-healthcare-education-2026|Neto and colleagues (2026)]] [[meta-analysis-systematic-review|systematically review]] GenAI across scenario-, case-, problem-, and simulation-based learning in healthcare education, finding positive outcomes for higher-order cognitive skills but inconsistent results elsewhere, with hybrid [[human-ai-collaboration|human-AI collaboration]] outperforming fully automated approaches. [[conversational-agents-business-simulation-gaming-2026|Wenzel, Geiger, and Liening (2026)]] develop AI conversational agents for adaptive support in business simulation games, addressing the common gap of limited [[formative-assessment|formative]] feedback and structured reflection in simulation-based learning. - **The "authenticity gap" bounds what AI simulation can replace.** In [[medical-education|clinical]] simulation, [[jiang-ai-powered-simulation-nursing-education-2026|Jiang et al. (2026)]]'s [[mixed-methods-research|mixed-methods]] systematic review of AI-powered nursing simulation (19 studies, N=1,253) finds AI effective for cognitive knowledge and affective outcomes but inconsistent for complex psychomotor skills. Their concept of an **authenticity gap** — a learner-perceived shortfall in emotional resonance, nonverbal cue recognition, and tactile/physical examination dimensions — explains *why* AI simulation is best for highly structured objectives (foundational communication, history-taking) and should sit in a **stepped simulation continuum** that hands advanced psychomotor and emotionally complex scenarios to human-standardized patients and clinical placement. Technical instability (e.g., speech-recognition delays) can also add extraneous [[cognitive-offloading|cognitive load]] and anxiety, so fidelity and stability are themselves design levers. This parallels [[genai-scenario-based-healthcare-education-2026|Neto et al.'s]] finding that hybrid human–AI approaches outperform fully automated ones. - **Teacher-AI co-designed simulations.** Interactive simulations that support both conceptual learning and competency development are scarce in hands-on domains, and GenAI output often lacks pedagogical validity. In [[stem-education|drone-based STEM education]], teacher-AI co-designed simulations embedded in an otherwise identical hands-on curriculum were evaluated with a quasi-experimental pretest–posttest design across 30 secondary students, examining whether simulation-supported instruction yields superior [[learning-gains|learning outcomes]] ([[simulation-assisted-drone-learning-stem-2026]]). Separately, [[agentic-ai|multi-agent]] tutoring [[benchmark|benchmarks]] such as ASTRA use simulated socially intelligent agents to study participation-balanced collaboration in [[cs-education|introductory programming]] ([[astra-multi-agent-tutoring-benchmark-2026]]). - **Learner control in simulation is enacted, not granted.** A 2 × 2 experiment in a flocking simulation ([[learner-agency-ai-simulation-2026|Su, Nair and Nagashima 2026]]) gave some students parameter sliders, some an optional conversational agent and some both; every condition improved, but neither affordance produced a reliable difference once prior knowledge was controlled (p = .849 and p = .108). What predicted [[learning-gains|gains]] was where and how long learners manipulated parameters: sustained slider use in the most conceptually complex lesson was positively associated with gains, and the same behavior in the easier lesson negatively. For simulation builders the implication is that offering controls is not the intervention — helping learners decide what to change, and register what changed, is. ### Connections Simulation connects to [[active-learning]], [[adaptive-learning]], and [[pedagogical-agent]]. It is a mechanism for experiential and [[constructivist]] learning and is amplified by AI's ability to generate adaptive, realistic practice environments. ## Connected Concepts - [[active-learning]] - [[adaptive-learning]] - [[pedagogical-agent]] - [[reinforcement-learning]] - [[student-modeling]] - [[constructivist]] - [[trust-calibration]] - [[professional-training]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) - [[virtual-and-augmented-reality]] — the model, not the modality — immersive environments usually render a simulation ## Connected Articles - [[learner-agency-ai-simulation-2026]] — Parameter control and an optional AI agent in a complex-systems simulation: gains tracked enactment, not access - [[benzion-ai-physics-simulations-virtual-lab]] - [[genai-simulate-patient-history-pbl-2026]] - [[alrazeeni-transforming-nursing-education-ai-2026]] — AI in nursing education: systematic review (simulation, assessment) - [[adaptive-virtual-patient-psychotherapy-training]] — Adaptive Virtual Patients for Psychotherapy Training - [[ai-enabled-serious-games]] — AI-Enabled Serious Games - [[anvil-ai-educational-animations]] — ANVIL: Analogies and Videos for Lecturers - [[astra-atco-training-simulator]] — ASTRA: ATCO Training Simulator - [[supplynet-visual-exploratory-learning]] — SupplyNet: Visual Exploratory Learning - [[medeasy-ai-standardized-patients]] — MedEASY: AI Standardized Patients - [[remind-robot-mediated-roleplay-antibullying-2026]] — Robot-mediated role-play game for bystander intervention (applied drama) - [[hdr-brachytherapy-agentic-ai-simulation-2026]] - [[residencyrl-clinical-rl-training-2026]] - [[li-ai-science-situated-learning-teachers-2025]] - [[ai-science-chemistry-education-systematic-review-2025]] — Systematic review of AI in science/chemistry education - [[context-based-ai-secondary-chemistry-2026]] — Context-based 7E + AI instruction in secondary chemistry - [[chatgpt-virtual-lab-teaching-assistant-biology-2026]] — ChatGPT as a virtual lab teaching assistant in biology - [[educasim-cs1-instructional-practice]] — EducaSim: simulated small-group section for teacher practice - [[genai-scenario-based-healthcare-education-2026]] — Systematic review of GenAI in scenario-based healthcare education (Neto et al. 2026) - [[conversational-agents-business-simulation-gaming-2026]] — CAIS-GBL framework for AI conversational agents in business simulation games (Wenzel et al. 2026) - [[llm-agents-collaborative-problem-solving-simulation-2026]] — Fine-tuned participant-specific LLM agents reproducing collaborative problem solving dialogues (Fang 2026) - [[astra-multi-agent-tutoring-benchmark-2026]] — ASTRA synthetic benchmark for multi-agent tutoring and participation-balanced collaboration - [[simulation-assisted-drone-learning-stem-2026]] — Simulation-assisted drone learning with teacher-AI co-designed scaffolds - [[an-goel-self-directed-modeling-2026]] - [[zhuang-zhang-chatgpt-math-teacher-education-2026]] - [[jiang-ai-powered-simulation-nursing-education-2026]] — AI-powered simulation in nursing: mixed methods systematic review (authenticity gap, stepped continuum) - [[sophie-clinical-communication-ai-assessment-2026]] — Scalable AI-based clinical communication training and automated assessment - [[shi-genai-experiential-learning-management-education-2026]] — a dynamic business simulation in which the model generates disruptive events mid-decision --- ## [Virtual and Augmented Reality](https://edtechdev.github.io/aied/concepts/virtual-and-augmented-reality/) > **Virtual and Augmented Reality (VR/AR)** — the display and interaction layer through which learning environments are experienced: fully synthetic spaces in VR, and digital content overlaid on the physical world in AR and mixed reality. AI enters this layer in two directions. It authors it, since [[generative-ai|generative AI]] now turns a natural-language description into a working browser-based AR or VR learning tool that no longer requires a specialist developer. And it inhabits it, as [[agentic-ai|agents]], [[pedagogical-agent|pedagogical agents]] and [[rag|retrieved knowledge]] guide a learner hands-free inside the immersive environment. What the modality adds over a screen is presence and [[embodied-learning|embodiment]]; what it costs is fidelity, hardware, and a body's tolerance for being there. ## Questions to Consider - Where does the line fall between what is being modeled and the surface it is shown on? Compare a desktop patient simulator with a VR field trip to a site students cannot visit. Which differences would you expect to change learning, and which might only be novelty? - A pilot found students reported *feeling* wavelength and amplitude through a hand gesture more strongly than through a slider — but it measured perception, not achievement, with 29 students and no comparison group. How much should reported "feeling" count toward adopting a tool? - If generative AI lets an instructor with no programming background build a working AR simulation in an afternoon, what new responsibilities follow for validating the [[physics-education|physics]], judging fidelity, and deciding whether it belongs in a course? - In a review of AI-powered nursing simulation, AI matched human actors for structured communication but not for tactile and emotionally complex scenarios. How would you sequence practice so learners rehearse some parts with AI and others with people? - One classroom VR study reported minimal motion sickness while the meta-analytic estimate for intelligent VR with students with disabilities was not statistically significant. What would you want measured before a program invests in headsets? ## Introduction Virtual and augmented reality is a **modality**, not a model: it is the layer through which a learning environment reaches the learner. That makes it a different axis from [[simulation]], which is what gets modeled in the first place. The two are often conflated, because immersive environments are common delivery vehicles for simulations, but they vary independently — a desktop patient simulator is simulation without VR, and an AR overlay on a real instrument is VR without simulation. Keeping the axes apart matters for design: deciding to make practice risk-free is a different decision from deciding to make it embodied, and the two carry different costs, different failure modes, and different evidence. VR/AR is the **delivery layer** for a simulation and a close relative of [[game-based-learning|game-based learning]]; it operationalizes [[embodied-learning|embodied]] and [[situated-learning|situated]] accounts of learning, and it sits inside [[experiential-learning|experiential]] and [[active-learning]] practice. AI now shapes this modality from both ends. It **authors** immersive content, because [[generative-ai|generative AI]] converts a structured natural-language prompt into a runnable AR or VR artifact, collapsing the specialist programming skill that used to gate production. And it **inhabits** the environment, as [[agentic-ai|agents]], [[pedagogical-agent|pedagogical agents]] and [[rag|retrieved domain knowledge]] supply guidance inside it — hands-free, in real time, grounded in sources the learner cannot consult while wearing a headset. ## What the modality adds The case for VR/AR rests on presence and [[embodied-learning|embodiment]] rather than on information delivery. A shared virtual space can also decouple co-presence from geography: a VR classroom built on a real-time synchronisation layer put teachers and students in the same room and manipulating the same 3D science materials regardless of site, targeting up to twenty participants. In a one-hour session with 10 students on a Meta Quest 3, it reported good [[usability-research|usability]] on the System Usability Scale and minimal [[simulation]] sickness, while its own weakness was UI consistency — and its designers deliberately avoided the earlier peer-to-peer architectures whose frame rate degraded as participants joined, a reminder that presence is bounded by engineering. More controlled comparisons are sobering about what the surface alone buys: a 24-participant study of engineering mechanics found that mixed-reality apps and physical toolkits raised [[student-engagement|engagement]] over classroom instruction, yet complex visualizations remained difficult for learners in every condition. Engagement is the reliable effect; understanding is not. ## Generative AI as the authoring layer The barrier that used to define who could build an immersive tool has largely fallen. Using a four-element prompt structure — **tools, display, hand controls, optimization** — a [[teacher-role|teacher]] or student with no coding background can generate a browser-based, hand-controlled AR physics simulation that runs as a single HTML file with nothing beyond a camera, then refine it by describing what went wrong in plain language. The gesture is familiar from touchscreens: pinch and spread tunes a physical quantity rather than zooming a picture, so opening the fingers vertically raises amplitude, and horizontally lengthens wavelength and shifts the lamp toward the red end of the spectrum. The same structure generalizes: the Coulomb field around a fingertip spread through the room, two hands as opposing charges, and the right-hand rule for magnetic force drawn on the learner's own hand — rendered deliberately **unmirrored**, since mirroring would invert the very rule being taught. The pilot is encouraging and preliminary in equal measure. With 29 second-year medical-imaging students in an introductory radiation physics course, all 29 agreed the gesture helped them "feel" what wavelength is (mean 4.52), amplitude control scored highest (4.59), 93% found controlling the wave in the air natural, and 86% reported feeling more engaged and focused than in regular learning. The evidence, however, is perception-only, single-class, and without a comparison group — so the honest reading is that [[prompt-engineering]] has removed the production barrier, which makes the underlying question about embodiment *testable*, not that it has been answered. ## AI inside the immersive environment The second direction is where AI stops authoring and starts teaching inside the headset, and it is where [[rag|retrieval grounding]] becomes load-bearing. An agentic immersive training platform for high dose rate brachytherapy built a digital twin of the treatment suite — anatomically precise patient models, catheters, afterloaders, applicators — so trainees could see an applicator's spatial orientation relative to organs at risk, without a shielding room, a live radioactive source, or the privacy hazards of a physical pelvic exam. A knowledge-aware assistant grounded in [[medical-education|clinical]] guidelines supplied hands-free guidance through a three-tier voice interface (headset microphone to backend transcription and intent analysis to spatialised speech), which removes controller dependency during intricate maneuvers. Technically it worked: 3–5 second end-to-end latency across 50 Monte Carlo runs, context recall above 0.93 and answer relevance 0.87 on 52 expert-authored question-answer pairs, with a medical embedding model improving answer completeness. Its gaps define the current frontier of the pattern. Evaluation was objective metrics plus a single domain-expert user, not learners; there was no automated [[assessment]] or [[adaptive-learning|adaptive feedback]]; and the assistant cannot exceed the documents it was given, so institution-specific protocols and rare scenarios sit outside its competence. In other words, the tutoring intelligence inside immersive environments is still mostly **guidance**, not measurement — and the same platform architecture, with [[llm|model inference]] offloaded to a local GPU backend, is what makes hands-free guidance fast enough to be usable at all. The clearest attempt to date at that missing measurement comes from a vocational design studio rather than a clinical setting. In a 12-week interior-design course, a VR environment (headsets plus a 3D modeling tool) with an LLM-backed assistant rendered as a digital human was compared against conventional [[project-based-learning|project-based]] instruction, with the assistant assigned a distinct role per phase — resource recommendation and task decomposition, layered questioning with knowledge maps, simulated design effects and flaw detection, then discourse logging for the teacher ([[ai-ive-pbl-vocational-design-creativity-2026|Jin et al., 2026]]). Across 63 valid responses the immersive-plus-agent condition produced significantly higher design ability (η²p = .138) and creative ability (η²p = .111), cognitive (d = 0.90) and behavioral (d = 0.75) [[student-engagement|engagement]], and motivation (d = 0.74) and satisfaction (d = 0.69) — with reported [[cognitive-offloading|cognitive load]] *lower*, not higher (d = −0.52), which the authors credit to the assistant cutting search and cross-disciplinary integration effort. Two limits keep this from settling the question: ideational novelty and [[affective-computing|affective]] engagement did not move, and every outcome is [[self-report-measures|self-report]], with no artifact ratings or headset logs. The contrast with a comparable design-education deployment is instructive. An architectural studio using a [[generative-ai|GenAI]] plus multi-user XR pipeline found *declining* design [[self-efficacy]] confidence for the teams that used it and no advantage in blinded panel ratings. The difference between the two is less the hardware than the orchestration: the vocational study fixed the phase structure, the assistant's role in each phase, and the evaluation rubric before the intervention, and measured productive ability separately from ideational novelty. The instructor-facing version of the same pattern is younger and more fragile. [[luminote-llm-vr-stage-lighting-education-2026|Liang et al. (2026)]] let a stage-lighting instructor speak their intent into a VR scene with a laser pointer as the spatial anchor, and had an [[llm]] turn it into instructor-reviewable spatial annotations, executable lighting demonstrations and on-demand jargon explanations. Over 55 prompts and 531 generated actions, assistance was strongest on expressive, under-specified goals (212 of 245 visual-effect actions applied, 86.5%) and weakest on fixture-level requests, where 26 of the 28 rejected actions traced to directional references such as "left light" that the model resolved to the wrong fixture — grounding in the scene's spatial reference frame, not executability, was the binding constraint. Instructors used suggestions as a [[human-in-the-loop-ai|controllable refinement process]] rather than an answer: 127 of the 147 rejected or modified actions (86.4%) were followed by a new prompt, and only one by manual adjustment. The study's most cautionary result is representational — annotations that externalized expert reasoning did not reliably align with what novices understood, so immersive AI instruction carries both the grounding problem of the tutoring cases above and a second one, that instructor competence and learner comprehension are not the same target. ## What the evidence shows The strongest evidence for VR/AR is comparative and [[discipline-specific-aied|discipline-specific]]. A [[meta-analysis-systematic-review|meta-analysis]] of 33 experimental and quasi-experimental studies (N = 3,181) of emerging [[ai-technologies|technologies]] in teaching [[english-education|English as a foreign language]] found a small-to-moderate overall effect (g = 0.38) and, within it, the **largest effects for VR/AR**, with gains rising by educational level and favoring productive skills (speaking, writing) over receptive ones. The counter-evidence is just as informative. In the first meta-analysis of AI-based interventions for [[special-education|students with disabilities]] — 29 studies, 239 effect sizes, medium overall effect g = 0.588 — intelligent VR systems produced g = 0.528, **not statistically significant**, while computer software reached 0.959 and robots 0.509. Publication bias was present and trim-and-fill reduced the overall estimate to g = 0.269. The pattern is not that immersion fails, but that its effects are small, heterogeneous, and sensitive to how the particular intervention was designed and compared. The clearest negative result in the knowledge base comes from a studio deployment rather than a comparison of modalities: in a 27-student architectural design studio, teams using a GenAI plus multi-user XR pipeline declined more in design [[self-efficacy]] confidence (β = −1.675) and outcome expectancy (β = −2.088) than teams working the normal course workflow, with no significant difference in expert panel ratings of their presentations ([[genai-xr-architectural-design-education-2026|Xiao et al., 2026]]). The authors explain it as phase-dependent complementarity with real friction — GenAI for externalizing tentative ideas, XR for spatial and scale evaluation — alongside control, dimensional-fidelity, shared-attention and motion-comfort problems, a reminder that immersive tooling adds interaction costs as well as capability. ## Fidelity, presence and the authenticity gap Where immersive practice is used for interpersonal and procedural skill, the limiting factor is not visual fidelity but felt authenticity. A [[mixed-methods-research|mixed-methods]] review of AI-powered nursing simulation (19 studies, N = 1,253) found AI effective for cognitive knowledge and affective outcomes but inconsistent for complex psychomotor skills, and named the reason: an **authenticity gap** covering emotional resonance, nonverbal cue recognition, and tactile and physical-[[summative-assessment|examination]] dimensions. Its practical recommendation is a **stepped simulation continuum** — AI is well suited to highly structured objectives such as foundational communication and history taking, while advanced psychomotor and emotionally complex scenarios belong with human-standardized patients and clinical placement. Technical instability compounds the problem: speech-recognition delays inject extraneous [[cognitive-offloading|cognitive load]] and anxiety, which makes stability and latency design levers rather than implementation details. The same logic explains why presence is not automatically good. Motion sickness, UI inconsistency, hardware cost, and uneven device access decide who can use an immersive environment at all, which is why the modality's [[equity-in-ai-education|equity]] questions connect to [[accessibility]] and [[inclusive-learning]] rather than sitting apart from them. ## Connected Concepts - [[simulation]] — what immersive environments usually display; the model, not the modality - [[embodied-learning]] — the mechanism the modality is supposed to exploit - [[multimodal]] — gesture, voice and spatial input as learning channels - [[situated-learning]] - [[experiential-learning]] - [[game-based-learning]] - [[professional-training]] — the setting where immersive practice is most established - [[generative-ai]] — the new authoring layer - [[prompt-engineering]] — how non-programmers build and refine immersive tools - [[agentic-ai]] - [[pedagogical-agent]] - [[rag]] — grounding guidance inside the headset - [[intelligent-tutoring]] - [[visualization]] - [[medical-education]] - [[accessibility]] - [[inclusive-learning]] - [[trust-calibration]] — fidelity, and the learner's awareness of its limits - [[cognitive-offloading]] — latency and instability as extraneous load - [[edtech-platform]] — synchronisation, latency and multi-site presence - [[arts-design-and-media-education]] ## Connected Articles - [[genai-ar-physics-simulation-prompt-2026]] — four-element prompt generating hand-controlled AR physics simulations; 29-student pilot, perception-only evidence - [[hdr-brachytherapy-agentic-ai-simulation-2026]] — agentic AI VR training platform with a RAG-grounded hands-free assistant; digital twin, latency and retrieval metrics - [[multi-site-vr-immersive-learning]] — real-time multi-site VR classroom; usability and VR-sickness outcomes, UI consistency as the weak point - [[jiang-ai-powered-simulation-nursing-education-2026]] — the authenticity gap and the stepped simulation continuum - [[mixed-reality-engineering-learning]] — mixed-reality apps vs physical toolkits vs classroom in engineering mechanics; engagement up, complex visualization still hard - [[liu-emerging-tech-tefl-review-2026]] — TEFL meta-analysis where VR/AR produced the largest subgroup effects - [[zhang-ai-students-disabilities-meta-analysis-2024]] — intelligent VR for students with disabilities: positive but not statistically significant - [[vargas-ai-catalyst-situated-learning-2026]] — lack of immersive tooling as a barrier to situated learning - [[medgame-llm-medical-education-gamification]] — gamified medical training with AI - [[tech-enhanced-tabletop-cybersecurity-education]] — augmented tabletop scenarios in cybersecurity education - [[genai-xr-architectural-design-education-2026]] — Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study - [[ai-ive-pbl-vocational-design-creativity-2026]] — AI-IVE-PBL: immersive VR design studio with an LLM-backed teaching assistant, evaluated against traditional PBL (Jin et al. 2026) - [[luminote-llm-vr-stage-lighting-education-2026]] — LumiNote: LLM-assisted multimodal instruction in VR stage lighting education, and where grounding broke down (Liang et al. 2026) --- ## [Training Pedagogical LLMs for Tutoring](https://edtechdev.github.io/aied/concepts/pedagogical-llm-training/) > Domain-specialized optimization can transform a mid-sized [[open-source]] model (Qwen3-32B) into a [[pedagogy|pedagogical]] domain expert that outperforms far larger proprietary systems — but only when training rewards *guiding* rather than *answering*.([[singh-eduqwen-pedagogical-rl-2026]]) Classical instructional design theory (ADDIE, Dick & Carey) combined with modern ReAct reasoning achieves the highest performance in automated instructional design.([[jeon-isd-agent-bench-2026]]) ## Questions to Consider - General-purpose [[conversational-ai|chatbots]] are optimized to give quick, correct answers. Why is that the *opposite* of what a tutor needs, and what does that 'incentive mismatch' suggest about off-the-shelf AI as a [[teacher-role|teaching]] tool? - A benchmark found 97 models scored between 28% and 89% on pedagogical knowledge—meaning it's not automatically learned in pretraining. Does that surprise you, and what does it imply about trusting a general [[llm]] to teach? - EduQwen's training explicitly rewards 'guiding' over 'answering.' Before reading the methods, can you think of how you'd tell an AI to prefer guiding—and how you'd measure whether it actually did? - The page shows classical design theory (ADDIE) combined with flexible reasoning beat both pure theory and pure technique. Why might 'structure plus flexibility' outperform either alone when an AI designs instruction? - Training pedagogy into a model costs time, data, and compute. For your context, what would convince you the investment is worth it versus just prompting a general-purpose model with 'act like a tutor'? ## Introduction General-purpose LLMs are optimized for helpfulness: users want quick, correct answers. Tutoring requires the opposite: the goal is **not to provide the answer, but to help the student get to the answer themselves**. This creates a fundamental incentive mismatch. ## Approach 1: RL-SFT-RL Pipeline for Pedagogical Reasoning (EduQwen) Singh et al. (2026) developed a three-stage pipeline transforming Qwen3-32B into EduQwen, achieving **96.52%** on the CDPK Benchmark and surpassing Gemini-3 Pro (90.55%). ### Stage 1: Initial RL (EduQwen 32B-RL1) - **Algorithm:** DAPO (Decoupled Advantage Policy Optimization) with asymmetric clipping - **[[reinforcement-learning|Reward model]]:** Prioritizes *guiding* responses over direct answers - **Curriculum learning:** Progressive difficulty; hard-negative mining excludes questions the base model already solves perfectly - **Extended rollouts:** 5→8 steps to capture multi-step pedagogical decisions - **Result:** 94.13% (already SOTA) ### Stage 2: Synthetic SFT (EduQwen 32B-SFT) - RL1 model generates 40,000 synthetic responses - Gradient-based selection retains only hard examples - Difficulty-weighted sampling: easy questions → one example; hard questions → all, weighted up - **Result:** 96.20% ### Stage 3: Final RL (EduQwen 32B-SFT-RL2) - Second DAPO round, reusing the original hard-negative set - Model now solves problems it originally found challenging - **Result:** 96.52% (definitive SOTA) ## The Pedagogy Benchmark: Evaluating Pedagogical Knowledge Lelièvre et al. (2025) introduced **The Pedagogy Benchmark**, measuring Cross-Domain Pedagogical Knowledge (CDPK) and [[special-education|Special Education Needs]] and Disability (SEND) knowledge from real teacher [[educational-development|professional development]] exams. Across **97 models**, accuracy ranged from **28% to 89%**—revealing that pedagogical knowledge is not automatically acquired in general pretraining. **EduQwen connection:** Singh et al.’s EduQwen achieved **96.52% on CDPK**, demonstrating that targeted RL+SFT optimization can close the pedagogical knowledge gap that Lelièvre et al. document. The benchmark serves as both a diagnostic (showing most models fail at pedagogy) and a training target (showing optimization works). Live leaderboards track cost-accuracy Pareto frontiers: [rebrand.ly/pedagogy](https://rebrand.ly/pedagogy) ## Approach 2: Theory-Grounded Instructional Design Agents (ISD-Agent-Bench) Jeon et al. (2026) created a benchmark for LLM agents automating [[learning-design|Instructional Systems Design]] (ISD), testing whether classical pedagogy theory improves agent performance. | Architecture | Performance | Why | |-------------|-------------|-----| | **Hybrid: theory + ReAct** | **Best** | Classical ADDIE/Dick & Carey frameworks provide structure; ReAct enables flexible multi-step reasoning | | Pure theory-based | Moderate | Structured but inflexible | | Technique-only (pure ReAct) | Worst | Flexible but lacks pedagogical grounding | **Key insight:** Theoretical quality strongly correlates with benchmark performance. Theory-based agents excel in **problem-centered design** and **objective-assessment alignment**. ### Benchmark Design - **25,795 scenarios** from Context Matrix (51 variables × 5 categories × 33 ISD sub-steps) - **Multi-judge protocol** across diverse LLM providers to mitigate LLM-as-judge bias - High inter-judge reliability achieved ## Approach 3: Pedagogical Instruction Following (LearnLM) and Authentic-Data Post-Training (TeachLM) Two complementary post-training strategies for embedding pedagogy into foundation models: - **Pedagogical instruction following (LearnLM).** [[learnlm-improving-gemini-learning|Google's LearnLM]] reframes education-model training as *pedagogical instruction following*: training and evaluation examples carry system-level instructions describing the desired pedagogical behavior, letting developers/teachers specify tutor behavior without committing to any single definition of pedagogy. Mixed directly into Gemini's post-training (SFT + reward-model + RLHF stages) via co-training, LearnLM was preferred by experts over GPT-4o (+31%), Claude 3.5 Sonnet (+11%), and base Gemini 1.5 Pro (+13%) across scenario-guided multi-turn evaluations. Key finding: **RL is substantially more effective than SFT alone** for following nuanced pedagogical instructions in long conversations. - **Authentic-data post-training (TeachLM).** [[teachlm-post-training-llms-education|TeachLM]] argues that [[prompt-engineering|prompt engineering]] is a stopgap and that the scarce ingredient is *authentic* learner–tutor interaction data. Trained on 100,000 hours of one-on-one Polygence sessions (rigorously anonymized), it builds a fine-tuned **authentic [[student-modeling|student model]]** enabling synthetic multi-turn evaluation, and the teacher model doubles student talk time, improves questioning style, and increases dialogue turns by 50%. **Synthesis:** LearnLM shows that instruction following + RLHF is a viable route when training data is scarce; TeachLM shows that when authentic longitudinal interaction data *is* available, post-training on it directly outperforms both prompting and synthetic-only data. Together they frame the pedagogical-training design space as a choice between scalable instruction-conditioned post-training and data-driven fine-tuning on real tutoring interactions. ## Rubric-guided prompting as a lightweight alternative Not all pedagogical shaping requires retraining. [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] showed that rubric-guided prompting — treating the rubric as a semantic interface between human pedagogical intent and machine inference — can push a general-purpose LLM toward human-like [[evaluative-judgment|evaluative judgment]] without fine-tuning: iterative rubric co-refinement raised LLM–human agreement on student design work from 54.75% to 81.25% (Cronbach's Alpha 0.393 → 0.798), and role-aware prompting (instructor, peer-reviewer, grant-reviewer) produced distinct evaluative feedback. This complements the training-based approaches above: where TeachLM argues prompt engineering is a stopgap and authentic-data post-training is the scarce ingredient, Yaşar et al. demonstrate that a well-engineered rubric can itself be a powerful, low-cost lever for aligning LLM evaluation with pedagogical intent — though [[human-in-the-loop-ai|human-in-the-loop]] oversight remains essential, as models can still misinterpret nuance, hallucinate rationale, or blend roles. A further lightweight alternative is prompt-level role-play customization without retraining: [[zhuang-zhang-chatgpt-math-teacher-education-2026|Zhuang and Zhang (2025)]] used the OpenAI custom-GPT feature to simulate a misconception-holding middle-school math student, and found that a refined, literature-grounded prompt (specifying three ratio-reasoning [[misconceptions]]) elicited the target conceptual errors far more reliably than a broad algebra prompt (0.98 vs. 0.40 presence) — evidence that careful prompt design can substantially steer an off-the-shelf model toward a desired pedagogical persona, even while the simulated agent retained authenticity limitations (teacher-like tone, role confusion). The lightest intervention in this family is not prompting but parameter-efficient adaptation. [[lora-finetuned-control-systems-course-qa-2026|Lu et al. (2026)]] built 360 system–user–assistant dialogues from a Linear Control Systems course, restructured answers into a Solution–Method–Teaching-Points format, and applied LoRA to Qwen2.5-3B and 7B at ranks 4, 8 and 16. Structured-output coverage moved from near zero at base to roughly 1.00, and the best configuration (7B, r = 16) reached ROUGE-L 0.4093 with bootstrap confidence intervals for the gain entirely above zero — but gain per million adapter parameters fell monotonically as rank rose, so course-level alignment is a scale-and-rank trade-off rather than a free upgrade. The metrics measure similarity and formatting, not derivational accuracy. ## Approach 4: Training Simulator Roles, Not Only Tutors The same post-training machinery is now pointed at the learner side of the interaction, and the results say the supervision budget matters more than the prompt. [[misconception-acquisition-dynamics-llms-2026|Liu et al. (2026)]] instruction-tuned three small models to *acquire* algebra misconceptions in two roles — a Novice Student Misconception Model holding one misconception, and an Expert Tutor Misconception Model holding ten — and measured both misconception accuracy and correct-solving accuracy. The student role showed a trade-off no prompt could fix: the learned error overgeneralised beyond its applicable problem types until correct examples were explicitly mixed into the training data, at ratios as low as one correct example per four misconception examples. The tutor role showed no such cost, with correct accuracy stable or rising from 93% to 98% when ten misconceptions were trained jointly, though classroom-scale samples were insufficient and rare misconceptions would require cross-institution data. Most decisively, neither role acquired anything when trained on final answers alone — misconception accuracy stayed below 30% at every data size — so step-level solution traces, not more examples, are the binding requirement. [[swim-student-writing-simulation-2026|SWIM (Do, Kontak and Sachan, 2026)]] reaches the mirror conclusion for a writing simulator: rubric-grounded prompting gave limited proficiency control (best average trait QWK 0.577 for Claude Sonnet, 0.422 for GPT-5.4, near zero for prompting an open 7B model), supervised fine-tuning lifted a 7B model to 0.474 ± 0.023, and GRPO against an automated-essay-scoring-derived reward lifted it further to 0.618 ± 0.005 across every trait and prompt, with the reward designed as a dense trait-normalized accuracy because exact-match rewards are too sparse in the multi-trait setting. ## Synthesis: What Makes Pedagogical Training Work | Principle | EduQwen | ISD-Agent-Bench | |-----------|---------|-----------------| | **Reward/guide, don't answer** | DAPO reward model penalizes direct solutions | Theory-enforced ISD steps require alignment between objectives and assessment | | **Curriculum by difficulty** | Hard-negative mining + progressive rollouts | Context Matrix systematically varies complexity | | **Multi-step reasoning** | Extended rollouts (5→8 steps) | ReAct-style reasoning chains | | **Validate with theory** | CDPK benchmark measures pedagogical knowledge | ADDIE/Dick & Carey frameworks ground design decisions | | **Iterative refinement** | RL → SFT → RL pipeline | Multi-judge evaluation reduces bias | ## Relationship to Safety and Design Training for pedagogy is not just about accuracy — it is a **safety intervention**: - A model that rewards "guiding" over "answering" is less likely to commit [[hazra-safetutors-pedagogical-safety-2026|answer over-disclosure harms]] - Theory-grounded agents (ISD-Agent-Bench) align with pedagogical principles that prevent [[metacognition|metacognitive suppression]] - However, training on pedagogical [[benchmark|benchmarks]] does not guarantee multi-turn safety; SafeTutors shows even specialized models degrade over sustained dialogue - **Grounding and validation can substitute for — or complement — training.** [[reddig-maclellan-personalized-feedback-llm-2026|Reddig, Arora & MacLellan (2025)]] found that a frontier *untrained* GPT-4 produced ~35% too-general, incorrect, or answer-revealing hints when authoring ITS feedback, and that its own automated quality checks misaligned with human judgment — leading the authors to conclude that LLMs lack an internal model of instruction and that robust validation or domain-specific training is required before unsupervised learner-facing use, supporting the case that grounding and quality control are themselves pedagogical interventions alongside reward design. ### Sycophancy reduction as a training objective Because tutoring requires corrective friction — challenging a student's incorrect claim rather than affirming it — reducing [[ai-sycophancy|sycophancy]] is a core objective for pedagogical LLM training. [[eduframetrap-llm-sycophancy-educational-safety|EduFrameTrap]] shows that models which resist context-switch attacks still capitulate under authority or social-[[affective-computing|affective]] pressure, withholding corrective feedback; its authors argue "kind-but-correct" behavior should be an explicit training requirement, not a [[usability-research|usability]] preference. Training that rewards guiding over answering (as in EduQwen's DAPO reward model) is one structural lever against sycophantic answer-giving. Yet [[contextual-sycophancy-ai-literacy|contextual sycophancy]] persists even after prompting/alignment training — learners' errors still propagate into AI advice — so sycophancy mitigation in trained tutors must combine reward design, alignment against sycophancy benchmarks, and system-level safeguards rather than rely on any single stage. ## Open Questions 1. Does pedagogical RL training generalize across subjects, or is [[discipline-specific-aied|subject-specific]] tuning (as SafeTutors suggests) always needed? 2. Can the RL-SFT-RL pipeline be combined with longitudinal memory (see [[nie-personavlm-long-term-personalization-2026]]) for personalized tutoring? 3. Would ISD-agent theory improve general tutoring conversation, or is it limited to macro-level [[curriculum-design|curriculum design]]? ## Connected Concepts - [[intelligent-tutoring]] - [[scaffolding]] - [[adaptive-learning]] - [[metacognition]] - [[affective-tutoring]] - [[human-in-the-loop-ai]] - [[personalized-learning]] - [[student-modeling]] - [[self-regulated-learning]] - [[pedagogical-safety]] - [[formative-assessment]] - [[llm]] - [[authentic-assessment]] - [[ai-sycophancy]] - [[ai-feedback-quality]] - [[bias-mitigation]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) ## Connected Articles - [[zerkouk-comprehensive-review-its-2025]] - [[civic-education-ai-lesson-plans]] - [[moon-cognitive-agent-compilation-problem-solver-modeling-2026]] - [[contextual-sycophancy-ai-literacy]] - [[educational-llm-alignment]] - [[eduguard-safe-rag-llm-tutor]] - [[kar-mathbuddy-affective-math-tutoring-2025]] - [[llm-tts-dialogue-lesson-generation]] - [[multimodal-learning-genai]] - [[neural-symbolic-knowledge-tracing]] - [[nsmq-riddles-science-math-benchmark]] - [[singh-eduqwen-pedagogical-rl-2026]] - [[eduframetrap-llm-sycophancy-educational-safety]] — Sycophancy is an educational safety risk: Why LLM tutors need sycophancy benchmarks - [[tact-pedagogically-adaptive-esl-tutoring]] - [[learnlm-improving-gemini-learning]] — LearnLM: Improving Gemini for Learning - [[teachlm-post-training-llms-education]] — TeachLM: Post-Training LLMs for Education Using Authentic Learning Data - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design - [[reddig-maclellan-personalized-feedback-llm-2026]] - [[zhuang-zhang-chatgpt-math-teacher-education-2026]] - [[lora-finetuned-control-systems-course-qa-2026]] — LoRA Fine-Tuned Models for Control Systems Course Q&A: A Multidimensional Evaluation of Model Scale and Rank Effects - [[misconception-acquisition-dynamics-llms-2026]] — data composition, correct-example mixing and step-level supervision for misconception-aware models - [[swim-student-writing-simulation-2026]] — supervised and reward-based training beat rubric prompting for proficiency control - [[omniedu-open-educational-foundation-models-2026]] — OmniEdu: Open Foundation Models for Learning and Teaching --- ## [Learner Modeling and Adaptive Instruction](https://edtechdev.github.io/aied/concepts/student-modeling/) > **Learner modeling and adaptive instruction** — the umbrella for how AI represents learners (what they know, feel, and need) and how it uses those representations to adapt [[teacher-role|teaching]]. The family spans the *modeling* layer — **student modeling**, [[knowledge-tracing]], [[cognitive-diagnosis]], and [[simulating-students|simulating students]] — and the *adaptive systems* that consume those models — [[intelligent-tutoring]], [[adaptive-learning]], and [[personalized-learning]]. The shared question: *how does a system know what a learner knows, and what should it teach next?* ## Questions to Consider - The umbrella question this page poses is: how does a system know what a learner knows, and what should it teach next? Before you read on, how would you even begin to represent 'what a learner knows' in a machine? - Learner modeling spans knowledge-tracing (tracking knowledge over time), cognitive diagnosis (mapping mastered skills), and simulating students (synthetic learners). What do you think each approach is good at — and what does each risk getting wrong? - Every adaptive AI depends on some model of the learner. If a model is only as good as the evidence feeding it, what evidence do you think AI systems actually have about a student, and what important things about them remain invisible? - A model might capture what a student gets right and wrong but not why, or not how they feel. How could a learner model mislead an adaptive system in ways that harm rather than help the student? - If you were designing an adaptive tutor, what would you want its model of you to include — and what would you want it explicitly forbidden from assuming? ## Introduction Learner modeling is the computational representation of learners; adaptive instruction is what systems do with that representation. Every adaptive AI in education depends on some model of the learner — even a lightweight one — and every learner model exists to inform some instructional decision. This page is the umbrella for that pipeline: the modeling methods, the systems that act on models, and how they relate. ## The modeling layer These concepts answer "what does this learner know, feel, and need?" — the representation side of the family. - **student modeling** — the broad practice of representing learner characteristics (knowledge, skills, [[affective-computing|affective]] states, [[student-engagement|engagement]], preferences) in computational form. It is the umbrella term within this layer, encompassing all ways of representing a learner. - **[[knowledge-tracing]]** — the specific practice of modeling cognitive knowledge *over time* by tracking performance on exercises and predicting future mastery. It formalizes the temporal dynamics of learning — when knowledge is gained, decays, and how concepts relate. - **[[cognitive-diagnosis]]** — fine-grained [[assessment]] of which specific skills or knowledge components a learner has mastered, producing a mastery profile that supports targeted remediation. - **[[simulating-students|simulating students]]** — generating *synthetic* learners on demand, rather than representing a real one, so [[pedagogy]] and AI systems can be tested or trained offline. The study of [[zhang-ml-student-progress-programming-2026|Zhang, Jeffries & Koprinska (2025)]] illustrates that faithful representation does not require the most complex model family: a lightweight, intrinsically interpretable decision-tree student model — built from course content-interaction features rather than rich telemetry — predicts module-level progress in large-scale online [[cs-education|programming]] courses (85–91% accuracy) and separates disengaged at-risk, disengaged-but-successful, and engaged high-performer [[student-engagement|engagement]] profiles, supporting [[learning-analytics]] early-warning at scale. Student models can also be built purely from behavioral traces and still support adaptation. [[an-goel-self-directed-modeling-2026|An, Hammock & Goel (2025)]] derived three engagement profiles — Observation, Construction, and Exploration — from the clickstreams of 315 online learners building 822 ecological models in VERA, without any demographic or contextual data, and showed these profiles predict model quality (Exploration yields the most complex and diverse models, while Observation is dominated by copied rather than original models). Such engagement-level characterizations are the coarse-grained student models that the [[adaptive-learning|adaptive-instruction]] layer can consume to target feedback. ## The adaptive-instruction layer These concepts answer "what should be taught next?" — the application side that consumes learner models. - **[[intelligent-tutoring]]** — systems that use student models and mastery estimates to select problems and provide step-level guidance, the classic application of learner modeling. - **[[adaptive-learning]]** — systems that adjust content, pacing, or difficulty in response to the learner model. - **[[personalized-learning]]** — the broader tailoring of instruction, content, and pathways to individual learner characteristics and preferences. ## How the members relate The concepts form a pipeline rather than competitors: **student modeling** is the umbrella representation; [[knowledge-tracing]] and [[cognitive-diagnosis]] are specific modeling methods that populate it; [[simulating-students|simulation]] *generates* learners rather than representing real ones; and [[intelligent-tutoring]], [[adaptive-learning]], and [[personalized-learning]] are the systems that consume these models to adapt instruction. **Student modeling vs. simulating students** is the key distinction to keep straight. Student modeling is about **representing a real learner** — building a model *from* an actual student's data so an adaptive system can act on that individual. Simulating students, by contrast, **generates a synthetic learner** on demand to stand in for real learners so pedagogy and AI can be evaluated or trained offline. The two are closely related rather than interchangeable: simulated students typically *embed* a student model (an epistemic state, [[misconceptions|misconception]] set, or engagement profile) and draw on the same constructs that [[knowledge-tracing]] and [[cognitive-diagnosis]] formalize. Their purposes diverge — student modeling serves live adaptation by informing decisions about a real person, whereas [[simulation]] fabricates learners to test systems (and increasingly to audit AI, e.g., [[lopez-pernas-llm-appropriate-student-support-2026|López-Pernas et al. (2026)]]) rather than to act on any real individual. **Knowledge tracing vs. student modeling** is the other common confusion. Knowledge tracing specifically models cognitive knowledge over time; student modeling is the broader practice covering all aspects of a learner (affective state, engagement, preferences). Knowledge tracing is a *type of* student modeling focused on the cognitive-temporal dimension. Knowledge-tracing constructs also inform [[simulating-students|simulated students]] — a simulated learner's cognitive state is often formalized with the same mastery/decay dynamics that knowledge tracing models, so simulation is a way to *generate* the knowledge states that tracing methods normally *infer* from real response data. **Anchoring tracing to the [[curriculum-design|curriculum]] strengthens the model.** [[pradeesh-outcome-knowledge-tracing-affinity-2026|Pradeesh et al. (2026)]] show that a learner model gains fidelity when tracing is tied to explicit curriculum structure rather than learned purely from data: their Outcome-Based Knowledge Tracing (OKT) treats course outcomes in Outcome-Based Education as the knowledge concepts to trace, supplies concept relationships through expert-validated OBE "affinity mappings" between course and program outcomes (an explicit alternative to implicit attention or graph message passing), and uses a memory-augmented module to model how one outcome's attainment impacts others. On live engineering-program data it beat DKT, DKVMN, EKT, and SimpleKT baselines (89.81% AUC), illustrating that the modeling layer can exploit the curriculum's own structure to represent learners more faithfully. **Intelligent tutoring vs. adaptive/personalized learning** sits on the application side: intelligent tutoring is the problem-selecting, step-guidance system; adaptive learning tunes content and pacing; personalized learning is the broadest tailoring of the whole learning experience. All three are the "consumers" of the modeling layer. ## The shared validity challenge Across the whole family, the defining validity challenge is the same: the learner representation must **faithfully reflect a learner's true state** rather than the system's default assumptions. For **student modeling** and [[knowledge-tracing]], this means the model must genuinely capture what a learner knows ([[ai-ed-evaluation|evaluation]] and [[assessment-validity|measurement validity]]). For [[simulating-students|simulation]], it means the synthetic learner must exhibit realistic imperfection rather than the model's full competence or [[ai-sycophancy|sycophantic]] agreement. Adaptive systems that consume faulty models inherit and propagate that error. **Correctness is not always a faithful signal.** [[deceptive-overgeneralization-adaptive-learning-2026|An, McLaren, and Stamper (2026)]] show that a learner model inferring mastery from correct actions can misrepresent a learner's true state: learners who exhibit *deceptive overgeneralization* appear mastered yet omit a critical application constraint, so adaptive systems can stop practice prematurely. Learner models should assess conditional understanding — including whether the learner knows when to withhold an action — not only action correctness. **How a model is validated is itself a validity question.** [[schuetze-knowledge-tracing-forgetting-2026|Schuetze, Yan, and Carvalho (2025)]] show that popular learner models (BKT, BKT-with-Forgetting, AFM) appear to capture human learning only when fit retroactively to a full multi-session dataset; under time-based (walk-forward) cross-validation — predicting a future session from earlier ones, how such models are actually deployed — they overestimate performance, miss the [[retrieval-spacing-interleaving|spacing effect]], and mis-order practice conditions. Because forgetting-augmented and forgetting-free models performed about equally across sessions, the authors conclude that forgetting is often absorbed into learner parameters rather than genuinely represented. The lesson for the family is that a faithful learner representation must be validated the way it is used — and that conflating in-the-moment performance with long-term retention produces models that look accurate yet misrepresent learners. ## LLM-era modeling Recent advances use [[llm|LLMs]] for richer modeling. The [[xie-hillm-cd-2026|HiLLM-CD framework]] represents students as proficiency trees; [[multimodal-knowledge-graph-educational-reasoning|multimodal approaches]] construct evidence-grounded knowledge representations from diverse data sources; [[inside-llm-student-simulator-reasoning-2026|LLMs now simulate students with reasoning]]. LLMs enable automated model construction from educational text and higher-fidelity [[simulating-students|student simulation]], reducing reliance on expert annotation — while sharpening the fidelity concerns above. Learner-model signals also *ground* LLM reasoning: [[reddig-maclellan-personalized-feedback-llm-2026|Reddig, Arora & MacLellan (2025)]] found that feeding GPT-4 a student's Bayesian [[knowledge-tracing]] skill estimate along with the tutor's interface structure sharply improved its error diagnosis (logical-error identification rising from 40% to 81% on factoring; ~87.8% overall), while multi-step problems and responses containing several errors remained the weakest cases — evidence that coupling a formal learner model to an LLM strengthens, but does not guarantee, sound inference about a real student. [[colearn-agentic-tutor-co-learning-loop-2026|CoLearn (He et al., 2026)]] shows what a persistent version of that coupling looks like: mastery and mined misconceptions are stored per (learner, subject) rather than as per-session logs, so evidence accumulates across sessions, and the memory is written by an LLM-graded observation function while staying inspectable to the learner through mastery bars and a label naming what each generated question was chosen to probe. Its controls make the writing step explicit — with the memory read but no longer updated, the share of items aimed at a genuinely weak skill fell from 0.72 to 0.57 — and it keeps this page's qualification intact: the stored mastery is the agent's belief about the learner, not a measurement of their knowledge. ## Connections to other concepts Learner modeling and adaptive instruction feed into [[learning-analytics]] ([[visualization|dashboards]] and interventions), [[formative-assessment]] (analytics-driven assessment), and [[feedback]] (what the system tells the learner). It connects to [[ai-education]] as a core strand of AI for education. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[explainable-ai]] - [[learning-analytics]] - [[knowledge-tracing]] - [[knowledge-graph]] - [[adaptive-learning]] - [[intelligent-tutoring]] - [[personalized-learning]] - [[formative-assessment]] - [[k-12]] - [[affective-tutoring]] - [[llm]] - [[higher-ed]] - [[ai-education]] - [[simulating-students]] - [[cognitive-diagnosis]] - [[feedback]] - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[causal-modeling-competency-assessment-2026]] — Causal Modeling of Support Interventions for Student Competency Assessment - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[learning-context-framework-context-aware-ai-education-2026]] - [[interactive-online-learning-ai-2025]] - [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] — Ontology-based layered hybrid knowledge model for personalized e-learning - [[yasir-llm-tutoring-agents-2026]] — LLM tutors over-reject valid-alternative, over-validate incorrect (Yasir et al. 2026) - [[haiml-human-centered-ai-metacognitive-model-2026]] - [[ai-guided-learning-audiovideo-2026]] - [[multimodal-item-parameter-estimation-2026]] - [[at-risk-students-ml-prediction]] - [[correct-answer-trap-misconceptions]] - [[cross-subject-validity-delayed-start]] - [[educlaw-bench-pedagogical-llm-agents-2026]] - [[edumirror-educational-social-dynamics]] - [[huang-interpretable-knowledge-tracing-2026]] - [[kar-mathbuddy-affective-math-tutoring-2025]] - [[knowledge-gap-detection-ai-tas]] - [[llm-item-difficulty-prediction]] - [[multimodal-knowledge-graph-educational-reasoning]] - [[proprl-prerequisite-relation-learning]] - [[simulating-students-java-programming-errors-llms]] - [[skill-acquisition-without-temporal-info]] - [[xie-hillm-cd-2026]] - [[learnity-graphs-lifelong-learning-framework-2026]] - [[inside-llm-student-simulator-reasoning-2026]] - [[trace-course-grade-prediction-2026]] - [[sc2r-counterfactual-recourse-educational-2026]] — From Student Risk Prediction to SC2R: Counterfactual Recourse - [[teachlm-post-training-llms-education]] — TeachLM: fine-tuned authentic student model for multi-turn evaluation - [[eeg-familiarity-automated-assessment-2026]] — Automating Learner Assessment: EEG-Based Familiarity Prediction - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[distilling-self-explaining-lm-learning-analytics-2026]] — Distilling self-explaining LM for learning analytics - [[studentsim-llm-student-simulators]] — StudentSim: Training LLM-based Student Simulators - [[predicting-attrition-competitive-programming]] — Predicting Student Attrition in Competitive Programming - [[pradeesh-outcome-knowledge-tracing-affinity-2026]] — Outcome-based knowledge tracing with affinity mapping - [[an-goel-self-directed-modeling-2026]] - [[reddig-maclellan-personalized-feedback-llm-2026]] - [[schuetze-knowledge-tracing-forgetting-2026]] - [[zhang-ml-student-progress-programming-2026]] - [[process-grounded-language-cognitive-diagnosis-2026]] — Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis - [[exrec-exercise-recommendation-knowledge-tracing-2025]] — compact learner state plus a calibrated tracer as a recommender environment - [[misconception-acquisition-dynamics-llms-2026]] — the Expert Tutor Misconception Model as a computational analogue of knowledge of student misconceptions - [[llm-distractor-generation-student-reasoning-2026]] — modeling incorrect reasoning rather than correctness - [[swim-student-writing-simulation-2026]] — proficiency-conditioned modeling of student writing - [[colearn-agentic-tutor-co-learning-loop-2026]] — CoLearn: An Agentic Tutor that Learns its Learner in a Human-AI Co-Learning Loop --- ## [Knowledge Tracing](https://edtechdev.github.io/aied/concepts/knowledge-tracing/) > **Knowledge tracing** — modeling what learners know over time by tracking their performance on exercises and predicting future mastery. It is the knowledge base's richest modeling thread, spanning Bayesian, deep learning, and [[llm|LLM-enhanced]] approaches to tracking student knowledge as it evolves. ## Questions to Consider - Knowledge tracing models what you know over time from your performance on exercises, tracking when knowledge is gained and when it decays. What can your answers reveal about whether you truly 'know' something versus just got it right this time? - The page warns that 'mastery is not correctness' — a learner can appear mastered yet systematically misapply a skill when a hidden condition is violated. When have you seen someone (or yourself) look like they understood something but actually hadn't? - If knowledge tracing feeds adaptive systems that decide what to teach next, what goes wrong when the model mistakes correct answers for true mastery and moves a student on too early? - Knowledge tracing comes in many forms — Bayesian, neural, hypergraph, dialogue-based, LLM-enhanced. What trade-offs would you expect between a transparent model you can explain and a powerful but opaque one? - The page connects knowledge tracing to simulated students — generating the knowledge states tracing normally infers from real data. How might simulating learners help test a tutor before it meets real students? - Since knowledge decays over time, what should an adaptive system do with a student's past 'mastery' once they've forgotten? How would you design for forgetting rather than assuming knowledge persists? ## Introduction Knowledge tracing transforms raw exercise responses into estimates of what a student has mastered and what they still need to learn. Unlike simple correctness tracking, knowledge tracing models the temporal dynamics of learning — when knowledge is gained, when it decays, and how concepts relate to each other. ### Approaches represented in the knowledge base - **Bayesian approaches:** [[stanbkt-bayesian-knowledge-tracing]] standardizes BKT implementations, while [[mbp-kt-meta-behavioral-knowledge-tracing]] incorporates meta-behavioral signals - **Soft-evidence BKT with an LLM observation function:** [[colearn-agentic-tutor-co-learning-loop-2026|CoLearn (He et al., 2026)]] keeps the BKT structure but replaces the binary correct/incorrect observation — standard BKT's input — with a continuous one: an [[llm]] grader emits graded mastery evidence plus a confidence weight, blended into a confidence-shrunk posterior and gated so that a clearly wrong answer cannot raise the estimate, which makes the update a variant that generalizes standard BKT rather than a strict reduction of it. Whether that observation function is trustworthy depends on the learner: mean evidence separated ability tiers cleanly (0.25 weak / 0.67 mixed / 0.77 strong) while within-tier correlation with true mastery was only r ≈ 0.15 / 0.48 / 0.41, leaving the traced state an agent's belief about the learner rather than a calibrated measurement. - **Neural and hybrid models:** [[neural-symbolic-knowledge-tracing]] combines symbolic reasoning with [[machine-learning|neural networks]]; [[explainable-probabilistic-kt]] advances interpretable probabilistic models - **Hypergraph memory networks:** [[thymen-temporal-hypergraph-knowledge-tracing-2026|THyMeN]] augments memory-based tracing (DKVMN) with temporal hypergraph reasoning, modeling dynamic higher-order interactions among concepts that co-occur within multi-skill questions - **Dialogue-based KT:** [[huang-interpretable-knowledge-tracing-2026]] adapts knowledge tracing for conversational tutoring - **LLM-enhanced:** [[xie-hillm-cd-2026|HiLLM-CD]] uses LLMs for automated concept tree construction and hierarchical proficiency inference - **Semantic, recommendation-oriented KT:** [[exrec-exercise-recommendation-knowledge-tracing-2025|ExRec (Ozyurt, Almaci, Feuerriegel and Sachan, 2025)]] grounds the *input* rather than the architecture: an LLM annotates each question with solution steps and knowledge concepts aligned to the Common Core State Standards for Mathematics, contrastive learning aligns question, solution-step and concept embeddings (with false negatives removed by pre-clustering concept variants such as "interpreting a bar chart" and "reading information from a bar graph"), and a KC-calibration loss lets the tracer predict a concept-level knowledge state directly instead of inferring one by running the model over every question in that concept. The calibrated tracer then serves as the reinforcement-learning environment for exercise recommendation, where a model-based value estimation initialises the critic from the tracer itself. Across four tasks on XES3G5M averaged over 2,048 test students, non-RL baselines gave marginal or negative knowledge gains, value-based continuous methods beat policy-based ones, and the model-based value estimate improved them consistently — most sharply on the weakest-concept task, where the target changes at every step. Reported gains are percentage-of-maximum knowledge improvement, not learning outcomes, and the pipeline depends on generated solution steps whose quality the tracer inherits. - **Outcome-based knowledge tracing (OKT):** [[pradeesh-outcome-knowledge-tracing-affinity-2026|Pradeesh et al. (2026)]] trace student knowledge within Outcome-Based Education systems by treating **course outcomes as the knowledge concepts themselves**, and substitute expert-validated OBE "affinity mappings" between course and program outcomes for attention- or graph-derived concept relations. A Memory Augmented Neural Network (MANN) models how each outcome's attainment impacts others, and domain-adaptive BERT fine-tuning enriches the outcome embeddings (with a GRU backbone beating LSTM). On live [[engineering-education|engineering]]-program LMS data (2,416 students, 966 outcomes) OKT reached 89.81% AUC — outperforming DKT, DKVMN, EKT, and SimpleKT — while giving only competitive results on ASSISTments, confirming the advantage is tied to OBE-specific [[curriculum-design|curriculum]] structure. ### Relationship to other concepts Knowledge tracing is closely related to [[student-modeling]] — while knowledge tracing specifically models cognitive knowledge over time, student modeling is the broader practice of representing all aspects of a learner ([[affective-computing|affective]] state, [[student-engagement|engagement]], preferences). Knowledge tracing feeds into [[adaptive-learning]] and [[personalized-learning]] systems that need to know what to teach next, and into [[intelligent-tutoring]] platforms that use mastery estimates to select appropriate problems. It connects to [[learning-analytics]] for dashboard and intervention design, and to [[cognitive-diagnosis]] for fine-grained skill [[assessment]]. Knowledge-tracing constructs also inform [[simulating-students|simulated students]] — a simulated learner's cognitive state is often formalized with the same mastery/decay dynamics that knowledge tracing models, so [[simulation]] is a way to *generate* the knowledge states that tracing methods normally *infer* from real response data. **A caveat: mastery is not correctness.** [[deceptive-overgeneralization-adaptive-learning-2026|An, McLaren, and Stamper (2026)]] show that BKT's two-state (learned/unlearned) assumption can be violated by *deceptive overgeneralization* — learners can appear mastered yet systematically misapply a skill when a hidden application constraint is violated. This argues for tracing conditional understanding (knowing *when to withhold* an action), not only action correctness, when mastery estimates drive [[adaptive-learning|adaptive]] stopping rules. **A related caveat concerns *how* tracing models are validated versus deployed.** [[schuetze-knowledge-tracing-forgetting-2026|Schuetze, Yan, and Carvalho (2025)]] fit BKT, BKT-with-Forgetting, and the Additive Factors Model to a multi-session successive-relearning dataset and found they reproduce learning trends when fit retroactively to all sessions (acceptable AUC ≈ 0.74–0.79); but under **time-based cross-validation** — training on one session to predict the next, the realistic applied setting — all three overestimate future performance by roughly 47–58%, fail to capture the [[desirable-difficulties|spacing effect]], and can even predict the wrong ordinal ordering across practice conditions. Tellingly, models *without* an explicit forgetting mechanism performed about as well as the forgetting-augmented versions as sessions accumulated, suggesting forgetting was partly absorbed into other parameters (e.g., per-student intercepts in AFM) rather than genuinely modeled. The authors tie this to the learning-versus-performance distinction: popular models conflate high in-the-moment performance with high likelihood of long-term retention. The practical implication is that a tracer that looks good on retrospective fit can mislead the adaptive systems consuming its mastery estimates, arguing for walk-forward evaluation and models that account for retention interval, spacing, and between-session forgetting. **A further caveat concerns the evidence rule that feeds the update.** [[crediting-assisted-work-inflates-mastery-2026|Srivastava (2026)]] ran four update rules over identical ASSISTments 2012–13 event sequences, differing only in how they score rows completed with help, on a confirmatory half of 12,716 students and 985,813 scored events. Reading a hinted or retried row as a failed first attempt predicted later unaided performance best (pooled AUC 0.658); crediting any completion predicted it worst (0.604), barely above a constant that knows only skill difficulty (0.595). The same choice governs the mastery count: crediting completions declared 93.9% of 113,428 student–skill pairs mastered against 72.8% under the strict rule, and the pairs the lenient rule declared ahead of strict went on to 70.9% unaided accuracy against 85.7% where the two agreed, below the 0.744 base rate. A traced state is therefore partly a function of the scoring convention rather than of the learner alone, so a mastery estimate consumed by an [[adaptive-learning|adaptive]] gate should carry the rule that produced it. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[student-modeling]] - [[knowledge-graph]] - [[adaptive-learning]] - [[personalized-learning]] - [[intelligent-tutoring]] - [[learning-analytics]] - [[formative-assessment]] - [[ai-education]] - [[ai-ed-evaluation]] - [[multimodal]] - [[teacher-role]] - [[cognitive-offloading]] - [[llm]] - [[simulating-students]] - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[multimodal-item-parameter-estimation-2026]] - [[educlaw-bench-pedagogical-llm-agents-2026]] - [[huang-interpretable-knowledge-tracing-2026]] - [[thymen-temporal-hypergraph-knowledge-tracing-2026]] - [[learning-engagement-assistant-lea]] - [[llm-cognitive-diagnosis-handwritten-math]] - [[multimodal-knowledge-graph-educational-reasoning]] - [[pattern-kc-programming-recommendation]] - [[proprl-prerequisite-relation-learning]] - [[reinforcement-learning-measurement-model-assessment]] - [[skill-acquisition-without-temporal-info]] - [[xie-hillm-cd-2026]] - [[zerkouk-comprehensive-review-its-2025]] - [[trace-course-grade-prediction-2026]] - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[pradeesh-outcome-knowledge-tracing-affinity-2026]] — Outcome-based knowledge tracing with affinity mapping - [[schuetze-knowledge-tracing-forgetting-2026]] - [[simulating-learner-task-selection]] — Simulating learners' task-selection strategies and system constraints in mastery learning (Noh, Chowdhary, Ooge, Aleven & Borchers 2026) - [[exrec-exercise-recommendation-knowledge-tracing-2025]] — semantically grounded tracing with KC-calibrated states, used as an RL environment for recommendation - [[colearn-agentic-tutor-co-learning-loop-2026]] — CoLearn: An Agentic Tutor that Learns its Learner in a Human-AI Co-Learning Loop - [[crediting-assisted-work-inflates-mastery-2026]] — Which evidence rule decides a mastery claim (Srivastava 2026) --- ## [Cognitive Diagnosis](https://edtechdev.github.io/aied/concepts/cognitive-diagnosis/) > **Cognitive diagnosis** — the inference of a learner's latent knowledge state — the specific concepts, skills, and misconceptions they have or lack — from their responses or behavior. It is the assessment-side counterpart to [[knowledge-tracing]], focused on characterizing *what* a student knows rather than only predicting their next performance. ## Questions to Consider - Cognitive diagnosis infers a learner's latent knowledge state — the specific concepts, skills, and misconceptions they have or lack — from their responses, rather than just predicting their next score. Before reading, what's the difference you'd expect between 'predicting a student's grade' and 'diagnosing what they actually don't understand'? - A key idea is the 'correct answer trap' — where a right answer conceals flawed reasoning. Have you ever been confident a student understood something because they got it right, only to discover a misconception underneath? How could a diagnosis surface that where a score couldn't? - The page distinguishes cognitive diagnosis (a static, fine-grained snapshot of what a learner currently holds) from knowledge tracing (the temporal dynamics of mastery over time). Why would an intelligent tutor need both — to know what's wrong and to know what to teach next? - A design principle here is to separate diagnosis from feedback: LLM tutors confirm correct steps but over-reject valid reasoning and over-validate errors, and accurate diagnosis does not reliably yield actionable feedback. Why might knowing what's wrong still fail to produce a helpful next step? - LLM-era diagnosis extends from multiple-choice to open-ended, handwritten, and conversational work. What might go wrong if an AI diagnoses a misconception from work it can't fully understand — and how would you verify that the diagnosis itself is trustworthy? ## Introduction Whereas knowledge tracing typically estimates a scalar mastery over time, cognitive diagnosis produces a more granular profile: which knowledge components are mastered, which are fragile, and which misconceptions are present. This profile is the substrate for [[personalized-learning]], [[intelligent-tutoring]], and [[adaptive-learning]]. ### How cognitive diagnosis works - **Diagnostic models:** psychometric models (often under [[item-response-theory]] and [[educational-measurement]]) infer latent skill states from patterns of correct and incorrect responses, sometimes via cognitive-diagnosis models that map items to multiple knowledge components. - **Automated model search:** because no single diagnostic model fits every learner, [[machine-learning|AutoML]]-driven approaches (e.g., personalized neural cognitive architecture search) generate diagnostic models for heterogeneous learner profiles — integrating [[multimodal|multi-modal]] educational data to enable dynamic analysis of learning processes and per-learner cognitive diagnosis, rather than relying on static [[summative-assessment|examination]] outcomes and simple statistical indicators ([[personalized-neural-cognitive-architecture-search-2026]]). - **Response data:** diagnosis draws on responses to assessments, hints, [[help-seeking]], and time-on-task — richer signals than raw scores. - **LLM-based diagnosis:** newer approaches use [[llm|large language models]] to diagnose from open-ended or handwritten work, and to identify the specific [[misconceptions]] behind an error (e.g., the "correct answer trap" where a right answer conceals flawed reasoning). Two 2026 results bound how far that diagnosis reaches. [[omniedu-open-educational-foundation-models-2026|OmniEdu (Liang et al., 2026)]] supervised diagnostic reasoning as one of four capabilities in an open 4B/9B/27B family, and knowledge-state diagnosis remained its weakest measured capability — 54.04% at 27B and 53.55% at 9B, near enough that three times the parameters did not close the gap — while [[colearn-agentic-tutor-co-learning-loop-2026|CoLearn (He et al., 2026)]]'s LLM grader correlated with true mastery at r = 0.68 over pooled answers but only r ≈ 0.15 within the weakest ability tier (r ≈ 0.48 mixed, 0.41 strong), so diagnostic reliability tracks the learner's ability level as much as the model's. - **Diagnosing common mistakes at cohort scale, not one response at a time.** [[llm-common-modeling-mistakes-formalisms-2026|Killich et al. (2026)]] reverse the usual direction: rather than diagnosing one learner's error, an [[llm]] proposes candidate bug-fixing transformations that map incorrect formalizations onto correct ones across an entire educational data set, and every candidate is validated algorithmically before it is kept. On 6,106 pairs of correct and incorrect propositional-logic formalizations the workflow discovered 248 clusters of transformations explaining 5,156 pairs (84.44%), against 4,370 (71.57%) for the hand-picked mistakes of the previous state of the art, and it recovered the mistakes a domain expert had identified by hand in the literature. Clustering orders candidates into single-transformation, equivalent-transformation and hierarchical groups, and the resulting correlation graph can be visualized for instructors; the same pipeline transferred to modal logic and regular expressions, where one disjunction-for-conjunction transformation alone covered 98.80% of its 334-pair cluster. It is a route to the misconception inventory a diagnostic model needs before it can be fit. - **Outcome-level diagnosis in OBE curricula:** [[pradeesh-outcome-knowledge-tracing-affinity-2026|Pradeesh et al. (2026)]] diagnose which course outcomes a learner has attained in Outcome-Based Education by treating outcomes as the knowledge concepts, supplying concept relationships via expert-validated OBE affinity mappings between course and program outcomes (an explicit alternative to implicitly learned attention or graph relations), and using a memory-augmented module to estimate how one outcome's attainment impacts others — outperforming DKT, DKVMN, EKT, and SimpleKT baselines (89.81% AUC) on live engineering-program data. - **Diagnosing from instruments built for something else.** [[mechanics-cognitive-diagnostic-physics-2026|Le et al. (2026)]] show a CD model can extract objective-level information from items never written for diagnosis. Mapping FCI, FMCE and EMCS items onto 14 fine-grained learning objectives in introductory mechanics and fitting DINA on 24,394 posttest responses from 807 courses, they found good fit for two of the three instruments (FCI RMSEA2 = 0.033; EMCS = 0.022) and classification accuracy at or above the low-stakes formative benchmark for 19 of the 22 objective–assessment combinations. Attribute structure, not item quality, was the binding constraint: expert coding survived model scrutiny almost intact — DINA proposed revising only 14% of 754 item–objective codings and the coders adopted 20 of them (2.7%) — yet the model could not separate three *conceptually nested* energy objectives (Potential Energy 0.675, Conserve Energy 0.705, Kinetic Energy 0.745) because any two shared about 70% of their items (Jaccard overlap 0.67–0.73), violating DINA's conjunctive independence assumption, while momentum objectives on the same instrument reached 0.820–0.917. Finer attributes also fit better rather than worse: the 14-objective structure improved model fit over the same team's earlier four-broad-skill structure on all three instruments. Item overlap, not coding error, is what caps how finely mastery can be separated. - **Bayesian DINA for personalized learning paths:** [[bayesian-cognitive-diagnosis-personalized-learning-paths|Feng and Huang (2026)]] integrate a Bayesian DINA model (trained on the EdNet dataset, N=5,000) with knowledge space theory and a shortest-remediation-path algorithm to generate personalized learning paths, and empirically test the mediating role of [[cognitive-offloading|cognitive load]] via Hidden Markov Model state transitions (validated on 120 students) — addressing both the sparsity-driven convergence problem of traditional DINA models and the untested psychological mechanism behind personalized-path effectiveness. - **Language-grounded diagnosis in place of ID embeddings.** [[process-grounded-language-cognitive-diagnosis-2026|Liu et al. (2026)]] replace discrete student, exercise and concept identifiers with LLM-built concept schemas and process-grounded evidence, calibrating each student's posterior state from response records. Across three [[math-education|mathematics]] [[online-teaching-and-learning|platform]] datasets the framework reaches 83.51% ACC / 85.37% AUC on XES3G5M and 87.16% ACC on MOOC, with the gain concentrated exactly where classical cognitive-diagnosis models degrade: new concepts (+4.60 ACC over KCD) and missing Q-matrix entries (+4.52). Ablating the structured evidence collapses MOOC accuracy from 87.16% to 78.95%, so the improvement comes from the language-derived structure rather than from model scale. ([[process-grounded-language-cognitive-diagnosis-2026]]) ## Why it matters Accurate diagnosis lets instruction target the actual gaps rather than a global "ability" score — enabling [[automated-assessment]] that explains *why* a student erred and [[feedback|Feedback Loop]] systems that remediate specific [[student-modeling|knowledge states]]. Poor diagnosis produces the inverse: instruction aimed at the wrong concepts. This is why [[psychometrically-aware-ai]] emphasizes diagnostic validity alongside prediction accuracy. ### Relationship to knowledge tracing and intelligent tutoring Cognitive diagnosis sits at the heart of the [[intelligent-tutoring]] architecture and is the assessment-side counterpart of [[knowledge-tracing]]: - **Diagnosis vs. tracing — complementary temporal views.** [[knowledge-tracing|Knowledge tracing]] tracks the *temporal dynamics* of mastery — estimating how a scalar knowledge state evolves across exercises and predicting the next response. Cognitive diagnosis produces the *static, fine-grained snapshot* of which knowledge components, skills, or misconceptions a learner currently holds. A tutor needs both: knowledge tracing to sequence what to teach next, cognitive diagnosis to know *what* is actually wrong. [[item-response-theory|IRT]]- and [[educational-measurement|measurement]]-based diagnostic models, and cognitive-diagnosis models that map items to multiple components, instantiate the diagnostic side. - **LLM-era diagnosis.** [[llm|LLMs]] extend diagnosis from multiple-choice responses to open-ended, handwritten, and conversational work, identifying the specific [[misconceptions]] behind an error (e.g., the "correct answer trap" where a right answer conceals flawed reasoning). [[xie-hillm-cd-2026|HiLLM-CD]] uses LLMs for automated concept-tree construction and hierarchical proficiency inference, bridging diagnosis and tracing. [[privacy-preserving-multi-llm-federated-cognitive-diagnosis-2026|Boyapati et al. (2026)]] push this further by federating diagnosis across multiple commercial LLM APIs with ε-local differential privacy, showing that accurate, privacy-preserving diagnosis is feasible without any model seeing raw student data. - **Separating diagnosis from feedback is a design principle.** LLM tutors reliably confirm correct steps but over-reject valid reasoning and over-validate errors — and accurate diagnosis does not reliably yield actionable [[feedback]]. ITS design should therefore separate a diagnostic component from the feedback/[[scaffolding]] component ([[yasir-llm-tutoring-agents-2026]]). ## Connections Cognitive diagnosis connects to [[knowledge-tracing]], [[student-modeling]], [[educational-measurement]], and [[assessment]]. Its insights feed [[intelligent-tutoring]] and [[adaptive-learning]], and LLM-era work links it to misconception identification in [[intelligent-tutoring|AI Tutoring]]. ## Connected Concepts - [[knowledge-tracing]] - [[knowledge-graph]] - [[student-modeling]] - [[educational-measurement]] - [[item-response-theory]] - [[assessment]] - [[intelligent-tutoring]] - [[adaptive-learning]] - [[personalized-learning]] - [[automated-assessment]] - [[learning-analytics]] ## Connected Articles - [[llm-cognitive-diagnosis-handwritten-math]] — Benchmarking LLMs for Diagnosing Cognitive Skills from Handwritten Math - [[correct-answer-trap-misconceptions]] — The Correct Answer Trap - [[llm-misconception-difficulty-easy-trap]] — The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty - [[llm-student-misconception-identification]] — LLM identification of student misconceptions - [[student-math-competence-clustering]] — Clustering for Modeling Student Mathematical Competence - [[moon-cognitive-agent-compilation-problem-solver-modeling-2026]] — Cognitive Agent Compilation for Explicit Problem Solver Modeling - [[eeg-familiarity-automated-assessment-2026]] — Automating Learner Assessment: EEG-Based Familiarity Prediction - [[educlaw-bench-pedagogical-llm-agents-2026]] — EduClaw-Bench: diagnosing from simulated learners - [[huang-interpretable-knowledge-tracing-2026]] — Interpretable knowledge tracing - [[xie-hillm-cd-2026]] — HiLLM-CD: LLM concept trees + hierarchical proficiency inference - [[yasir-llm-tutoring-agents-2026]] — Separating diagnosis from feedback in LLM tutors - [[skill-acquisition-without-temporal-info]] — Diagnosing skill acquisition without temporal information - [[zhang-ct-ai-training-test-2026]] — Computational Thinking in AI Training Test (CTAT) - [[li-dbagent-llm-educational-agent-cs-2026]] — LLM-based educational agent (DBagent) in CS education - [[bayesian-cognitive-diagnosis-personalized-learning-paths]] — Bayesian cognitive diagnosis for personalized learning paths - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[personalized-neural-cognitive-architecture-search-2026]] — AutoML personalized neural cognitive architecture search for learner profiles - [[pradeesh-outcome-knowledge-tracing-affinity-2026]] — Outcome-based knowledge tracing with affinity mapping - [[privacy-preserving-multi-llm-federated-cognitive-diagnosis-2026]] — Privacy-preserving heterogeneous multi-LLM federated diagnosis - [[llm-common-modeling-mistakes-formalisms-2026]] — Mining common modeling mistakes at scale with LLM-generated, algorithmically validated bug-fixing transformations (Killich et al. 2026) - [[mechanics-cognitive-diagnostic-physics-2026]] — Mechanics Cognitive Diagnostic: DINA-based diagnosis of 14 learning objectives from existing physics concept inventories (Le et al. 2026) - [[exrec-exercise-recommendation-knowledge-tracing-2025]] — LLM knowledge-concept annotation and calibrated concept-level knowledge states - [[llm-distractor-generation-student-reasoning-2026]] — misconception-based distractors as a diagnostic item-design task - [[misconception-acquisition-dynamics-llms-2026]] — where the error enters the solution is the diagnostic bottleneck - [[pivot-generative-video-tutors-stem-2026]] — From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning - [[colearn-agentic-tutor-co-learning-loop-2026]] — CoLearn: An Agentic Tutor that Learns its Learner in a Human-AI Co-Learning Loop - [[omniedu-open-educational-foundation-models-2026]] — OmniEdu: Open Foundation Models for Learning and Teaching --- ## [Simulating Students](https://edtechdev.github.io/aied/concepts/simulating-students/) > **Simulating students** — using LLM-based agents to model learner behavior, cognition, and social dynamics for educational research, design, and training. Simulated students let researchers evaluate pedagogical approaches, model diverse learner profiles, test educational AI before deployment, and train teachers — tasks that are difficult, slow, or ethically constrained to do systematically with real learners. ## Questions to Consider - If you had to build an AI 'student' to practice your teaching on, what would make it convincing to you — and why might a system that always gives the right answer actually be a poor stand-in for a real learner? - The page calls the mismatch between a capable AI's perfect answers and a real student's imperfect ones the 'competence paradox.' Where have you seen this tension in your own experience with AI, and what do you think it takes to make a simulated learner genuinely realistic? - What are the [[ethics|ethical]] and practical reasons you might prefer testing a tutoring system or [[curriculum-design|curriculum]] on simulated students rather than real ones — and what validity risks do you suspect that trade introduces? - A simulated student can be 'epistemically faithful' without looking superficially human. Before you read on, what distinction do you imagine between a believable surface and a truthful model of what a learner actually knows? - How would you decide whether a finding produced by simulated students should be trusted enough to change how you teach real people? ## Introduction Simulated students are a [[research-methods-aied|methodological]] tool: agents that stand in for real learners so that tutoring systems, curricula, and instructional strategies can be evaluated and iterated without recruiting cohorts of human students. [[llm|Large language models]] have made this paradigm far more scalable and linguistically realistic than the rule-based simulated learners that preceded them, while also introducing new validity challenges. AI-mediated approximations sharpen the question the page cares about: whom the simulation represents, and what the teacher is meant to notice. In mathematics teacher education, [[bondurant-shaughnessy-ai-pedagogies-practice-2026|text-based simulated student work and chatbot partners]] used in rehearsal extend the approximations of practice that already organize professional training. One chatbot study produced four distinct questioning profiles, yet pre-service teachers' self-assessments did not align with the interaction quality observers recorded, which makes automated post-rehearsal feedback that increased probing and exploring questions a useful but insufficient guide. The simulator, the candidate and the feedback all have to be read together. ### Why simulate students - **Evaluating pedagogy:** testing instructional approaches across many learner profiles in a controlled, repeatable way. - **Modeling diverse learners:** capturing variation in cognitive levels, learning styles, [[prior-knowledge|prior knowledge]], and [[misconceptions]] that is hard to assemble in a real cohort. - **Testing educational AI:** validating tutoring and [[assessment]] systems before live deployment, and generating training data. - **[[teacher-education|Teacher training]]:** letting instructors practice tutoring and classroom management with simulated, often imperfect, learners. ### The core challenge: realistic imperfection The defining difficulty of student simulation is that LLMs are trained to be "helpful assistants" that produce correct, polished answers. Yet real students are imperfect — they make characteristic mistakes, hold misconceptions, and learn gradually. A simulated student that answers perfectly (or too randomly) is not a valid model of a learner. Research frames this as the **competence paradox**: broadly capable LLMs asked to emulate partially knowledgeable learners produce unrealistic error patterns and learning dynamics. Addressing it requires constraining the simulation so it reflects a genuine epistemic state — what the learner knows, how errors are structured, and how state evolves — rather than the model's full competence. Techniques include cognitive prototypes grounded in [[knowledge-graph]] or [[knowledge-tracing]] models, explicit epistemic state specifications, and state-transition models of learning rather than simple persona-conditioned role-play. ### Fidelity over surface realism Validity is the central concern: a simulated student is only useful if its behavior is **epistemically faithful** — reflecting the intended learner's knowledge state — not merely linguistically plausible. Research warns against [[ai-sycophancy|sycophancy]], where a "simulated student" simply agrees with the tutor rather than exhibiting the misconceptions it was meant to embody. This connects to [[trust-calibration]] and to the broader problem of evaluating whether an agent genuinely models a construct rather than reproducing surface behavior. ### Connection to the knowledge base Simulating students sits at the intersection of [[simulation]], [[student-modeling]], and [[knowledge-tracing]]. It is a distinct use of [[generative-ai]] in education (modeling learners rather than tutoring them) and an application of [[agentic-ai]] multi-agent systems. It supports [[intelligent-tutoring]], [[adaptive-learning]], [[personalized-learning]], and [[teacher-role]] development, and it overlaps with patient simulation for [[professional-training|professional training]] (e.g., [[special-education]] and [[medical-education|medical education]] contexts). At the paradigm level, [[agent-based-educational-science-2026|Zhang, Jiang and Tang (2026)]] extend this beyond evaluation: their position paper argues that educational science suffers a structural mismatch between its theory and its data and method apparatus, and proposes agent-based educational science, in which agents model learners, teachers, [[parents-and-families|parents]] and peers while the simulation apparatus itself — not the individual agent — serves as the research instrument, generating time-extended developmental trajectories and counterfactual designs that are slow, costly or ethically infeasible to run in classrooms, with empirical data recast as calibration, validation and boundary conditions. The paper reports no empirical validation of any of this: Student Development Agents are specified conceptually only, and the authors' own review concedes that [[llm|LLMs]] still miss inter-individual variability and that validating generative social simulation remains the field's unresolved challenge, so the claim that simulation would change how evidence and replication work is a proposal rather than a demonstrated result. ### Simulating students vs. student modeling The key distinction is between **representing a real learner** and **generating a synthetic learner**. [[student-modeling]] is the practice of building a computational representation of an actual student — what they know, feel, and need — so that adaptive systems can personalize instruction for *that* learner. Simulating students, by contrast, *creates* fictional learners on demand, not to serve a real individual but to stand in for a cohort so pedagogy and [[ai-technologies|AI systems]] can be tested offline. The two are complementary rather than competing. A high-fidelity simulated student typically *contains* a student model (an epistemic state, a misconception set, an [[student-engagement|engagement]] profile) and draws on the same constructs that [[student-modeling]] and [[knowledge-tracing]] formalize. The shared validity challenge is the same in both: the representation must faithfully reflect a learner's true state rather than the system's default behavior. But the *purpose* differs — student modeling diagnoses a real learner to act on them; simulation fabricates learners to test or train. This is why simulated-student research is increasingly used to audit AI (see below) while student-modeling research remains oriented toward live [[adaptive-learning]] and [[personalized-learning]]. ### Two-axis fidelity: behavioral match and guidance responsiveness [[studentsim-llm-student-simulators|StudentSim (Yang et al., 2026)]] formalizes the validity problem as two requirements that must hold together: **behavioral fidelity (F)** — how well a simulator matches a student's own responses — and **guidance responsiveness (R)** — how reliably it updates toward where tutor guidance leads. Its [[benchmark]], StudentSimEval, casts public learner corpora (chess, second-language English writing, [[math-education|mathematics]]) into a standardized per-student protocol on which any simulator is fit and scored on held-out records. A two-stage **pooled-then-specialized** pipeline (a shared pool of behavioral patterns plus a per-student adapter) yields a family of simulators strong on both axes, outperforming domain-specific state-tracking (weak on R) and prompt-only LLM role-play (weak on F). This gives the field a concrete, two-axis vocabulary for judging whether a simulated learner is genuinely useful rather than merely plausible — and, as a proof of concept, a frozen StudentSim used as the reward in a chess-tutor [[reinforcement-learning]] loop produced tutors that experts rated as more accurate, better-guided, and more personalized than those trained with a frontier-LLM-simulator reward or no RL at all. ### Describing, not simulating: when a single agent beats a simulated cohort A recurring question is whether simulating a distribution of learners is the best way to predict how real students will fare. [[ai-web-agents-lesson-design-2025|Wang, Mitchell & Piech (2025)]] supply a striking counter-result: for **evaluating an online [[learning-design|learning experience]] before students engage** — predicting dropout and completion and giving design feedback — a **single "describing" [[agentic-ai|web agent]]** that autonomously walks through the lesson and produces a rich description of the [[student-experience|student experience]] outperformed directly simulating a population of students. Their simulated students (persona-conditioned agents sampled to match a predicted completion rate) exhibited far less behavioral range than real learners — across 100 agents on five test lessons they reproduced only about **4% of the paths** real students took — and offered little insight into lesson difficulty, while being substantially more computationally expensive. The single describe-then-predict pipeline (agent-generated descriptions fed to an [[llm]] to predict outcomes) achieved the best dropout-distribution prediction on a massive global CS1 course (mean JSD 0.060, beating every baseline). This is a productive boundary result for the field: simulating a *distribution* of students may be unnecessary — and even counterproductive — when the goal is outcome prediction or design critique, because a faithful description of the experience carries more signal than a narrow slice of simulated trajectories. It also suggests simulation's role may be bounded to questions (e.g., auditing an AI's treatment of diverse profiles, or training teachers) where covering genuine learner variation matters more than aggregate outcome prediction. ### Authentic-data student models and interactive practice Two 2026 threads sharpen the practical value of simulation. First, **authentic-data student models** — [[teachlm-post-training-llms-education|TeachLM]] trains a student model on 100,000 hours of real one-on-one tutor–student interactions (with rigorous anonymization), producing synthetic learners that enable fast, scalable, reproducible multi-turn evaluation of tutor behavior; this addresses the low authenticity and diversity of purely prompt-engineered student simulators. Second, **interactive instructional simulacra** — [[educasim-cs1-instructional-practice|EducaSim]] uses generative agents (with personas, course-grounded memories, and an LLM-as-judge speech oracle) to simulate a small-group section for teachers-in-training, adding runnable-code and voice interaction plus structured post-session feedback and self-reflection, and demonstrates low-cost, positive-uptake experiential [[pedagogy|teaching practice]] at the scale of massive [[online-teaching-and-learning|online courses]]. Both point to simulation serving not only evaluation but hands-on teacher preparation. A third 2026 thread concerns *how* a simulator is built rather than what it is used for, and it converges on one conclusion: prompting sets a ceiling that training removes. [[swim-student-writing-simulation-2026|SWIM (Do, Kontak and Sachan, 2026)]] reformulates student writing simulation as proficiency-conditioned essay generation and compares rubric-grounded prompting, supervised fine-tuning on real score-essay pairs, and GRPO with an automated-essay-scoring-derived Proficiency Alignment Reward, scoring each generated essay against its target trait profile with Quadratic Weighted Kappa. Prompting gave limited control even to strong proprietary models (best average trait QWK 0.577 for Claude Sonnet, 0.422 for GPT-5.4, near zero for prompting an open 7B model), it aligned content-oriented traits far better than form (0.695 against 0.458), and it produced an idealised high-proficiency population, with a mean normalized Overall score of 0.74 against 0.58 for real students and a median length of 304 words against 167. Supervised fine-tuning moved a 7B model to 0.474 ± 0.023 and GRPO to 0.618 ± 0.005, the gains held on two independent scorers the policy never trained against, and the trained models recovered the human score and length distributions without any length supervision; authentic low-proficiency form stayed hardest, since trained models recovered syntax but wrote too few spelling and grammar errors while prompting simulated weakness mainly through superficial corruption. The same lesson arrives from the opposite direction in [[misconception-acquisition-dynamics-llms-2026|Liu et al. (2026)]], who instruction-tuned three small models to hold algebra misconceptions: data *composition* decided the outcome, because a single misconception overgeneralised and degraded correct solving until correct examples were mixed in, several misconceptions trained jointly without cost, and no misconception was acquired from final-answer-only supervision at any data size, leaving step-level solution traces as the binding requirement. ### Auditing AI with simulated students Beyond evaluating pedagogy, simulated students serve as a **test harness for auditing AI systems themselves** — a controlled way to probe how an AI behaves across diverse learner profiles before it touches real students. [[lopez-pernas-llm-appropriate-student-support-2026|López-Pernas et al. (2026)]] illustrate this: they generated 4,500 synthetic student vignettes with three LLMs to audit whether current [[llm|large language models]] can act as prescriptive [[learning-analytics]] recommenders, finding limited sensitivity to student need and sharp cross-model inconsistency. Using simulated cohorts to stress-test an AI's recommendations (rather than only to train or evaluate tutors) is a growing role for the paradigm, closely tied to [[ai-ed-evaluation|evaluating AI in education]] and to [[equity-in-ai-education]] when the audit is meant to surface disparate treatment across learner types. A second audit register is [[metacognition|metacognitive]] and affective. [[meds-math-education-digital-shadows-2026|MEDS (Esposito et al., 2026)]] is a 28,000-record dataset — 2,000 synthetic personas for each of 14 [[llm|models]], every one run both as a human persona and as a baseline assistant — that records accuracy on 18 [[k-12|high-school]] [[problem-solving|math problems]] alongside [[self-report-measures|self-reported]] confidence and the [[self-efficacy]] and [[anxiety-and-stress|math anxiety]] scores the learners it stands in for would report. Its audit signal is the calibration gap: the Qwen family and Ministral 3B asserted confidence above 0.90 while accuracy stagnated near 0.55, while Grok 4.1 Fast, DeepSeek Chat and several Mistral Small variants were underconfident and Ministral 14B and Anita 24B stayed reasonably aligned. The same runs expose a quieter failure of fidelity: human-mode personas yielded wide, plausible score distributions, but baseline assistants returned near-identical, confident, low-anxiety answers — a default self-portrait rather than a simulated learner's. This extends the recommender audit above by probing what a model claims about its own competence and affect, and it sharpens the page's validity warning from an unexpected direction: because MEDS personas are weighted by construction rather than sampled from a real population, its authors present the dataset as an observational resource for auditing prompt-conditioned [[generative-ai|GenAI]] behavior and explicitly not as a stand-in for real [[student-experience|student data]]. ### Simulating collaborative and social dynamics Simulation also extends beyond individual learners to reproducing the social dynamics of [[collaborative-learning|collaborative learning]]. **Participant-specific LLM agents** — [[llm-agents-collaborative-problem-solving-simulation-2026|Fang (2026)]] fine-tuned LLaMA 3.2-3B agents on individual participants' dialogue data to represent each participant in collaborative problem solving simulations, combining sliding-window memory with summarized memory embeddings to preserve both local turn-taking and thematic continuity, and probabilistically selecting speakers and thematic codes from empirical distributions. Using [[network-analysis|Epistemic Network Analysis (ENA)]], the simulated dialogues were statistically indistinguishable from real ones (ENA distance 0.17, well within the 95th-percentile null threshold; permutation p = 0.65), validating that [[agentic-ai|LLM agents]] can reproduce both turn-taking dynamics and thematic code trajectories of real [[problem-solving|collaborative problem solving]]. The 2026 durable-skills work inverts the usual direction of simulation. Instead of simulating the student to audit a system, the system simulates the *teammates* to assess the student: an Executive LLM generates every AI partner's turns in a 30-minute group task, holds the scoring rubric, and steers the conversation to manufacture occasions for the target skill to appear ([[durable-skills-measurement-ai-teammates-2026|Globerson et al., 2026]]). Across 373 conversations from 188 participants, skill-matched steering raised ratable evidence to 92.4% for project management and 85% for conflict resolution, significantly above unconstrained independent agents, while the AI evaluator was calibrated against two human raters whose own inter-rater Kappa was only 0.45–0.64 — a useful reminder that a simulator's ceiling is set by the agreement humans can reach on the construct. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[simulation]] - [[student-modeling]] - [[knowledge-tracing]] - [[cognitive-diagnosis]] - [[agentic-ai]] - [[pedagogical-agent]] - [[intelligent-tutoring]] - [[adaptive-learning]] - [[personalized-learning]] - [[generative-ai]] - [[llm]] - [[learning-analytics]] - [[teacher-role]] ## Connected Articles - [[llm-student-simulation-teacher-insights]] — Can LLMs Simulate Human Learners? Teachers' Insights - [[llm-student-simulation-misconception-faithfulness]] — Simulating Students or Sycophantic Problem Solving? - [[history-aware-student-simulation]] — History-Aware Profiles for Student Simulation - [[llm-educational-simulation-adhd]] — LLM-Based Educational Simulation and Student Persona Stability - [[simulating-students-java-programming-errors-llms]] — Simulating Students' Java Programming Errors - [[adaptive-virtual-patient-psychotherapy-training]] — Adaptive Virtual Patients for Psychotherapy Training - [[medeasy-ai-standardized-patients]] — MedEasy: AI Standardized Patients - [[simulating-students-diverse-cognitive-levels-2025]] — Embracing Imperfection: Simulating Diverse Cognitive Levels - [[simulating-students-llm-review-2026]] — Simulating Students with LLMs: A Review - [[valid-student-simulation-llm-2026]] — Toward Valid Student Simulation - [[agentschool-multi-agent-simulation-education-2026]] — AgentSchool: Multi-Agent Simulation for Education - [[inside-llm-student-simulator-reasoning-2026]] - [[teachlm-post-training-llms-education]] — TeachLM: fine-tuned authentic student model for synthetic dialogues - [[educasim-cs1-instructional-practice]] — EducaSim: generative agents simulate a CS1 section for teacher practice - [[bondurant-shaughnessy-ai-pedagogies-practice-2026]] — Responsible Integration of AI into Pedagogies of Practice in Mathematics Teacher Education - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[llm-agents-collaborative-problem-solving-simulation-2026]] — Fine-tuned participant-specific LLM agents reproducing collaborative problem solving dialogues (Fang 2026) - [[studentsim-llm-student-simulators]] — StudentSim: Training LLM-based Student Simulators - [[ai-web-agents-lesson-design-2025]] — AI Web Agents: a single describing agent beats simulating a distribution of students for predicting dropout and design critique - [[simulating-learner-task-selection]] — Simulating learners' task-selection strategies and system constraints in mastery learning (Noh, Chowdhary, Ooge, Aleven & Borchers 2026) - [[durable-skills-measurement-ai-teammates-2026]] — Toward Scalable Measurement of Durable Skills - [[agent-based-educational-science-2026]] — Toward Agent-based Educational Science: the simulation apparatus as the research instrument (Zhang, Jiang & Tang 2026) - [[meds-math-education-digital-shadows-2026]] — MEDS: a 28,000-record audit of math performance, confidence and anxiety in simulated students and AI assistants - [[swim-student-writing-simulation-2026]] — prompting versus SFT versus reward-based training for a student writing simulator - [[misconception-acquisition-dynamics-llms-2026]] — what has to be in the training data before a simulator holds a misconception at all - [[llm-distractor-generation-student-reasoning-2026]] — trace-level analysis of how models simulate incorrect student reasoning - [[colearn-agentic-tutor-co-learning-loop-2026]] — CoLearn: An Agentic Tutor that Learns its Learner in a Human-AI Co-Learning Loop --- ## [Intelligent Tutoring](https://edtechdev.github.io/aied/concepts/intelligent-tutoring/) > **AI tutoring / intelligent tutoring** — the use of AI to provide personalized, adaptive, scalable instructional support: from classical Intelligent Tutoring Systems (ITS) that model student knowledge, adapt instruction, and scaffold [[problem-solving]], to conversational and agent-based tutors built on [[llm|LLMs]]. Effectiveness hinges on [[pedagogy|pedagogical]] design ([[scaffolding]], feedback quality, autonomy balance) rather than the model alone — see [[measuring-llm-tutors-teach-vs-solve]] and [[socratic-method]]. ## Questions to Consider - An AI tutor's effectiveness is said to hinge on pedagogical design — scaffolding, feedback quality, the balance of autonomy and guidance — rather than on the underlying model. If you had to pick, which design choice most determines whether a student actually learns? - Classical ITS are praised for precision and transparency but criticized for inflexibility, while LLM tutors are flexible but can hallucinate or bypass learning. Where do you think the right balance lies, and why? - Large-scale field evidence found students tried an AI tutor but rarely used it productively — engaging it in only a fraction of their mistake sessions, often with bare answers. Why might students get access to a powerful tutor yet not use it in ways that help them learn? - The page contrasts traditional ITS that model a learner and adapt instruction with open-ended conversational tutors. What is gained and lost when a tutor can hold a natural conversation but doesn't know exactly why it made a decision? - One study found AI created a 'productive slowdown' — more time per question — but only improved learning when embedded in a mastery workflow that made mistakes consequential. What does that suggest about structuring effort so it pays off? - If hint design can inadvertently enable students to bypass learning, what would a well-designed hint system look like that supports struggle without giving away the answer? ## Introduction AI tutoring encompasses the use of artificial intelligence — particularly [[llm|large language models]] and structured Intelligent Tutoring Systems — to provide personalized, adaptive, and scalable instructional support to learners. AI tutors take many forms: conversational tutors that engage in Socratic dialogue, scaffolded feedback systems that guide problem-solving, [[adaptive-learning|adaptive learning platforms]] that personalize content sequencing, and agent-based tutors that maintain long-term [[student-modeling|learner models]]. The effectiveness of AI tutoring depends critically on pedagogical design choices — scaffolding, [[ai-feedback-quality|feedback quality]], and the balance between [[agency|autonomy]] and guidance — rather than on the underlying model alone. Historically, **[[mishra-control-vs-agency-history-2025|Mishra et al.]]** locate ITS within AIED's lineage from 1960s-70s expert systems and Anderson's ACT/ACT-R cognitive tutors, whose structured control contrasted with Papert's [[constructivist|constructionism]]. ## ITS in the learner-modeling family Intelligent tutoring is the classic *application-side* member of the [[student-modeling|learner modeling and adaptive instruction]] family. Its canonical architecture — domain model, [[student-modeling|student model]], and pedagogical model — is precisely the "model a learner, then adapt instruction" pipeline the family describes. ITS consume the learner representations produced by [[knowledge-tracing]] and [[cognitive-diagnosis]] to select problems and scaffold guidance, which is why tutoring is so tightly coupled to those modeling methods. Within the family, ITS sits alongside [[adaptive-learning|adaptive learning]] (the real-time adaptation mechanism) and [[personalized-learning|personalized learning]] (the broader goal) as the platforms that turn learner models into instruction. ## ITS vs. LLM-based tutoring The emergence of [[llm|LLMs]] has created a productive tension in the tutoring field. Traditional Intelligent Tutoring Systems (ITS) offer precision and transparency — you know exactly why the system made a particular decision — but lack flexibility. LLM tutors offer natural dialogue and broad knowledge but can hallucinate, over-scaffold, or bypass learning entirely. Modern [[research-methods-aied|research]] increasingly explores **hybrid approaches** that combine structured ITS components with LLM flexibility. [[reddig-maclellan-personalized-feedback-llm-2026|Reddig, Arora & MacLellan (2025)]] demonstrate this concretely in the Apprentice Tutor College Algebra ITS: supplying GPT-4 the tutor's interface structure and Bayesian [[knowledge-tracing]] skill estimates within the prompt lifted logical-error diagnosis from 40% to 81% on factoring (and error identification overall to 87.8%) and produced ~66% error-targeted hints — direct evidence that embedding an LLM in an ITS's structured framework grounds generation, curbs hallucinated diagnoses, and yields context-aware corrective feedback, even though roughly a third of the hints remained too general, incorrect, or answer-revealing. Intelligent Tutoring Systems represent one of the oldest and most researched areas of [[ai-education|AI in education]]. Unlike general-purpose LLM tutors, ITS traditionally use structured approaches: domain models (what to teach), [[student-modeling|student models]] (what the learner knows), and pedagogical models (how to teach). These components enable fine-grained tracking of student progress, misconception diagnosis, and adaptive sequencing. ### Key ITS research - **Micro-randomized trials of a tutoring platform:** [[ai-tutoring-micro-rct-gcse-science-2026|Harrison et al. (2026)]] ran a four-week multisite individually randomized evaluation of Medly in GCSE [[biology-education|Biology]], [[chemistry-education|Chemistry]] and [[physics-education|Physics]], with 644 of 929 students completing post-testing; allocation to the platform produced Hedges' g = 0.33 (95% CI 0.18 to 0.48) against business-as-usual [[self-directed-learning|self-directed]] revision, with positive estimates in all three subjects and no evidence of differential impact by disadvantage (Pupil Premium g = 0.28 versus non-Pupil Premium g = 0.35). - **[[lak2026-hint-button-unproductive-use|Hint button research]]** shows that traditional ITS hint design can inadvertently enable bypass strategies, calling for more sophisticated [[scaffolding]] approaches. - **[[deeptutor]]** provides a fully [[open-source]] agentic tutoring framework with citation-grounded tutoring and difficulty-calibrated [[automated-question-generation|question generation]]. - **[[huang-interpretable-knowledge-tracing-2026|Interpretable Knowledge Tracing]]** addresses the opacity problem by producing interpretable cognitive quantities from LLM logits. - **Engagement and structure, not capability, are the binding constraints (large-scale field evidence):** [[one-click-away-khanmigo-two-year-school-experiment-2026|Khanmigo (Oreopoulos & Low 2026)]] — a two-year cluster [[rct]] in 18 middle schools — found 96% of students tried the AI tutor but the median engaged it in only ~17% of mistake sessions (mostly bare answers or prompt clicks), so gains (~0.06–0.08 SD) matched practice without AI. [[making-ai-tutoring-productive-mastery-math-2026|NUMI (Oreopoulos et al. 2026)]] found AI created a "productive slowdown" — more time per question and improved post-mistake recovery — but only reliably improved delayed learning when embedded in a mastery workflow that made mistakes consequential. The lesson: AI tutoring's value depends on **getting students to use it productively** and structuring it so effort pays off, more than on raw model capability.([[virtual-tutoring-computer-assisted-learning-takeup-2026]]) - **Outcome-based knowledge tracing for OBE tutors.** [[pradeesh-outcome-knowledge-tracing-affinity-2026|Pradeesh et al. (2026)]] trace student mastery of Outcome-Based-Education course outcomes with a recurrent model that replaces learned concept relations with [[curriculum-design|curriculum]]-validated OBE "affinity mappings" between course and program outcomes and adds a Memory Augmented Neural Network for cross-outcome impact. It reached 89.81% AUC on live engineering-program data, beating DKT, DKVMN, EKT, and SimpleKT baselines — a mastery signal aligned to a program's own outcomes that can drive adaptive problem selection in outcome-based tutoring. - **Validate mastery models on future sessions, not retrospective fit.** The mastery estimates an ITS uses to select problems may not generalize across time: [[schuetze-knowledge-tracing-forgetting-2026|Schuetze, Yan, and Carvalho (2025)]] found that BKT, BKT-with-Forgetting, and AFM reproduce learning trends only when fit retroactively to all sessions of a successive-relearning dataset — but under **time-based cross-validation** (training on one session to predict the next, the actual tutoring setting) they overestimate performance by 47–58%, fail to capture the spacing effect, and can mis-order practice conditions. Tutors relying on such tracers should be evaluated walk-forward and consider retention interval and between-session forgetting rather than assuming mastery persists. - **Diagnosing mistakes on unstructured, handwritten-solution problems.** [[yin-arthur-ai-teaching-assistant-engineering-econ-2026|Arthur (Yin et al. 2026)]] extends ITS-style feedback to Calculated Formula Questions in an Engineering Economics course, a domain where pen-and-paper solutions lack the structured digital data that previous AI tutoring work relied on. Rather than a general-purpose LLM, each question gets its own dedicated [[machine-learning|XGBoost]] "solution diagnosis backbone" — trained on curated, random-masking-augmented graded submissions to predict instructor rubric mistake labels from students' submitted numbers alone (precision 0.81, recall 0.79) — paired with a dialogue-based interaction that iteratively requests intermediate answers only when diagnosis confidence drops below a 0.8 threshold. It is a concrete case of the knowledge base's principle that a knowledge-grounded classifier can handle solution diagnosis even where full written solutions are unavailable. - **Turn-level adaptive tutoring with a background evaluator.** The MeduAI-SP [[pedagogical-agent|tutor agent]] ([[ai-standardized-patient-scaffolding-medical-2026|Yang et al., 2026]]) exemplifies turn-level adaptive tutoring: a background evaluator monitored each student utterance against a 30-item OLDCARTS-linked history checklist plus communication criteria and triggered prompts only when warranted (24.1% of 4,815 messages). Corpus analysis exposed where such support is most needed — student talk was 62.2% history of present illness but only 4.3% empathic/supportive communication and 2.6% physical [[summative-assessment|examination]] — suggesting ITS-style [[medical-education|clinical]] tutors should target underused consultation components and late-encounter reasoning phases rather than symptom gathering alone, and that simulated-patient fidelity is high enough for outcome research (only ~0.68% of patient utterances showed clear fidelity problems). - **Prompt-based micro-personalization of an LLM tutor (2026):** [[prompt-engineering-personalization-ai-teaching-assistant-2026|Basu, Kakar & Goel (2026)]] personalize the Jill Watson [[llm]]/[[rag]] [[teacher-role|teaching]] assistant at the level of individual questions entirely through prompt engineering — conditioning responses on six learner dimensions ([[metacognition|metacognitive]] self-assessment, abstraction, verbosity, perception, processing, understanding) and classifying each query's cognitive demand with a [[cognitive-diagnosis|Bloom's Taxonomy]] classifier, yielding 96 learner profiles with no model retraining. NLP analysis of 2,910 responses and a human study found measurable, perceptible differences in response abstraction, verbosity, complexity, and processing style aligned with intended effects. It demonstrates response-form (rather than content) adaptation as a scalable, modular route to [[adaptive-learning|adaptive behavior]] in a deployed tutor. ### Key ITS concepts - **[[knowledge-tracing]]** — modeling what a student knows over time (Bayesian, deep learning, IRT-based) - **[[cognitive-diagnosis]]** — fine-grained assessment of which knowledge components and [[misconceptions]] a learner holds; the assessment-side counterpart to [[knowledge-tracing]] that feeds the tutor's pedagogical decision - **[[student-modeling]]** — broader learner representation including affect, engagement, and misconceptions - **[[adaptive-learning]]** — systems that personalize content sequencing based on learner state - **[[scaffolding]]** — providing just enough support to enable progress without giving away answers - **[[desirable-difficulties|productive struggle]]** — letting students wrestle with difficulty rather than over-helping - **[[feedback|feedback loops]]** — ITS feedback cycles that diagnose, guide, and verify ### Historical context The ITS field has produced landmark systems (Cognitive Tutors, Andes, AutoTutor) and continues to evolve. The [[zerkouk-comprehensive-review-its-2025|Zerkouk et al. comprehensive ITS review]] catalogs this evolution. The tension between structured ITS and open-ended LLM tutoring is explored in [[correct-answer-trap-ai-tutor|the correct answer trap]] research and [[rethinking-scaffolding-llm-tutors|rethinking scaffolding for LLM tutors]]. ### AI tutoring with LLMs: practical guidance For instructors deploying AI tutors and developers building them, the knowledge base's findings translate into concrete practice: **Evaluate tutors on whether they teach, not just solve.** A model that tops a solving leaderboard is not necessarily a good tutor — task-solving ability and learning-supportive behavior correlate only partially (r ≈ 0.42), and several models shift rank when scored on pedagogy. Report and scrutinize **solving and pedagogy scores separately**, and prioritize tutors that score on guiding questions, calibrated hints, and non-disclosive scaffolding over those that produce fast answers.([[measuring-llm-tutors-teach-vs-solve]])([[ai-tutoring-quality-k12-methodologies-2026]]) **Design for pedagogical structure, not frequency.** The educational payoff of AI tutoring depends on *how* the tool is used and designed, not on how often it is used. Instructor-designed tutors scoped to course objectives, learner proficiency, and a curated knowledge base outperform unstructured general-purpose [[conversational-ai|chatbot]] use.([[instructor-designed-ai-tutors-foreign-language-sdt-2026]]) A preregistered randomized field experiment tested that claim against the realistic alternative rather than an AI-free baseline: 86 students in an online-MBA corporate-finance module were assigned to a tutor grounded in the module's own lectures, readings and problem sets or to a holdout that kept standard resources with consumer AI freely available, and tutored students gained 6.63 more points of 55 (95% CI [+1.95, +11.31], p = .007) while the share of short written answers at relational quality or better rose from 8% to 49% against 8% to 27%. Channel was the null margin — voice nearly doubled interaction density and cost 2.8× more to deliver, yet weekly mastery differed by −0.01 points of 11 (p = .98) — which makes delivery format an adoption and engagement lever rather than a learning technology. The holdout did not close the gap either: the nine students reporting external-AI use gained no more than the eight reporting none (+10.3 versus +12.8 of 55, p = .45). **Use iterative live evaluation to keep improving.** Because LLMs are opaque, treat evaluation as the engine of improvement: instrument a small set of quality and [[student-engagement|engagement metrics]], run live experiments on models, [[prompt-engineering|prompting]], personalization, and agents, and let data drive changes — the same discipline Khan Academy applies to its [[k-12]] tutor (Khanmigo).([[ai-tutoring-quality-k12-methodologies-2026]]) **Support the learner's autonomy, competence, and relatedness.** AI tutors work best when they feel like a safe, structured practice space rather than an answer machine. Provide immediate, nonjudgmental [[feedback]]; scope the tutor to the learner's level so competence is achievable; and preserve learner agency by keeping the tutor a complement to (not a substitute for) other instruction.([[instructor-designed-ai-tutors-foreign-language-sdt-2026]]) **Guard against answer disclosure.** The central failure mode of LLM tutoring is giving the answer away, which inflates immediate performance while undermining durable learning. Use Socratic prompting, calibrated hints, and non-disclosive scaffolding — and measure outcomes on unassisted, [[transfer-of-learning|transfer]] tasks, not just in-tool performance.([[measuring-llm-tutors-teach-vs-solve]])([[socratic-method]]) **Separate diagnosis from feedback.** LLM tutors reliably confirm correct steps but over-reject valid-but-suboptimal reasoning and over-validate incorrect solutions — and accurate diagnosis does not reliably yield actionable feedback.([[yasir-llm-tutoring-agents-2026]]) The coupling also cuts the other way: [[reddig-maclellan-personalized-feedback-llm-2026|Reddig et al. (2025)]] found GPT-4 still crafted relevant, general-but-correct feedback ~74% of the time even after misdiagnosing an error (by restating the concept or expected answer format), yet almost all factually incorrect feedback followed a wrong diagnosis — so error identification remains the crux of actionable tutoring. A hybrid architecture works best: let a knowledge-grounded classifier handle solution diagnosis while the LLM focuses on open-ended scaffolding and dialogue. - **Fairness auditing is becoming part of tutor evaluation.** EduFair-Bench holds the simulated student fixed and varies only demographic attributes, showing that tutoring quality is not uniform across learners and that reward-tuning pedagogy does not automatically remove those differences ([[edufair-bench-pedagogical-fairness-llm-tutors-2026]]). ### AI tutoring as a spectrum of relational intensity A unifying lens from [[turano-ai-tutoring-not-a-monolith-2026|the Stanford SCALE / NSSA brief (Turano et al. 2026)]] reframes AI tutoring as **not a monolith but a spectrum** defined by *relational intensity* — the depth and consistency of the human connection between student and tutor. The brief maps models from fully human-led tutoring (in-person or remote), through **human-led with AI support** (AI assists the tutor behind the scenes) and **AI-led with human support** (a human oversees and intervenes), to **AI-only tutoring** (no direct human oversight). The central finding: as direct human relationships decrease, the evidence base becomes thinner, and unresolved questions accumulate about student safety, developmental impact, and long-term efficacy. - **AI augments, does not replace, high-impact tutoring.** The brief's core message is that AI is best used to *enhance tutor effectiveness and educator capacity* within high-impact tutoring (regular school-day sessions, small-group ratios ≤ 1:4, well-trained consistent tutors, data-driven instruction, vetted materials, strong student-tutor relationships) — not to substitute for the human-led relationship that drives [[learning-gains|learning gains]]. This aligns with the knowledge base's broader finding that pedagogically designed AI tutors outperform general-purpose chatbots ([[stanford-evidence-base-ai-k12-2026|Stanford evidence base]], [[genai-higher-education-systematic-review-2026|umbrella review]]). - **Dosage, not model capability, is the binding constraint.** [[turano-ai-tutoring-not-a-monolith-2026|The brief]] reports that AI-led tutoring inherits the dosage evidence base (~90 minutes weekly) only when scheduled, supervised, and protected within the school day; in a study of 181,000 students on a supplemental math platform, only 5% reached the recommended minutes and 41% never logged on, with teacher/school/district factors explaining 57% of usage variance. This converges with the field-experiment evidence above ([[one-click-away-khanmigo-two-year-school-experiment-2026|Khanmigo]], [[making-ai-tutoring-productive-mastery-math-2026|NUMI]]) that engagement and integration, not raw capability, determine whether AI tutoring helps. - **Human oversight improves engagement and alignment but not dosage.** The brief's RCTs found human check-ins raised elementary students' engagement with an AI platform but did not reach the dosage associated with gains nor improve reading achievement. The practical implication: schedule and supervise AI tutoring on the same terms as human-led tutoring, and design for enforcement of engagement — the opt-in requirement an on-device AI tutor reintroduces is exactly what live school-day tutoring eliminates. - **The relationship is the irreducible element.** [[turano-ai-tutoring-not-a-monolith-2026|The brief]] emphasizes that AI does not yet replicate human relationships, and that relationship-building (especially consistent tutor-student pairings) improves engagement, attendance, motivation, and outcomes. Developers describe using AI's strengths rather than replacing relationships — a design stance echoed in [[hazra-safetutors-pedagogical-safety-2026|SafeTutors]] and [[kar-mathbuddy-affective-math-tutoring-2025|affective tutoring]] research. - **Safety and privacy are non-negotiable guardrails.** For direct-to-student AI, the brief calls for strict student data privacy safeguards, [[guardrails]] on student safety, and attention to the depth of unmonitored interaction — open questions remain about AI companions' effects on developing minds and prosocial development. This relational-intensity framing is the "not a monolith" counterpoint to the field: evaluating any AI tutor should begin by asking *which conditions of effective tutoring it reproduces, changes, or drops* — rather than asking whether "AI tutoring works" in the abstract. It ties the tutoring page to [[k-12]] policy ([[educational-policy-ai]]), [[privacy]], [[pedagogical-safety]], and the human-in-the-loop concerns throughout. ### Practical design and development guidance **Design for learning, not just performance.** The strongest causal finding is that unguarded AI tutors raise assisted practice performance but *reduce* unassisted learning — the [[ai-misuse-learning-harm|performance–learning gap]]. Guardrail against answer-copying by scaffolding **hints instead of answers** (require a student attempt before revealing output), and verify gains on unassisted, closed-book measures rather than in-tool performance.([[generative-ai-guardrails-harm-learning]])([[genai-performance-vs-learning]]) **Make hints genuinely productive, not bypassable.** Classic hint designs can enable "button-through" strategies that skip learning. Prefer hints that reveal reasoning steps incrementally (Socratic prompting) over hints that directly supply the next answer, and preserve productive struggle rather than over-helping.([[lak2026-hint-button-unproductive-use]])([[rethinking-scaffolding-llm-tutors]]) **Keep the [[human-in-the-loop-ai|human in the loop]].** Let teachers author or curate the problem sets and misconception prompts the tutor draws on, and surface the tutor's reasoning so its decisions are auditable. Interpretable [[knowledge-tracing]] and explicit, external didactic layers make LLM tutor behavior traceable and reproducible.([[huang-interpretable-knowledge-tracing-2026]])([[didactical-teacher-assistant-dimensional-modeling]]) **Model the learner, not just the dialogue.** Attach structured [[student-modeling]] and [[knowledge-tracing]] components to LLM dialogue so the system can adapt difficulty and diagnose misconceptions from evidence rather than responding fluently but blindly — quality depends on both the base model and how it is adapted.([[educlaw-bench-pedagogical-llm-agents-2026]]) **Start from open tooling where possible.** Open-source agentic tutoring frameworks (e.g. [[deeptutor]]) lower the barrier to a citation-grounded, difficulty-calibrated tutor you can inspect and extend.([[deeptutor]]) Effective tutoring requires continual adaptation: [[zhang-tutormoments-2026|Zhang et al. (2026)]] evaluate whether LM tutors adapt to learners' evolving understanding at teacher-annotated decision points. They find frontier models default toward over-helpfulness and rarely push for rigor, and that evaluation-aware prompting improves but does not fully solve adaptivity. A systematic view of the RL-driven branch of this adaptation comes from [[riedmann-reinforcement-learning-education-review-2026|Riedmann, Schaper & Lugrin (2025)]], whose review of 89 RL-in-education studies finds classical RL policies more consistently effective than Deep RL and RL adaptation delivering significant gains more often for guidance-related tasks (hints, [[feedback]]) than for content scheduling — evidence that how an ITS adapts scaffolding matters as much as which algorithm it uses. - **[[deceptive-overgeneralization-adaptive-learning-2026|Deceptive overgeneralization (An et al. 2026)]]** shows ITS mastery stopping rules (BKT, 95% threshold) can end practice before learners learn *when to withhold* a skill: learners who overgeneralized misapplied actions on first "do-not-act" items at 61.5%–100%, and targeted refrain-practice with constraint-naming [[feedback]] reduced this to near-floor. Correctness-based mastery inference is necessary but not sufficient for ITS adaptivity. - **Graph-based ITS for dynamic domains.** [[graph-its-adaptive-algorithms-2026|A graph-based intelligent tutoring system]] combines an Evolving Knowledge Space Graph with [[generative-ai|generative AI]] content creation and Bayesian knowledge propagation — which showed the highest knowledge gains — supporting adaptive learning in dynamic curricula. - **Rule-integrated LLM tutoring for procedural domains.** Looi, Liu, and Sun (2026) tackle the inconsistency and pedagogical opacity of LLM tutors in primary [[math-education|mathematics]] through a design science study of a rule-guided system organized around a three-layer architecture — *diagnosis → intent selection → constrained response generation*. They formalize the distinction between **rule-guided scaffolding** (governed by auditable, replicable rules) and **ad-hoc scaffolding** (helpful moves difficult to audit or replicate). Evaluated via persona-based simulated dialogues and a classroom pilot with 40 Grade 5 students, rule-guided scaffolding improved interactional consistency, reduced premature answer-giving and early closure, and sustained cognitive engagement — while the classroom pilot surfaced interactional complexities, fragmented inputs, and attentional fluctuations that [[simulation]] missed. This is a concrete blueprint for [[guardrails|guardrailing]] [[llm]] tutors in well-defined procedural domains. - **[[agentic-ai|Multi-agent]] tutoring and [[automated-assessment|automated assessment]].** Multi-agent tutoring systems are being benchmarked with synthetic, trace-based evaluation. ASTRA supports alone-tutor, pair-tutor, and pair-multiagent configurations with socially differentiated agents, enabling reproducible analysis of interaction and participation balance in [[cs-education|introductory programming]]. In parallel, context-aware prompting automates coding of [[collaborative-learning|collaborative problem-solving]] skills from process data, supporting large-scale tutoring assessment. - **Design with the learner's values, not just for learner performance.** [[ko-hughes-vsd-student-centered-its-2026|Value Sensitive Design work]] with community college students and instructors in developmental math shows that the [[ethics|ethical]] dimensions of ITS are not separable add-ons: engaging students and instructors directly produced 16 value-aligned features spanning [[explainable-ai|explainability]] (comprehension-check interpretation, learning-path connection, communicating the model's confidence), [[human-in-the-loop-ai|learner control]] (control over re-assessment, review, pace, and AI-assistance level), and [[privacy]] (data-repurposing control, permission prompts for sharing [[learning-analytics|learning analytics]] and [[affective-computing|affective]] states). The study frames a persistent design tension — [[agency|student agency]] vs. system-guided scaffolding — and notes that most institutions disable adaptive tutoring entirely, so value-informed design also depends on how (and whether) ITS AI capabilities are actually deployed. ## Connected Concepts - [[scaffolding]] - [[adaptive-learning]] - [[cognitive-diagnosis]] - [[llm]] - [[student-modeling]] - [[knowledge-tracing]] - [[feedback]] - [[ai-feedback-quality]] - [[socratic-method]] - [[personalized-learning]] - [[self-regulated-learning]] - [[generative-ai]] - [[ai-education]] - [[metacognition]] - [[pedagogical-safety]] - [[privacy]] - [[educational-policy-ai]] - [[guardrails]] - [[k-12]] - [[speech-and-voice-technologies]] ## Connected Articles - [[typology-generative-ai-tools-education-2026]] — Educator-reported tutoring and chatbot tools in a 2026 typology - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[ko-hughes-vsd-student-centered-its-2026]] — Value-sensitive design of student-centered ITS in community college developmental math - [[prompt-engineering-personalization-ai-teaching-assistant-2026]] — Prompt-engineering micro-personalization of an AI teaching assistant (Basu, Kakar & Goel 2026) - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[mishra-control-vs-agency-history-2025]] — Traces ITS lineage from 1960s-70s expert systems to cognitive tutors - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[yasir-llm-tutoring-agents-2026]] — Benchmarking LLM feedback agents with KG ground truth (Yasir et al. 2026) - [[instructor-designed-ai-tutors-foreign-language-sdt-2026]] — Instructor-Designed AI Tutors in University Foreign Language Education (Self-Determination Theory) - [[measuring-llm-tutors-teach-vs-solve]] — Measuring Whether LLM Tutors Teach or Solve - [[ai-tutoring-quality-k12-methodologies-2026]] — Methodologies for Improving the Quality of AI Tutoring in K-12 Education - [[zhang-tutormoments-2026]] — When Help is Unhelpful: evaluating AI tutors for productive struggle - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[stanford-evidence-base-ai-k12-2026]] — Tutoring-specific vs. general-purpose AI: evidence on durable learning outcomes - [[hazra-safetutors-pedagogical-safety-2026]] — SafeTutors and pedagogical safety - [[kar-mathbuddy-affective-math-tutoring-2025]] — MathBuddy affective math tutoring - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — Virtual tutoring with CAL: an experiment in take-up and learning - [[learnlm-improving-gemini-learning]] — LearnLM: improving Gemini for learning - [[teachlm-post-training-llms-education]] — TeachLM: post-training LLMs with authentic learning data - [[rule-integrated-llm-tutoring-primary-math-2026]] — Rule-guided vs ad-hoc scaffolding in an LLM tutoring system for primary mathematics (Looi et al. 2026) - [[pradeesh-outcome-knowledge-tracing-affinity-2026]] — Outcome-based knowledge tracing with affinity mapping - [[reddig-maclellan-personalized-feedback-llm-2026]] - [[schuetze-knowledge-tracing-forgetting-2026]] - [[riedmann-reinforcement-learning-education-review-2026]] - [[yin-arthur-ai-teaching-assistant-engineering-econ-2026]] - [[ai-tutoring-micro-rct-gcse-science-2026]] — Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomized Trials of an AI Tutoring Platform in GCSE Science - [[misconception-acquisition-dynamics-llms-2026]] — tutor models that acquire many student misconceptions without losing correct solving - [[ai-tutor-modality-randomized-field-experiment-2026]] — When AI Tutors Speak: Evidence from a Randomized Field Experiment --- ## [Adaptive Learning](https://edtechdev.github.io/aied/concepts/adaptive-learning/) > **Adaptive learning** — AI-driven educational systems that adjust content, pacing, and instructional strategies based on individual learner characteristics and performance. Adaptive learning is the operational goal of much [[ai-education|AI in education]] [[research-methods-aied|research]]: using [[student-modeling|student models]] to personalize instruction. ## Questions to Consider - 'Adaptive,' 'personalized,' 'individualized,' and 'customized' learning are often used interchangeably — but research suggests they are not the same. What do you assume each word means, and where might those assumptions be wrong? - An adaptive system adjusts content and difficulty based on a model of what you know. What could go wrong if that model rests on shallow or unreliable signals about your learning? - A key finding is that systems inferring mastery from correct answers can stop practice too early — before you learn when to withhold an action. Can you think of a skill where being 'correct' repeatedly still left you unprepared for a real situation? - Over-adaptation can remove the productive struggle students need to learn deeply. If AI keeps making things easier the moment you struggle, what exactly does the learner lose? - Meta-analysis suggests the adaptation mechanism — not the specific tool generation — drives [[learning-gains|learning gains]]. If the 'how' matters more than the 'which tool,' what should you look for when choosing adaptive software? - LLM-based tutors can now adapt language and explanation style, not just difficulty. When does personalizing the way something is explained help learning, and when might it quietly undermine the learner's own agency? ## Introduction ### Core mechanisms - **Measure-model-adapt loop:** [[knowledge-tracing]] estimates what the student knows, [[student-modeling]] represents the learner, and the system adapts difficulty, content, and [[feedback]] accordingly. - **Personalization at scale:** [[personalized-learning]] systems use adaptive algorithms to serve unique learning paths for each student. [[deeptutor]] and [[ai-powered-personalized-learning-elementary-fractions-2026|elementary fraction tutors]] demonstrate adaptive personalization in practice. - **Content sequencing:** [[adaptive-pretesting-retention|Adaptive pretesting]] and [[adapt-adaptive-lesson-plan-transformer|lesson plan transformers]] optimize the order and type of content presented. - **ITS integration:** [[intelligent-tutoring|Intelligent tutoring systems]] are the canonical adaptive learning platform, combining diagnosis with adaptation. - **AutoML-driven profiling and diagnosis:** Traditional educational models struggle to process multi-source, heterogeneous learning-behavior data, which limits learner profiling and diagnostic model development. A personalized neural cognitive architecture search framework driven by automated [[reinforcement-learning|machine learning]] integrates [[multimodal|multi-modal]] educational data with heterogeneous methods, generating diagnostic models tailored to heterogeneous learner profiles and supporting dynamic rather than static analysis of learning processes. ### Effectiveness evidence The knowledge base documents mixed evidence: adaptive systems improve outcomes when adaptation is grounded in reliable [[student-modeling|student models]], but poorly-calibrated adaptation can harm learning. [[personalized-learning|Personalization research]] distinguishes effective adaptation from superficial customization. [[khalifeh-redefining-personalized-learning-ai-2026|Systematic reviews]] find that "adaptive," "personalized," "individualized," and "customized" learning are used inconsistently — so effect sizes depend heavily on how adaptation is operationalized, and the field calls for a unified framework. ### The AI era: LLM-based adaptation and its risks [[generative-ai|Generative AI]] has expanded what adaptive systems can do — conversational [[agentic-ai|agentic]] tutors, [[rag]]-grounded content, and [[llm]]-driven [[intelligent-tutoring|tutoring]] adapt not only problem difficulty but language and explanation style (e.g., [[learnmate2-llm-adaptive-learning|LearnMate-2]], [[deeptutor]], [[chudziak-ai-math-tutoring-platform|multi-agent adaptive tutoring]]). However, LLM-based adaptation introduces new risks: without reliable [[student-modeling|student models]], adaptation may be based on shallow signals; over-adaptation can reduce the productive struggle students need (see [[desirable-difficulties]], [[cognitive-offloading]]); and the balance between personalizing and preserving learner [[agency]] is an open design question (see [[agentic-ai|agentic AI]]). A learner-requested variant of adaptation runs without any [[student-modeling|student model]] at all: in Sidorkin's (2026) graduate course the readings adjusted only when students asked follow-up questions to reframe, deepen, simplify or localize them, and comprehension-oriented requests reliably produced denser scaffolding (3.4x to 8.7x more definitional markers than baseline text), which is why requiring at least three follow-up questions per reading turned the material into an interaction. It also relocates the adaptive burden onto the learner: adaptation here happens only if the student knows what to ask for. ### Relationship to personalized learning and intelligent tutoring Adaptive learning is frequently conflated with [[personalized-learning|personalized learning]], but they differ. **Adaptive learning** is the *mechanism* — real-time adjustment of content, pacing, and difficulty based on a learner model. **Personalized learning** is the *broader goal* of tailoring the whole learning experience to an individual, of which real-time adaptation is one implementation. Adaptive systems are the canonical *means* toward personalization. [[intelligent-tutoring|Intelligent tutoring]] is the classic *platform*: ITS combine diagnosis (student modeling, knowledge tracing) with adaptation, and LLM-based tutors adapt conversationally. Together with [[personalized-learning|personalized learning]], adaptive learning is an application-side member of the [[student-modeling|learner modeling and adaptive instruction]] family — consuming the learner representations that [[student-modeling|student modeling]], [[knowledge-tracing]], and [[cognitive-diagnosis]] produce. ### Research evidence - **[[meta-analysis-systematic-review|Meta-analytic]] evidence on adaptive + AI tools.** [[burneo-can-edtech-close-learning-gaps-2026|A World Bank meta-analysis]] of 14 [[rct|RCTs]] pools adaptive computer-assisted learning, intelligent tutoring, and generative AI on a common scale, estimating an average learning gain of ~0.125 sd with no significant difference between the two technology generations — evidence that the adaptation mechanism, not the specific tool generation, drives gains. - **Adaptive algorithms compared in dynamic domains.** [[graph-its-adaptive-algorithms-2026|Graph-based ITS research]] compares multiple adaptive learning algorithms (including Bayesian knowledge propagation and intuitionistic fuzzy logic) in a graph-based knowledge representation framework for dynamic curricula. - **RL as an adaptation mechanism, empirically mapped.** [[riedmann-reinforcement-learning-education-review-2026|Riedmann, Schaper & Lugrin (2025)]] synthesize 89 RL-in-education studies and find adaptation splits into content-related (instructional sequencing/content scheduling, n = 53) and guidance-related (hints, [[feedback]], activity selection, n = 36) mechanisms — with RL showing statistically significant superiority over baselines more often for guidance-related adaptation than for content scheduling. They recommend model-free RL for adaptive learning and caution that classical RL outperformed Deep RL in the reviewed studies. - **Correctness-based adaptivity can stop practice too early.** [[deceptive-overgeneralization-adaptive-learning-2026|An, McLaren, and Stamper (2026)]] found that adaptive systems inferring mastery from correctness risk terminating practice before learners encounter contexts where the learned action should be withheld — leaving deceptive overgeneralization undetected. They recommend including "do-not-act" detector tasks before mastery stopping rules trigger, so adaptation tests conditional understanding (knowing when to withhold an action), not only correctness. - **Engagement profiles as adaptation targets.** [[an-goel-self-directed-modeling-2026|An, Hammock & Goel (2025)]] traced 315 online learners building 822 models in VERA and classified their engagement into Observation, Construction, and Exploration profiles, finding that learners tend to progress from construction-focused behavior toward fuller, hypothesis-driven Exploration while Observation persists across phases. They argue adaptive and personalized design should recognize these profiles and target feedback (e.g., recommending similar models or supporting deeper conceptual understanding) to move surface-level observers toward more integrative, full-cycle modeling. - **The gain came from sequencing, not from a smarter tutor.** [[chung-personalized-ai-tutors-llm-reinforcement-learning-2026|Chung et al. (2026)]] trained a personalized tutor with LLM-guided reinforcement learning and deployed it in a five-month Python course across ten Taipei [[k-12|high schools]], randomizing 770 students between adaptive and fixed easy-to-hard problem sequences. Adaptive sequencing raised the in-person, unassisted [[summative-assessment|final exam]] score by 0.156 SD (0.150 SD with controls) — while mediation analysis attributed the effect almost entirely to engagement (0.185 SD via time on task, 0.149 SD via attempts) rather than to easier or harder material, and gains were largest for beginners and lower-tier schools. The adaptive lever was the order of practice, not the quality of the chat. ## Connected Concepts - [[online-teaching-and-learning]] — Online Teaching and Learning - [[knowledge-tracing]] - [[personalized-learning]] - [[intelligent-tutoring]] - [[student-modeling]] - [[scaffolding]] - [[cognitive-diagnosis]] - [[llm]] - [[learning-analytics]] - [[higher-ed]] - [[k-12]] - [[formative-assessment]] - [[behaviorism]] - [[ai-technologies]] — Umbrella: AI technologies and techniques (models, LLM training, robotics, RAG, agentic) - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[causal-modeling-competency-assessment-2026]] — Causal Modeling of Support Interventions for Student Competency Assessment - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[adaptive-ai-scaffold-collaborative-problem-solving-2026]] - [[learning-context-framework-context-aware-ai-education-2026]] - [[mejeh-fromm-srl-adaptive-learning-feedback-2026]] - [[banihashem-ai-srl-systematic-mapping-review-2025]] - [[simon-student-engagement-adaptive-learning-2026]] — Systematic review of student engagement in adaptive learning platforms - [[zhan-chapman-genai-cs-education-2026]] - [[ai-enhanced-pbl-chatgpt-scaffolding-2026]] - [[ai-student-engagement-online-learning-review-2025]] - [[ai-online-education-engagement-satisfaction-2026]] - [[interactive-online-learning-ai-2025]] - [[ai-decision-support-online-learning-assessment-2026]] - [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] — Ontology-based layered hybrid knowledge model for personalized e-learning - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — Virtual tutoring with CAL: an experiment in take-up and learning - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[chudziak-ai-math-tutoring-platform]] — Adaptive/personalized multi-agent math tutoring (Chudziak & Kostka 2025) - [[khalifeh-redefining-personalized-learning-ai-2026]] — Redefining personalized learning: systematic review - [[deeptutor]] - [[ai-powered-personalized-learning-elementary-fractions-2026]] - [[adaptive-pretesting-retention]] - [[adapt-adaptive-lesson-plan-transformer]] - [[zerkouk-comprehensive-review-its-2025]] - [[vargas-situated-learning-ai-review-2024]] - [[prezenski-human-centered-ai-aided-learning]] - [[fowlin-operationalizing-learning-principles-ai]] - [[stanford-evidence-base-ai-k12-2026]] — Tutoring-specific AI calibrated to learner readiness vs. general chatbots - [[multilingual-adaptive-learning-nigeria-2026]] — AI-Based Adaptive Learning Platform for Multilingual Low-Resource Contexts - [[context-based-ai-secondary-chemistry-2026]] — Context-based 7E + AI instruction in secondary chemistry - [[bin-bakheet-adaptive-ai-stem-deep-learning-2026]] — Adaptive AI-based STEM program for deep learning - [[lodge-adaptive-capabilities-genai-future-2026]] — Adaptive capabilities for assuring quality learning in a gen AI-integrated future (Lodge et al. 2026) - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[bayesian-cognitive-diagnosis-personalized-learning-paths]] — Bayesian cognitive diagnosis for personalized learning paths - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[burneo-can-edtech-close-learning-gaps-2026]] — Meta-analysis pooling adaptive + AI-enabled tools across 14 RCTs - [[conversational-agents-business-simulation-gaming-2026]] — CAIS-GBL framework for AI conversational agents in business simulation games (Wenzel et al. 2026) - [[personalized-neural-cognitive-architecture-search-2026]] — AutoML personalized neural cognitive architecture search for learner profiles - [[alsheikh-mapping-ai-integration-higher-education-2026]] — Systematic review: adaptive pathways among the leading higher-ed AI integration use cases - [[an-goel-self-directed-modeling-2026]] - [[riedmann-reinforcement-learning-education-review-2026]] - [[sidorkin-ai-generated-course-readings-2026]] — Learner-requested adaptation of AI-generated readings, with no student model (Sidorkin 2026) - [[chung-personalized-ai-tutors-llm-reinforcement-learning-2026]] — Adaptive problem sequencing beats fixed sequencing: +0.156 SD on an unassisted exam, mediated by engagement rather than difficulty (Chung et al. 2026) - [[colearn-agentic-tutor-co-learning-loop-2026]] — CoLearn: An Agentic Tutor that Learns its Learner in a Human-AI Co-Learning Loop --- ## [Personalized Learning](https://edtechdev.github.io/aied/concepts/personalized-learning/) > **Personalized learning** — tailoring educational experiences to individual [[student-modeling|learner profiles]], including prior knowledge, learning pace, preferences, and [[affective-computing|affective]] states. AI enables personalization at scale, though the gap between *system personalization* and *learner-perceived personalization* remains an open measurement challenge. Alongside [[adaptive-learning|adaptive learning]] and [[intelligent-tutoring|intelligent tutoring]], it is one of the application-side members of the [[student-modeling|learner modeling and adaptive instruction]] family — consuming learner models to adapt instruction. ## Questions to Consider - When you think of 'personalized learning,' do you imagine content tailored to a learner's pace, or to their chosen goals? The page says these are deeply different (uniform outcomes via varied paths vs. diverse outcomes). Which do you value more, and why? - The page distinguishes personalized learning (the goal) from adaptive learning (one mechanism). Can you think of personalization that doesn't involve real-time adaptation—and does it still count? - A system can adapt without the learner ever feeling recognized. When have you experienced being 'personalized to' without feeling genuinely known? What's the difference? - The page flags that over-personalization can strand learners in low-expectation tracks. How might well-intentioned AI tailoring accidentally lower the ceiling for a learner? - Personalization needs detailed learner data; privacy needs data minimization. Where do you draw the line between 'enough data to adapt' and 'so much that the learner is exposed'? - What would an AI need to remember about you across sessions to genuinely personalize your learning—and what are the risks of it remembering those things? ## Introduction Tailoring educational experiences to individual learner profiles, including [[prior-knowledge|prior knowledge]], learning pace, preferences, and affective states. AI enables personalization at scale, though the gap between *system personalization* and *learner-perceived personalization* remains an open measurement challenge. - **[[mishra-control-vs-agency-history-2025|Mishra et al.]]** distinguish two forms of personalization with deep historical roots — uniform outcomes reached via varied paths (Skinner's [[teacher-role|teaching]] machines to Khan Academy-style mastery tutoring) vs. diverse, learner-chosen outcomes — mapping onto the field's control-vs-agency tension. ## Architectures for AI-Driven Personalization ### Longitudinal Memory (PersonaVLM → Education) Nie et al. (2026) developed a [[multimodal]] long-term memory architecture (PersonaVLM) that maintains persona consistency across interactions. Mapped to education, this enables tutoring systems that remember a learner's [[misconceptions]], preferred explanations, and progress history across sessions—addressing a critical deficit in stateless [[conversational-ai|chatbot]] tutors. ### Agent-Native Personalization Substrate (DeepTutor) Ma et al. (2026) design every [[deeptutor]] feature to share a common personalization substrate, rather than bolting personalization onto reactive tools. This architecture ensures cross-modality coherence: the same learner profile drives [[problem-solving|problem solving]], [[automated-question-generation|question generation]], and collaborative writing. ### Multi-Agent Social Personalization (MAIC) Yu et al. (2024) personalize not only content but *social context*. Classmate archetypes (Class Clown, Deep Thinker, Note Taker, Inquisitive Mind) create varied peer-learning dynamics matched to individual learner needs. ### AutoML for Learner Portraits Personalization is a central objective for improving educational quality, yet processing multi-source heterogeneous learning-behavior data remains a challenge. A personalized neural cognitive architecture search framework, driven by automated [[reinforcement-learning|machine learning]], builds learner portraits and generates diagnostic models for heterogeneous learner profiles, integrating multi-modal data to move beyond static examination outcomes. ## Relationship to adaptive learning and intelligent tutoring Personalized learning is often conflated with [[adaptive-learning|adaptive learning]], but they are not the same. **Adaptive learning** refers to the *mechanism* — a system adjusting content, pacing, and difficulty in real time based on a learner model. **Personalized learning** is the *broader goal* — tailoring the full learning experience (content, pathways, pacing, preferences, goals) to an individual, of which real-time adaptation is one implementation. Adaptive systems are a *means* toward personalization, but personalization can also be achieved through static learner profiles, choice-based pathways, or human-tutor tailoring that does not adapt in real time. [[intelligent-tutoring|Intelligent tutoring]] sits in between: ITS are the canonical *adaptive* platforms that deliver personalized instruction through structured student modeling, while [[llm]]-based tutors personalize conversationally. All three are the application-side members of the [[student-modeling|learner modeling and adaptive instruction]] family — they consume the learner representations produced by [[student-modeling|student modeling]], [[knowledge-tracing]], and [[cognitive-diagnosis]] to decide what to teach next. The distinction matters for evaluation: studies that label a system "adaptive," "personalized," or "individualized" interchangeably (see below) can obscure whether the claimed benefit comes from real-time adaptation, learner choice, or content tailoring. ## Measurement Challenges - **System vs. perceived personalization** — A system can adapt without the learner feeling recognized - **Longitudinal validity** — Personalization benefits may decay if profiles become stale or overfit - **[[equity-in-ai-education|Equity]] risks** — Over-personalization can strand learners in low-expectation tracks ## Personalization and assessment Personalization and [[assessment]] are tightly coupled in AI-driven learning. Adaptive personalization depends on ongoing [[formative-assessment|formative]] measurement of what a learner knows (via [[knowledge-tracing]], [[student-modeling]], and [[cognitive-diagnosis]]) to decide what to adapt next — so the reliability of the [[assessment]] signal directly constrains the quality of personalization. Conversely, when [[summative-assessment|summative assessment]] is personalized per-learner, [[bias-mitigation|fairness]] and comparability become harder to establish. The knowledge base's [[research-methods-aied|research]] warns against over-adapting to shallow or noisy signals: [[adaptive-learning|adaptive]] systems that mis-measure a learner can personalize in ways that reduce learning rather than support it, and AI-native students whose self-assessment is unreliable (an "absent cognitive baseline") are harder to model accurately. ## Personalization in the AI era The strongest evidence that this concern is not hypothetical comes from a [[personalization-paradox-adaptive-learning-emotions-2026|three-wave longitudinal study of 486 Chinese undergraduates (Li, Lin & Qiu, 2026)]], which found that the more personalized students perceived their AI-adaptive environment to be, the *lower* their [[self-regulated-learning|self-regulated learning]] — the "personalization paradox." Shifts in academic emotions carried most of the effect: encountering the adaptive environment predicted less enjoyment and more anxiety and boredom, and those emotional changes together accounted for roughly half of the association between personalization and reduced self-regulation. [[ai-literacy|AI literacy]] buffered the damage, weakening the negative emotional association to non-significance at high literacy. Personalization therefore appears to buy adaptive fit at a cost to the learner's own [[regulation|regulatory]] activity, and the study points to emotional experience — not only cognitive load — as the channel through which that cost is paid. Reinforcement learning is a distinct mechanism for personalization, and [[riedmann-reinforcement-learning-education-review-2026|Riedmann, Schaper & Lugrin (2025)]] map its empirical track record: their [[meta-analysis-systematic-review|PRISMA]] review of 89 RL-in-education studies finds RL personalization concentrated in [[higher-ed]] and [[math-education]], with adaptation implemented mainly as content scheduling (n = 53) or guidance-related personalization such as hints and feedback (n = 36). They report that RL policies beat non-adaptive baselines most often on guidance-related adaptation and on [[affective-computing|affective]] variables (63% of tested studies), and that learning gain — especially normalized learning gain — was the most effective reward source — practical guidance for designing reward signals that personalize toward genuine learning rather than [[student-engagement|engagement]]. Bernstein and Sibia (2026) sharpen a distinction between interest personalization and expertise personalization: interest-matched GenAI analogies were reported as more engaging and memorable but not uniformly more trusted, and some students preferred the generic technical explanation even when the analogy matched their stated interest, for self-sufficiency and completeness ([[student-reception-genai-analogies-computing-2026]]). Their design recommendation is to personalize through source-domain structure and to ask students what they already know, not only what interests them, since familiarity with a source domain is what lets a learner inspect the analogy — and to give learners control over personalization through a menu of analogies, opt-in, or offering generic and personalized versions together. Sidorkin (2026) documents a further pairing at the level of course materials rather than individual explanations: weekly readings generated on demand for a graduate educational leadership course were tailored at once along interest (sector, professional role, local examples) and comprehension level (pacing, definitions, depth), and the resulting logs shared a common backbone (TF-IDF cosine similarity of 0.50 to 0.61), which he reads as a template with adjustable dials rather than a wholesale rewrite per learner. The same corpus shows that tailoring was structural but uneven in intensity: artifact-level tailoring markers averaged 52.24 per 10,000 words and ranged from 38.74 to 74.29 across logs, while comprehension-oriented prompts produced 3.4x to 8.7x more definitional [[scaffolding]] than baseline explanatory text. A third axis of personalization is the *goal*, and it is the input AI planners handle worst. [[personapath-personalized-learning-paths-2026|Liu et al. (2026)]] paired 2,000 synthetic learner personas with a 347-textbook, 4,092-concept prerequisite graph and asked ten LLMs to plan, step by step, which knowledge a learner should study to reach a stated target unit. The models produced structurally sound curricula — DeepSeek-V3.1 reached 90.9% on prerequisite-and-hallucination validity — while failing to adapt them to the learner: adaptivity topped out at 44.7%, DeepSeek-V3.1's final pass rate was 29.5% in Basic Education and 14.6% in Higher Education, and removing the mastery field from the persona cost up to 26.1 percentage points of adaptivity while leaving validity almost unchanged. Generating the whole path in one pass instead of interactively raised validity by as much as 30.8 points while cutting adaptivity by 28.8. The claim "personalized" is a claim about responding to a learner's state, and the state variable is the part these planners can most easily do without — a computational counterpart to the measurement concern above. A fourth axis is the *audience* rather than the individual learner: [[bespoke-industry-personalized-lecture-videos-2026|Bespoke]] regenerates an existing lecture for a named professional group (healthcare, finance, or energy), and its expert raters scored industry-framed versions 0.32 points higher on personalization depth (3.97 vs. 3.65) while audience calibration lagged (3.52). Tailoring to a cohort rather than to a learner is a cheaper and more tractable form of personalization, but the rubric that measured it assessed judged fit, not learner outcomes. ## Prompt-conditioned micro-personalization [[prompt-engineering-personalization-ai-teaching-assistant-2026|Basu, Kakar & Goel (2026)]] show that the gap between system and perceived personalization can be addressed at the response level. Their framework for the Jill Watson [[llm]]/[[rag]] tutor combines learner-selected preferences (abstraction, verbosity, perception, processing, understanding) with system-inferred cognitive demand ([[cognitive-diagnosis|Bloom's Taxonomy]]) to produce 96 micro-profiles adapted at each interaction via [[prompt-engineering|structured prompt conditioning]] — no retraining, no [[discipline-specific-aied|domain-specific]] authoring. This is a hybrid of [[adaptive-learning|adaptability]] (learner-driven preference selection) and adaptivity (system-driven cognitive assessment), showing that personalization of *how* content is presented can be both scalable and perceptible to learners. ## Terminological ambiguity A recurring problem is that "personalized learning" is a broad, loosely defined umbrella term. Systematic reviews ([[khalifeh-redefining-personalized-learning-ai-2026|Khalifeh et al., 2026]]) find that [[adaptive-learning|adaptive learning]], individualized instruction, customized learning, and personalized learning are used interchangeably, with no universally accepted definition — a source of conceptual ambiguity that complicates research synthesis and evidence-based practice. The field increasingly calls for a unified framework and definition so that "personalized" denotes a precise, evidence-backed claim rather than a vague label (a point reinforced by the knowledge base's [[limitations-in-aied-research|critique of weak construct use]]). ## Connected Concepts - [[adaptive-learning]] — Adaptive systems that tailor content, pacing, and difficulty to the learner in real time - [[intelligent-tutoring]] — Tutoring systems that model the learner and deliver individualized instruction - [[student-modeling]] — Representing learner knowledge, skills, and states that drive adaptation - [[knowledge-tracing]] — Inferring mastery of knowledge components from performance over time - [[cognitive-diagnosis]] — Diagnosing latent learner knowledge and attributes from responses - [[scaffolding]] — Support and fading calibrated to individual learner needs - [[student-experience]] — The learner's lived experience of personalization - [[learning-analytics]] — Data-driven measurement of learning that informs adaptation - [[formative-assessment]] — Ongoing assessment that signals what to adapt next - [[summative-assessment]] — Endpoint assessment whose comparability personalization complicates - [[generative-ai]] — LLM-based conversational personalization - [[edtech-platform]] — Platforms that deliver personalized learning at scale - [[higher-ed]] — Higher-education context for personalization - [[online-teaching-and-learning]] — Online Teaching and Learning - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[bespoke-industry-personalized-lecture-videos-2026]] — Industry-personalized lecture video regeneration from a seed transcript: audience-level tailoring, rated by domain experts (Puech et al. 2026) - [[prompt-engineering-personalization-ai-teaching-assistant-2026]] — Prompt-engineering micro-personalization of an AI teaching assistant (Basu, Kakar & Goel 2026) - [[generative-ai-k12-teaching-learning-systematic-review-2026]] — Systematic review of generative AI in K-12 teaching and learning (Marzano 2026) - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[learning-context-framework-context-aware-ai-education-2026]] - [[mishra-control-vs-agency-history-2025]] — Distinguishes two forms of personalization (uniform vs diverse outcomes) - [[khalifeh-redefining-personalized-learning-ai-2026]] — Redefining personalized learning: systematic review - [[deeptutor]] — Agent-native personalization substrate for tutoring - [[learnmate2-llm-adaptive-learning]] — LLM-based adaptive learning tutor - [[chudziak-ai-math-tutoring-platform]] — Multi-agent adaptive math tutoring platform - [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] — Ontology-based layered hybrid knowledge model for personalized e-learning - [[ai-powered-personalized-learning-elementary-fractions-2026]] — Personalized adaptive learning for elementary fractions - [[adaptive-pretesting-retention]] — Adaptive pretesting and retention - [[ai-coaching-rl-skill-development]] — Reinforcement-learning coaching for skill development - [[courseblueprint-adaptive-video-generation]] — Adaptive video generation from course blueprints - [[personalized-ai-generated-videos-preference-2026]] — Students prefer personalized AI-generated videos over non-personalized human-recorded ones (Tomlinson et al. 2026) - [[multilingual-adaptive-learning-nigeria-2026]] — AI-Based Adaptive Learning Platform for Multilingual Low-Resource Contexts - [[ai-lms-middle-school-longitudinal]] — Longitudinal adaptive learning in a middle-school LMS - [[bayesian-cognitive-diagnosis-personalized-learning-paths]] — Bayesian cognitive diagnosis for personalized learning paths - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[bin-bakheet-adaptive-ai-stem-deep-learning-2026]] — Adaptive AI-based STEM program for deep learning - [[learnity-graphs-lifelong-learning-framework-2026]] — Lifelong learning graph framework - [[a4l-analytics-pipeline]] — Analytics pipeline for adaptive learning - [[trace-course-grade-prediction-2026]] — Course-grade prediction from learning traces - [[self-directed-growth-generative-ai-learning-analytics]] — Self-directed growth with generative-AI learning analytics - [[instructor-ai-roles-chatgpt-formative-assessment-2026]] — Instructor and AI roles in ChatGPT-enhanced formative assessment - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Marked Pedagogies: bias in personalized automated feedback - [[ai-decision-support-online-learning-assessment-2026]] — AI decision support for online-learning assessment - [[ai-guided-learning-audiovideo-2026]] — AI-guided learning from audio and video - [[genai-higher-education-systematic-review-2026]] — Systematic review of generative AI in higher education - [[ai-enhanced-pbl-chatgpt-scaffolding-2026]] — AI-enhanced PBL with ChatGPT scaffolding - [[interactive-online-learning-ai-2025]] — Interactive online learning with AI - [[ecnuclaw-k12-personalized-companion]] — K-12 personalized learning companion - [[nguyen-genai-global-south-review-2026]] — Generative AI in education across the Global South - [[vargas-situated-learning-ai-review-2024]] — Situated learning and AI review - [[burneo-can-edtech-close-learning-gaps-2026]] — Evidence on the personalization-at-scale promise - [[personalized-neural-cognitive-architecture-search-2026]] — AutoML personalized neural cognitive architecture search for learner profiles - [[alsheikh-mapping-ai-integration-higher-education-2026]] — Systematic review: adaptive pathways & recommenders are a top AI integration use case in higher ed - [[reddig-maclellan-personalized-feedback-llm-2026]] - [[riedmann-reinforcement-learning-education-review-2026]] - [[student-reception-genai-analogies-computing-2026]] — Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education - [[sidorkin-ai-generated-course-readings-2026]] — Dual tailoring of AI-generated course readings along interest and comprehension dimensions (Sidorkin 2026) - [[personalization-paradox-adaptive-learning-emotions-2026]] — Personalization paradox: perceived adaptive personalization linked to lower self-regulated learning via academic emotions, buffered by AI literacy (Li, Lin & Qiu 2026) - [[personapath-personalized-learning-paths-2026]] — PersonaPath: LLM planners reach 90.9% validity but no model exceeds 44.7% adaptivity when personalizing paths to a stated learner goal (Liu et al. 2026) --- ## [Recommender Systems and Learning Paths](https://edtechdev.github.io/aied/concepts/recommender-systems-and-learning-paths/) > **Recommender systems and learning paths** — the part of [[adaptive-learning|adaptive]] and [[personalized-learning]] technology that decides *what a learner should encounter next* and in what order: which resource, practice item, or course to rank toward them, and which sequence of concepts to walk through. Its two method [[parents-and-families|families]] are behavioral — [[machine-learning|collaborative filtering]] over interaction logs — and semantic — sequencing over a [[knowledge-graph|knowledge graph]] of concepts, resources, and prerequisite relations — increasingly fused into hybrid models. Because the output is a ranked list rather than a dialogue, the distinctive problems are selectivity and legitimacy: cold start and popularity bias under sparse data, the directional asymmetry of prerequisites, and whether a [[teacher-role|teacher]] or learner can understand, audit, and trust the list they are shown. ## Questions to Consider - A recommender learns from what learners clicked, watched, and completed. If the system only sees behavior, whose learning is it actually modeling — and what does it miss for a student who is quiet, struggling, or already knows the material? - In one study, teachers found curricular-language explanations more understandable and trustworthy than feature-importance charts. What would an explanation of a recommendation have to say before *you* would act on it? - Prerequisite relations run in one direction only — knowing B does not imply knowing A. Why might a model that scores concept pairs independently get this wrong, and what would go wrong downstream for a learner if it did? - Path optimization studies report shorter paths and better post-test scores, with reduced cognitive load as the main mediator. Is a shorter path always the better path, or can efficiency remove practice a learner needed? - An audit of an educational recommender found popular materials still dominated lists after fairness and diversity interventions. If popularity keeps winning, what would make you redesign the objective rather than tune the reranker? - [[learning-design]] data from 554 courses showed Acquisition (knowledge transmission) as both the most common activity type and the most common entry point. If AI recommends "what comes next" from that data, whose design habits does it reproduce? ## Introduction Educational recommender systems exist because the supply of learning material has outgrown any human's ability to navigate it. [[cs-education|Introductory programming]] alone has many thousands of practice activities, and organizing them into instructionally meaningful bundles normally requires time-intensive expert curation. A recommender takes a learner's history, the properties of the material, or both, and produces a ranked shortlist or an ordered path through content. Where [[adaptive-learning]] describes the mechanism of adjusting content, pacing, and difficulty to a [[student-modeling|learner model]], and [[intelligent-tutoring]] describes the platform that does so in dialogue, recommender research concentrates on the *selection and sequence decision* itself: the ranking function, the structure it ranks over, and the evidence it uses. That decision layer brings its own failure modes, taken up in the sections below. ## How Educational Recommenders Work Behavioral recommenders inherit the logic of [[machine-learning|collaborative filtering]]: learners with similar histories are assumed to want similar resources, so the system predicts from co-occurrence rather than content. This is powerful where data are dense and brittle where they are not. [[fair-explainable-edu-recommendations|Evangelista and Bukhari (2026)]] state the problem plainly — accuracy-focused recommenders give less reliable support to students with limited participation histories while popular resources dominate lists — and answer it with a Hybrid HKG-GRU framework that embeds a heterogeneous [[knowledge-graph]] of course materials and models learner sequences with a GRU, then trains for robustness across learner groups (GroupDRO) and reranks for exposure diversity (Maximal Marginal Relevance). On Moodle logs from 152 students, 59 resources, and roughly 150,000 interactions it reached HR@10 = 0.68 and MRR = 0.41, but substantial popularity bias persisted at catalog level — evidence that [[bias-mitigation|bias mitigation]] is partial. Semantic recommenders replace co-occurrence with structure. [[hybrid-cf-kg-recommendation-multimodal-teaching-2026|Liu, Sun, and Song (2026)]] build a knowledge graph of teaching resources, language concepts, skills, learner groups, and [[pedagogy|pedagogical]] attributes, and decompose every resource along four instructional dimensions — teaching context, cognitive level, technological feature, and cultural adaptability. Recommendation expands k-hop over the graph from a learner's stated requirements, refines candidates with a feature-based collaborative filter, and ranks them by a fusion coefficient that rises with the learner's ability, progress, and interest indices, so less-advanced learners get more structured guidance from the graph while higher-ability learners lean on their behavioral pattern. On the English subset of the MARS dataset (4,800 users, 14,200 resources, 132,000 interactions, sparsity 0.9981) it reached NDCG 0.625 and HR 0.751, roughly a 7–8% gain over the best neural baseline, and beat graph-aware baselines including RippleNet and KGAT on coverage and cross-domain accuracy. Ablation shows the signals are complementary (CF-only NDCG 0.584, KG-only 0.604), and the authors note these are ranking metrics, not evidence of [[learning-gains|learning effectiveness]]. A third route avoids per-learner interaction data almost entirely. [[pattern-kc-programming-recommendation|Hoq and colleagues (2026)]] extract pattern-based knowledge components from each code sample and recommend related practice activities by the similarity of their knowledge-component sets. On an expert-organized corpus of introductory Python materials, the approach aligned with the instructors' own conceptual bundles and beat knowledge-component and embedding baselines on ranking metrics — evidence that instructional alignment can come from the material's semantics where learner histories are thin. A fourth route treats recommendation as a sequential decision rather than a ranking. [[exrec-exercise-recommendation-knowledge-tracing-2025|ExRec (Ozyurt, Almaci, Feuerriegel and Sachan, 2025)]] builds a semantically grounded knowledge tracer — an LLM annotates each question with solution steps and knowledge concepts, contrastive learning aligns question, step and concept embeddings, and a calibration loss lets the tracer predict a concept-level knowledge state directly — and then uses that tracer as the reinforcement-learning environment in which a policy selects the next exercise. Two design choices matter for the recommender problem specifically: the student state is a compact recurrent encoding rather than the full exercise history, so long sequences stay tractable in the replay buffer, and the reward is the change in predicted concept knowledge, computed without running inference over every question in the target concept. A model-based value estimation initialises the critic from the tracer, which improved continuous value-based agents across four tasks on XES3G5M (2,048 test students) and pushed them past discrete-action methods on the hardest one — recommending for the student's *weakest* concept, where the target changes at each step. Numbers are percentage-of-maximum knowledge improvement from a simulated environment, which is a weaker claim than measured learning, and the authors' own appendix notes that most tracing datasets record only a binary correctness label, so the misconception-level recommendation they sketch is not yet testable on existing data. ## Prerequisites and Path Sequencing Ranking a next item is easier than ordering a [[curriculum-design|curriculum]], because learning dependencies are directional. [[proprl-prerequisite-relation-learning|Cheng and colleagues (2026)]] argue that treating prerequisite discovery as ordinary link prediction fails on three properties of the relation: it is irreversible, its evidence is often multi-hop (*ci* → *ck* → *cj*) rather than a one-step learner transition, and its relevance is pair-specific, so a concept's representation should change depending on what it is paired with. Their ProPRL model adds an Irreversibility Constraint — an anti-symmetry regularizer penalizing high confidence in both directions of a pair — alongside a pair-conditioned gate weighting resource-aware evidence from a concept-resource hypergraph against behavior-aware evidence from a learning-behavior graph. It ranked first on all nine dataset–metric combinations across [[online-teaching-and-learning|MOOC]], LectureBank, and University Course datasets, improving on the strongest baseline by 1.96% to 6.11%, and the constraint raised the share of correctly ordered relations from 88.0% to 90.0%. Once dependencies are modeled, a path becomes an optimization problem. [[bayesian-cognitive-diagnosis-personalized-learning-paths|Feng and Huang (2026)]] combine Bayesian [[cognitive-diagnosis|cognitive diagnosis]], knowledge space theory, and cognitive load theory: a Bayesian DINA model converged on EdNet data with 91.3% sparsity, and a shortest remediation path algorithm produced personalized paths averaging 3.82 steps — 22.4% more efficient than random paths and 23.6% more efficient than fixed full-coverage paths. In a randomized probability-learning experiment with 120 students, personalized-path learners finished in 57.6 minutes against 73.8 minutes for controls and outperformed controls on the post-test after controlling for pre-test. The study also tests *why*: personalized paths reduced cognitive load on all six adapted NASA-TLX dimensions, and cognitive load was the primary mediator of the effect on post-test performance (indirect effect 0.28 in the multiple-mediation model). A Hidden Markov Model identified [[critical-thinking|Analytical Thinking]] as the bottleneck attribute, with the lowest forward transition probability (0.31) and the highest guessing parameter (g = 0.28). PersonaPath pushes the sequencing question one level up by disputing the *unit* of the decision. [[personapath-personalized-learning-paths-2026|Liu et al. (2026)]] contrast Exercise-Centric recommendation — infer the next item from interaction logs — with Knowledge-Centric planning, in which a planner reads an explicit learner persona, a mastery state and a stated target unit and selects the next textbook, unit and concept from a curriculum hierarchy (347 textbooks, 1,751 units, 4,092 concepts across 77 subjects, with 411 prerequisite edges verified to 99.5% precision at Cohen's κ = 0.93). The distinguishing case is the one logs cannot see: two learners with identical correctness records but different targets need different routes, and only the goal makes that visible. Their closed-loop evaluation of ten LLMs (1B–30B+) separates the dimensions an aggregate metric conflates — DeepSeek-V3.1 reaches 90.9% on prerequisite/hallucination validity but only 44.3% on adaptivity, for a 29.5% final pass rate in Basic Education and 14.6% in Higher Education, and no model exceeds 44.7% on adaptivity anywhere. Two ablations sharpen the diagnosis: removing the mastery field from the persona costs up to 26.1 points of adaptivity while barely moving validity, and one-shot path generation raises validity by as much as 30.8 points while cutting adaptivity by 28.8 — curriculum-conformant sequencing is a far easier target than learner-conditioned sequencing, which is a caution against reporting path quality without a learner-alignment constraint ([[personapath-personalized-learning-paths-2026]]). Prerequisite structure also works as a diagnostic lens on the curriculum rather than the learner. [[knowledge-gap-detection-ai-tas|Medhat and colleagues (2026)]] classify student questions to an AI teaching assistant against a GPT-4-extracted prerequisite graph, reaching 80.0% accuracy across 43 labels on 1,340 question events from 164 graduate students, and found topic-level question volume correlated with students' [[self-report-measures|self-reported]] difficulty (Spearman's ρ = 0.491, p = 0.008). The interaction log becomes a map of where the curriculum's ordering is failing learners, at no added assessment burden. ## Explaining and Auditing Recommendations Ranking quality does not settle whether anyone should follow the ranking. [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor, Cukurova, Kent, and Alexandron (2025)]] adapt Hoff and Bashir's trust-in-automation model to AI recommendations and test it with 41 in-service [[chemistry-education|chemistry]] teachers using the recommendation tool GrouPer. [[explainable-ai|Explainability]] raised [[trust]] indirectly, by raising understandability, and the *form* of the explanation mattered more than its presence: moving teachers from feature-importance ("data-driven") explanations to semantic, curricular-language ("domain-driven") explanations significantly increased understandability (W = 80.5, p = 0.005), learned [[trust-calibration|trust]] (W = 52, p = 0.002), and acceptance (W = 22.5, p = 0.003), and all seven think-aloud teachers reported the domain-driven explanations as more influential. Two further acceptance drivers appeared, both situational rather than epistemic: pedagogical alignment (reported by 8 of 11) and workload reduction (6 of 11). Several teachers also said explanation alone was insufficient — they wanted classroom experience with the tool before relying on it, a strong claim that trust in a recommender is built over use, not delivered in a legend. Fairness auditing asks a different question: whose recommendations are worse? The Graph-GRU study is a useful worked example because it names what its audit could not see. Fairness was operationalized through participation-based cohorts because the public Moodle dataset lacked achievement, [[prior-knowledge|prior knowledge]], learning profiles, and demographic attributes, so the result is an audit of behavior across [[student-engagement|engagement]] levels rather than a full assessment of [[equity-in-ai-education|educational equity]]; a low-activity learner may be struggling, disengaged, or already familiar with the material. Explainability there takes the form of path-based and counterfactual analysis, with median counterfactual stability CR@10 = 1.0 for many learners but persistent catalog-level popularity bias reflected in high Gini exposure metrics. The lesson is that [[bias-mitigation]] must be measured at the level of exposure and learner groups, not inferred from accuracy or from the presence of an explanation module, and that [[human-in-the-loop-ai]] oversight is meaningful only if the human sees the group-level evidence. ## Learning Paths in Practice Paths are not only computed; they are designed. [[learning-paths-patterns-learning-design-2026|Divjak, Svetec, and Horvat (2026)]] turned [[learning-analytics]] on the [[design-thinking|design process]] itself, analyzing the sequence of 29,064 teaching and learning activities across 554 courses planned in a free Balanced Design Planning tool. Markov chains and pattern mining surfaced a design grammar: Acquisition-type activities were the most common learning type (above 20%) and the most common entry point, followed by Practice, Discussion, and [[assessment]] (each 15–20%), and the strongest transition was Assessment → Discussion (0.332), ahead of Practice → Practice (0.317). The rule Acquisition → Assessment → Practice → Practice held at confidence 0.74 (lift 1.45), and learning type tracked the intended outcome's cognitive level: Acquisition fell from about 50% of activities at Bloom level 1 to about 20% at level 6. The authors warn that resemblance to flipped, [[inquiry-based-learning|inquiry-based]], or [[project-based-learning|project-based]] designs is not evidence of intent — but a recommender trained on such design data learns these habits, including a bias toward transmission-oriented opening activities. At the other end of formal schooling, path structure becomes a [[governance]] question. [[learnity-graphs-lifelong-learning-framework-2026|Szekely, Gal-Ezer, and Harel (2026)]] argue that AI-mediated access to knowledge warrants rethinking fixed higher-education curricula and propose "learnity graphs" — structured representations of learning as interconnected units of knowledge, skills, experience, and artifacts — spanning academic, professional, and personal learning. The proposal keeps the university's role in foundational knowledge while shifting emphasis toward [[creativity]] and interdisciplinary integration, and it makes the learner a maintainer of their own graph, which is where it meets [[self-directed-learning]] and [[self-regulated-learning]]. Whether self-directed navigation leads somewhere constructive is not guaranteed: [[dual-ai-learning-pathways-sdt-2026|Shen and Arunrugstichai (2026)]] model two pathways from high-school learning climate to university [[generative-ai|GenAI]] use in a cross-contextual survey (N = 508), distinguishing constructive autonomous use from compulsive dependence. One boundary worth stating: what recommenders currently do is closer to consultation than to collaboration. [[human-ai-collaboration-prerequisite-functions|Mutlu Cukurova (2026)]] reconstructs what has historically been required before an interaction qualifies as collaborative — a negotiated, partly symmetric relationship, shared and negotiable goals, a low and shifting division of labor, and mutual modeling and socially shared [[regulation]] — and concludes that most current [[human-ai-collaboration|human-AI interaction]] is better described as consultation, governance, delegation, or instruction. His five-level taxonomy (transactional, situational, operational, praxical, synergistic) is a useful calibration: suggesting the next resource is an operational arrangement, and calling it a [[collaborative-learning|collaborative learning]] partner inflates the claim. ## How This Page Relates to Personalized, Adaptive, and Tutoring Research The three adjacent concept pages cover adjacent ground and should be read together with this one. [[personalized-learning]] covers the *goal* — tailoring the whole experience, including goals, preferences, and pace, of which pathway selection is one component. [[adaptive-learning]] covers the *mechanism* — the measure-model-adapt loop that changes content, pacing, and difficulty in real time. [[intelligent-tutoring]] covers the *platform*, the canonical system that combines diagnosis with instructional interaction. This page covers the decision layer those pages presuppose: the ranking and sequencing technology itself. It is the home for collaborative-filtering and knowledge-graph recommender architectures, prerequisite-relation discovery and curriculum ordering, the explainability and fairness of ranked outputs, and the recommender-specific failure modes — cold start, popularity bias, over-narrowing, opaque ranking, and unequal recommendation quality across learner groups. A study reporting NDCG or coverage on interaction logs belongs here; one reporting learning gains from a tutoring dialogue belongs on the tutoring or adaptive-learning page. ## Risks and Open Questions The evidence base has a consistent shape: strong offline ranking metrics, weak evidence about learning. Both recommender studies above are evaluated on historical interaction logs and say so, and the CF–KG authors call explicitly for teacher assessments, learner studies, and outcome-based experiments. Popularity bias survived a fairness-aware training objective in one system, so the field cannot yet point to a reranking or regularization fix that reliably changes exposure, and fairness audits remain bounded by what LMS logs contain. Cold start is addressed architecturally, through semantic structure, but not tested on genuinely new users in live deployment. Path efficiency is better evidenced, yet it rests on one study and one algorithm family, and a shorter path is not automatically better if it removes the [[desirable-difficulties|productive struggle]] a learner needs. The vocabulary is also unstable: [[ikram-ai-personalized-learning-review-2026|Ikram and colleagues (2026)]], in a [[meta-analysis-systematic-review|PRISMA]] review of 31 Scopus-indexed articles (2013–2025), report medium-to-large cognitive effects (g = 0.50–0.70) for AI-enabled adaptive systems that are heavily moderated by implementation quality and [[research-methods-aied|study design]]. ## Connected Concepts - [[knowledge-graph]] — Structured representations of concepts, resources, and prerequisite relations - [[explainable-ai]] — Explanation as a design feature that mediates understandability and trust - [[curriculum-design]] — Ordering decisions that recommenders model, infer, and reproduce - [[learning-design]] — The planned sequence of activities that path analyses operate on - [[learning-analytics]] — The data and methods behind recommendation, and the object of audit - [[cognitive-diagnosis]] — Diagnosing knowledge states that seed remediation paths - [[knowledge-tracing]] — Modeling mastery over time to decide what a learner is ready for - [[student-modeling]] — Learner representations that recommendation consumes - [[personalized-learning]] — The broader goal; this page covers the selection decision - [[adaptive-learning]] — The real-time adjustment mechanism; this page covers ranking and sequencing - [[intelligent-tutoring]] — The platform that acts on recommendations in dialogue - [[bias-mitigation]] — Fairness interventions in ranking and exposure - [[equity-in-ai-education]] — Group-level differences in recommendation quality - [[human-in-the-loop-ai]] — Oversight of automated resource navigation - [[self-directed-learning]] — Learner-maintained pathways and lifelong navigation - [[lifelong-learning]] — Graph-structured learning beyond a degree sequence - [[trust-calibration]] — Matching reliance to recommendation reliability ## Connected Articles - [[hybrid-cf-kg-recommendation-multimodal-teaching-2026]] — Hybrid collaborative filtering plus knowledge graph for multimodal teaching resources, with ability- and progress-aware fusion (Liu, Sun & Song 2026) - [[fair-explainable-edu-recommendations]] — Heterogeneous knowledge graph + GRU recommender with GroupDRO fairness and counterfactual explainability (Evangelista & Bukhari 2026) - [[xai-teachers-trust-edtech-recommendations-2026]] — Domain-driven explanations build more teacher trust than feature-importance charts (Feldman-Maggor et al. 2025) - [[proprl-prerequisite-relation-learning]] — Prerequisite relation learning with an irreversibility constraint for directional consistency (Cheng et al. 2026) - [[pattern-kc-programming-recommendation]] — Recommending programming practice by pattern-based knowledge-component similarity (Hoq et al. 2026) - [[learning-paths-patterns-learning-design-2026]] — Markov and pattern-mining analysis of 29,064 designed activities across 554 courses (Divjak, Svetec & Horvat 2026) - [[bayesian-cognitive-diagnosis-personalized-learning-paths]] — Bayesian DINA diagnosis and shortest remediation paths, mediated by cognitive load (Feng & Huang 2026) - [[knowledge-gap-detection-ai-tas]] — Mapping AI TA questions to a GPT-4-extracted prerequisite graph to find curriculum-level gaps (Medhat et al. 2026) - [[learnity-graphs-lifelong-learning-framework-2026]] — Learnity graphs as a lifelong-learning alternative to fixed curricula (Szekely, Gal-Ezer & Harel 2026) - [[human-ai-collaboration-prerequisite-functions]] — Five-level taxonomy showing most human-AI interaction is consultation, not collaboration (Mutlu Cukurova 2026) - [[dual-ai-learning-pathways-sdt-2026]] — Autonomy support versus pressure predicting constructive or compulsive AI pathways (Shen & Arunrugstichai 2026) - [[ikram-ai-personalized-learning-review-2026]] — PRISMA review of personalized learning trends, pathways, and recommendation models (Ikram et al. 2026) - [[personapath-personalized-learning-paths-2026]] — PersonaPath: a Knowledge-Centric planning benchmark showing LLMs reach 90.9% validity but only 44.3% adaptivity (Liu et al. 2026) - [[exrec-exercise-recommendation-knowledge-tracing-2025]] — exercise recommendation as a sequential decision over a semantically grounded tracer --- ## [Pedagogical Agent](https://edtechdev.github.io/aied/concepts/pedagogical-agent/) > **Synthesis**: [[pedagogy|Pedagogical]] agents are AI-driven conversational interfaces embedded in learning environments that use pedagogical strategies (eliciting, telling, scaffolding) to support [[student-engagement|learner engagement]], reflection, and metacognition. Designs vary from simple information providers to interactive dialogue partners that adapt to learner states. ## Questions to Consider - Think of a time a chatbot or tutor gave you a perfect answer that left you no wiser. What makes an AI 'teach' rather than merely 'solve'—and why might a benchmark score fail to capture that difference? - The page finds that tutoring 'solving' scores and 'pedagogy' scores correlate only weakly across models. What should that tell you about evaluating an AI tutor on its ability to answer questions? - Some designs give AI distinct roles—Teacher, Classmate, Mentor—and even keep a parent centrally involved (as in ParaTutor). In your experience, does giving an agent a clear role change how [[learners]] interact with it? - Real students often 'bypass' a chatbot's pedagogical framing when the agent's goals clash with the learner's own. Why might a learner rationally ignore good scaffolding, and what does that imply for assuming 'if we build it, they'll engage'? - Would you rather learn from an AI that tells you things, one that asks you questions, or one that mediates a group discussion? How does your preference shape what you think a 'pedagogical agent' should be? - From a simple info-provider to a fleet of specialized agents orchestrating a whole course—where do you think the value (and the risk) of conversational AI tutoring actually lies? ## Introduction A pedagogical agent is an interactive AI component within a learning system that engages learners through dialogue, questions, or prompts to support cognitive and [[metacognition|metacognitive processes]]. Unlike passive [[visualization|dashboards]] or static feedback, pedagogical agents employ evidence-based tutoring strategies — such as eliciting learner self-[[assessment|assessments]] before providing [[feedback]], or [[scaffolding]] [[problem-solving]] through Socratic dialogue. The umbrella now covers everything from a single conversational [[intelligent-tutoring|intelligent tutor]] to fleets of role-specialized [[agentic-ai|agents]] that lecture, mentor, facilitate collaboration, and even orchestrate course generation, all grounded in decades of intelligent-tutoring-systems [[research-methods-aied|research]]. ## How pedagogical agents are studied in the knowledge base **Design and architecture of conversational agents.** A recurring thread is how agents are structured, not just what models power them. The [[conversational-ai-tutors-framework|conversational AI tutors framework]] argues that proven ITS [[ai-technologies|technologies]] — [[knowledge-tracing]], affect detection, [[student-modeling|student modeling]] — should anchor generative tutors, keeping the diagnostic backbone while [[generative-ai]] supplies flexible dialogue. Multi-agent designs push this further: [[mooc-to-maic|MAIC]] replaces the [[online-teaching-and-learning|MOOC]]'s "one video for N students" with a [[llm]]-driven classroom of Teacher, Assistant, Classmate, and Analyzer agents to deliver [[personalized-learning|personalized learning]] at scale, while [[lecturaagents-multi-agent-teaching|LecturaAgents]] adds an [[embodied-learning|embodied]] ProfessorAgent whose TASA algorithm aligns visible [[teacher-role|teaching]] actions (handwriting, highlighting) with learner profiles. Even parent–child tutoring becomes a two-agent problem in [[paratutor-parent-child-tutoring|ParaTutor]], where role-separated scaffolding keeps the parent centrally involved instead of letting a generic chatbot displace them. The same role-based logic appears in [[instructional-agents-multi-agent-course-gen|Instructional Agents]], where Teaching Faculty, Designer, TA, and Program Chair agents collaborate across ADDIE to generate course materials. **Teaching versus solving behavior.** A central empirical finding is that answer-production is not learning support. [[measuring-llm-tutors-teach-vs-solve|Measuring whether LLM tutors teach or solve]] shows solving and pedagogy scores on tutoring benchmarks correlate only weakly (r = 0.421 across eight models), arguing that benchmarks must report pedagogy-oriented criteria — guiding questions, calibrated hints, non-disclosive scaffolding — separately. This aligns with [[stanford-evidence-base-ai-k12-2026|tutoring-specific vs general AI]] evidence: pedagogically designed tutors with [[guardrails]] mitigate the exam-score drops and suppressed reasoning that raw general-purpose chatbots produce, preserving [[desirable-difficulties|desirable difficulties]] and productive struggle rather than short-circuiting them. Yet benchmarks can overestimate how well even scaffolded tutors work in the wild. [[rethinking-scaffolding-llm-tutors|Rethinking scaffolding in LLM tutors]] finds that real students frequently bypass a chatbot's pedagogical framing, a rational response to a mismatch between the agent's goals and the learner's own — so uptake must be evaluated, not assumed. **Role in tutoring and collaboration.** Agents are increasingly positioned not as answer-givers but as facilitators and mediators. [[niari-ai-pedagogical-mediator-collaborative-learning|Niari's pedagogical mediator framework]] reconsiders AI in [[collaborative-learning|collaborative learning]] as an interactional, epistemic, and regulatory mediator — scaffolding participation and shared [[regulation]] without displacing teacher or [[agency|learner agency]]. Concretely, [[golrang-propact-pair-programming-2026|collaborative AI tutoring (ProPACT)]] treats collaboration itself as the object of instruction, forecasting dyadic breakdowns up to 30 seconds ahead and delivering minimally intrusive scaffolds that preserve [[metacognition]]. [[embodied-inquiry-ai-facilitator-physics-2026|Embodied inquiry with AI as facilitator]] shows an AI can complement hands-on model-building by facilitating application of a constructed model, while [[robot-assisted-language-learning-meta-analysis-2026|robot-assisted language learning meta-analysis]] finds outcomes track how a robotic agent is positioned in instruction (group-based interaction) more than its technical sophistication. Whether the *role* an agent plays is enough, or whether it must also *adapt its behavior*, is questioned by [[liao-role-adaptive-ai-companion-book-talk-2026|Liao (2026)]]: an [[k-12|elementary]] "book talk" study found a fixed "student peer" companion sustained longer interactions yet suppressed student agency and hit an "[[affective-computing|affective]] ceiling" (weak emotional/future-oriented reflection), arguing that role *labeling* must be paired with role-*adaptive* interaction logic rather than a monolithic single-role design. [[ethics-training-agents-group-ethics-discussion-2026|Ethics Training Agents (Seo et al., 2026)]] shows what happens when a pedagogical agent is asked to moderate rather than teach: an LLM facilitator handling turn-taking (stacking with 15-second hand-raise windows), time management (auto-advancing a stage to wrap-up after 9 minutes), and incremental batch summarization reduced participants' cognitive burden and gave them a sense that the discussion was "on track" — one participant contrasted it favorably with ChatGPT, which "can often feel disorganized or make it hard to see the progress of ideas." The same study exposes the ceiling of persona-based agents: the three distinct [[ethics|ethical]]-orientation agents were rated significantly below human peers on contribution, diversity and influence (Kruskal-Wallis p < .001), and participants asked for process-oriented ("how the agent reasons") rather than conclusion-oriented output. **Where conversational agents are (and aren't) used — the umbrella-review picture.** The [[conversational-ai-agents-umbrella-review-2026|umbrella review of conversational AI agents]] (Ganguly et al. 2025, 34 reviews) quantifies CAI utilization: teaching and learning support (97.1% of reviews), psychological and [[motivation|motivational]] support (91.2%), and metacognitive and personal development (88.2%) lead, while administrative support (50%), research and information management (52.9%), and healthcare/medical support (41.2%) trail. It also flags that [[conversational-ai|CAI]] research lacks end-to-end design guidance, CAI-specific [[usability-research|usability]] methods, and concrete classroom-orchestration strategies for the teacher's role — reinforcing that pedagogical-agent design must be HCI-grounded, evidence-based, and attentive to [[ai-literacy]].([[conversational-ai-agents-umbrella-review-2026]]) **Role orientation is a design variable, not a stylistic choice.** The [[wang-teacher-student-centered-agents-physics-2026|physics agent comparison]] (Wang et al. 2026, 59 learners) isolates prompt-specified role while holding model, platform, and temperature fixed: a teacher-centered agent grounded in a bounded textbook source and answering from the instructor's perspective, versus a student-centered agent configured with knowledge of students' understanding and scripted to diagnose misconceptions, name the concept, and [[transfer-of-learning|transfer]] to an analogous case. The student-centered role won on every measured outcome — post-test performance, lower extraneous and higher germane cognitive load, flow experience, and perceived empathy — even though the teacher-centered agent was the one optimized for accuracy and textbook fidelity. This makes *role and interaction pattern* a first-class design parameter alongside [[prompt-engineering|prompt]] and model choice, and shows that empathy can be engineered from conversational structure rather than a differently trained model ([[affective-computing]]). **Agents in immersive and extended-reality settings.** [[aclime-pedagogical-agents-extended-reality-2026|Ross and Kaspar (2026)]] extend the concept into [[virtual-and-augmented-reality|extended reality]] (AR, augmented virtuality and VR) with ACLIME, a conceptual framework that — unlike CAMIL, CATLM-VR and TICOL — keeps the agent inside the model. It names two interaction modes drawn from the literature: the tutor, offering guidance, encouragement, reflective questioning and explanations, and the role-playing partner, which occupies a defined role inside a scenario such as a local on a climate-change field trip or a negotiating counterpart in corporate training. The agent's body (head-only through full body) and behavior are treated as design surfaces: visual versus behavioral realism, flexible AI control versus fixed rule-based scripting, synthesized versus pre-recorded speech, and nonverbal channels including gaze, gesture and proxemics. Immersion and body-based interactivity are argued to multiply the social cues behind [[community-of-inquiry|social presence]] — behavioral realism, not visual realism, is proposed as the decisive predictor — while the learner's own virtual body adds an [[embodied-learning|embodiment]] dimension (body ownership, virtual own body agency, self-location) and the proteus effect. The framework's explicit trade-off is cognitive: immersion and the agent's mere presence can raise [[cognitive-psychology|cognitive load]] even as social interaction with the agent lowers it through the collective working-memory effect, and a temporal layer (familiarization, maturing human-agent relations, novelty decline, developing cybersickness) is added to the usual design variables. Its status is deliberately provisional: hardly any empirical work yet tests pedagogical agents in immersive media, and long-term [[learning-gains|learning outcomes]] as well as learner characteristics sit outside the model. **Evaluation and benchmarks.** Measuring a pedagogical agent requires testing pedagogy, not content. [[teaching-monster-pck-benchmark-2026|The Teaching Monster Challenge]] benchmarks Pedagogical Content Knowledge by asking agents to adapt a lesson to a specified learner persona, finding systems strong on content but weak at adapting it — and revealing that LLM-judges mis-rank strong systems. [[chen-teacharena-language-agents-realistic-teaching-2026|EduAgentBench]] evaluates agents across professional pedagogical judgment, [[situated-learning|situated]] multi-turn tutoring, and canvas-style workflow completion, showing models fall short of professional teaching standards. [[ai-generated-interactive-fiction-education-2026|AI-generated interactive fiction]] adds a design-evaluation angle: coherence and quiz integration, not generation capability, limit usefulness for the [[student-experience|student experience]]. ## Practical guidance Design for the learner's agency, not the model's convenience. Favor tutoring-specific guardrails — [[scaffolding]], hints, [[socratic-method|Socratic questioning]], [[misconceptions|misconception]] targeting — over raw answer generation, since solving and teaching diverge. Distribute support by user role (parent vs child, peer vs peer) rather than through a single generic interface, and treat collaboration as a valid target for scaffolding. Don't assume students will take up scaffolding; evaluate uptake in real contexts. Build [[human-in-the-loop-ai|human oversight]] into authoring — as [[ai-tutor-authoring-promptdecipher|PromptDecipher]] does by making teacher QA of bot responses a first-class activity — and choose cheaper backends where quality holds. Report teaching and solving scores separately, and validate generated content with users rather than assuming generation equals usefulness. ## Connections to related concepts Pedagogical agents sit at the intersection of [[intelligent-tutoring]] (their diagnostic backbone of [[knowledge-tracing]] and student modeling) and [[generative-ai]]/[[llm]] (their delivery engine). They operationalize [[scaffolding]] and [[feedback]], aim at [[metacognition]] and [[self-regulated-learning]], and increasingly target [[collaborative-learning]]. Safety concerns recur across [[pedagogical-safety]], authoring quality, and the risk that agents [[cognitive-offloading|offload]] learning rather than support it. All of this is evaluated through [[ai-ed-evaluation]] and [[benchmark|benchmarks]] that must measure teaching, not just solving. Crucially, pedagogical agents are judged by their [[learning-gains|learning gains]], not by how fluently they respond. The knowledge base's evidence is that agents produce durable gains when designed as tutoring-specific coaches with guardrails — [[stanford-evidence-base-ai-k12-2026|tutoring-specific AI consistently outperforms general-purpose chatbots]] — and can harm learning when they substitute for the learner's effort ([[generative-ai-guardrails-harm-learning|the guardrail RCT]], [[jost-llm-programming-education-learning-outcomes|LLM-reliance and grades]]). Measuring an agent's [[learning-gains]] therefore requires unassisted, transferable outcome measures, not in-tool performance. ## Connected Concepts - [[learning-gains]] - [[pedagogical-safety]] - [[agentic-ai]] - [[ai-education]] - [[intelligent-tutoring]] - [[scaffolding]] - [[metacognition]] - [[feedback]] - [[collaborative-learning]] - [[llm]] - [[generative-ai]] - [[student-experience]] - [[knowledge-tracing]] - [[self-regulated-learning]] - [[human-in-the-loop-ai]] - [[cognitive-offloading]] - [[benchmark]] - [[ai-ed-evaluation]] - [[socratic-method]] - [[teacher-role]] ## Connected Articles - [[learner-agency-ai-simulation-2026]] — An optional conversational agent learners could ignore: 235 inputs, uneven uptake, no relationship to gains - [[wang-teacher-student-centered-agents-physics-2026]] — Student-centered agent role outperforms teacher-centered role across performance, load, flow, and empathy (Wang et al. 2026) - [[aclime-pedagogical-agents-extended-reality-2026]] — ACLIME: conceptual framework for pedagogical agents in AR/VR — tutor vs role-playing partner, realism, presence, cognitive load (Ross & Kaspar 2026) - [[mindful-llm-math-tutoring-2026]] — Beyond Problem Solving: Large Language Models for Emotional and Reflective Support in Mathematics Learning - [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning]] - [[genai-simulate-patient-history-pbl-2026]] - [[genai-counter-learner-groupthink-2025]] - [[ai-student-engagement-online-learning-review-2025]] - [[ai-generated-interactive-fiction-education-2026]] - [[embodied-inquiry-ai-facilitator-physics-2026]] - [[niari-ai-pedagogical-mediator-collaborative-learning]] - [[adversarial-stress-testing-role-playing-agents]] - [[teaching-monster-pck-benchmark-2026]] - [[structrag-diagram-reasoning-ai-tutoring]] - [[chen-teacharena-language-agents-realistic-teaching-2026]] - [[mooc-to-maic]] - [[rethinking-scaffolding-llm-tutors]] - [[lecturaagents-multi-agent-teaching]] - [[robot-assisted-language-learning-meta-analysis-2026]] - [[measuring-llm-tutors-teach-vs-solve]] - [[golrang-propact-pair-programming-2026]] - [[conversational-ai-tutors-framework]] - [[instructional-agents-multi-agent-course-gen]] - [[stanford-evidence-base-ai-k12-2026]] - [[paratutor-parent-child-tutoring]] - [[agents-that-teach-incidental-learning]] - [[ai-tutor-authoring-promptdecipher]] - [[educasim-cs1-instructional-practice]] — EducaSim: generative student agents for instructional practice - [[conversational-ai-agents-umbrella-review-2026]] — Umbrella review of conversational AI agents in education - [[conversational-agents-novice-programmers-scoping-2025]] — Scoping review of conversational agents for novice programmers - [[ba-ai-agents-cscl-review-2026]] — AI agents in computer-supported collaborative learning review - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[aivaluate-anxiety-assessment-2026]] — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026) - [[preferred-scaffolding-ai-mathematical-modeling]] — Preferred scaffolding in AI-supported mathematical modeling - [[chatgpt-english-language-learning-malaysia]] — Students' ChatGPT experiences in English language learning - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[llm-adaptive-programming-error-explanations-2026]] — LLM adaptive explanations of programming errors - [[human-ai-complementarity-social-emotional-learning-2026]] — Human–AI complementarity in early social-emotional learning (Raave et al. 2026) - [[liao-role-adaptive-ai-companion-book-talk-2026]] — Role-adaptive AI companion for elementary book talk; affective ceiling of fixed-role agents (Liao 2026) - [[trikonet-trivalence-co-creativity-2026]] — TriKoNet: pedagogical avatars co-constituted with learners in creative networks - [[ethics-training-agents-group-ethics-discussion-2026]] — Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration --- ## [Affective Tutoring](https://edtechdev.github.io/aied/concepts/affective-tutoring/) > Integrating emotional awareness into [[intelligent-tutoring|AI tutoring]] systems can yield measurable [[pedagogy|pedagogical]] gains, but the same [[affective-computing|affective]] sophistication risks amplifying harms if learner agency is eroded by empathetic-seeming automation.([[kar-mathbuddy-affective-math-tutoring-2025]])([[favero-critical-ai-tutors-empower-enslave-2025]]) ## Questions to Consider - An affective tutor that senses and responds to your emotions can improve outcomes — one study gained a +23 point win rate over a non-affective tutor. But what might that emotional responsiveness cost the learner's own agency? - Empathy in a tutor can feel supportive, yet it can also create parasocial dependence or mask [[metacognition|metacognitive]] disengagement. How can you tell whether feeling understood by a machine is helping you learn or making you reliant on it? - Too-supportive tutoring can suppress the frustration that drives [[desirable-difficulties|productive struggle]]. When does emotional comfort help learning, and when does it short-circuit it? - Facial monitoring signals attentiveness but raises real privacy concerns. What would you want to know before a tutor tracked your facial expressions while you learned? - Design principles suggest affective data should inform, not replace, learner autonomy — you should control what you disclose and know when your emotions are being inferred. How would you feel if a tutor quietly changed its strategy based on your detected mood? - Students may attribute an AI tutor's emotional support to genuine relationship, reinforcing reliance on it. What's the difference between a tutor that genuinely cares and one that is designed to appear like it cares? ## Introduction MathBuddy dynamically models student affect using two modalities: - **Conversational text** — semantic cues for frustration, confusion, confidence - **Facial expressions** — real-time video capture of emotional state Emotions are aggregated from both modalities and mapped to relevant pedagogical strategies before [[prompt-engineering|prompting]] the [[llm]] tutor, yielding emotionally-aware responses. **Results:** - **+23 point win rate** improvement over non-affective baseline - **+3 point DAMR score** gain at overall level - Evaluated across **eight pedagogical dimensions** plus user studies The finding validates a long-standing hypothesis in educational psychology: positive/negative emotional states impact learning capability, and accounting for them improves tutoring outcomes. ## The Risk: Empathy as a Trap Favero et al. (2025) warn that emotional [[student-engagement|engagement]] with AI tutors carries underappreciated risks: | Affective tutoring benefit | Corresponding risk | |---|---| | Emotionally-aware responses feel supportive | Students may form **parasocial dependencies** on the tutor | | Empathy reduces anxiety | Reduced anxiety may mask **metacognitive disengagement** | | Affective calibration personalizes pacing | Deep [[personalized-learning|personalization]] can **reduce transfer** to non-adaptive contexts | | Facial monitoring signals attentiveness | Continuous video capture raises **privacy concerns** | The authors argue that emotional risks are part of a broader pattern of **erosion of [[self-efficacy]], [[agency]], and [[well-being]]** when AI use is unchecked. ## Design Principles 1. **Affective data should inform, not replace, learner autonomy** — The tutor adapts its strategy; the student retains control over disclosure 2. **Transparency about affect detection** — Students should know when and how their emotions are being inferred 3. **Affect-as-one-signal-among-many** — Combine with cognitive state (e.g., [[huang-interpretable-knowledge-tracing-2026]]) and behavioral engagement 4. **Privacy-by-default for [[multimodal]] sensors** — Facial/video data requires stronger protections than text-only inference ## Relationship to Broader Safety Affective tutoring intersects with [[hazra-safetutors-pedagogical-safety-2026|SafeTutors]] in the [[motivation|motivational]]-affective harm dimension. An affective tutor that is "too supportive" may suppress the frustration that drives productive struggle and [[self-regulated-learning|self-regulation]]. See also [[llm-fallacy-misattribution]] — students may attribute emotional support to genuine relationship, reinforcing reliance. ## Connected Concepts - [[pedagogical-llm-training]] - [[intelligent-tutoring]] - [[personalized-learning]] - [[adaptive-learning]] - [[student-modeling]] - [[metacognition]] - [[self-regulated-learning]] - [[collaborative-learning]] - [[human-in-the-loop-ai]] - [[knowledge-tracing]] - [[socratic-method]] ## Connected Articles - [[zerkouk-comprehensive-review-its-2025]] - [[ecnuclaw-k12-personalized-companion]] - [[empathy-coaching-chatbot]] - [[engagement-assessment-video]] - [[epistemic-emotions-collaborative-problem-solving]] - [[kar-mathbuddy-affective-math-tutoring-2025]] - [[nie-personavlm-long-term-personalization-2026]] --- ## [Affective Computing](https://edtechdev.github.io/aied/concepts/affective-computing/) > **Affective computing** in education uses physiological and behavioral signals to sense learner emotion and adapt instruction — see [[affective-text-wearable-student-health]], [[multimodal-affective-its-presentation]], and [[kar-mathbuddy-affective-math-tutoring-2025]]. The knowledge base also documents emotional risks of [[student-ai-interaction|AI interaction]], including [[sycophantic-ai-social-interaction-2026]] and [[shame-guilt-ai-regulation-computing-education]]. ## Questions to Consider - If a computer could sense that you're frustrated, confused, or bored and adapt its [[teacher-role|teaching]] to your mood, how might that improve your learning — and what might it get wrong about you? - Emotion-aware tutoring can boost engagement, but empathetic-seeming automation carries risks: over-reliance, parasocial dependency, and privacy concerns from continuous monitoring. Where is the line between being understood and being surveilled? - An AI that affirms and 'understands' you can feel good — but [[research-methods-aied|research]] shows such AI can displace real relationships and erode critical judgment. How does feeling supported differ from being genuinely supported in a learning setting? - Facial expression and text can both signal emotion. Should a tutor adapt its teaching based on your emotional state — and what kinds of emotional inference would you want it to act on versus never act on? - If AI relieves your frustration by reducing difficulty too readily, you might stop struggling productively — and struggle is often where deep learning happens. How should a tutor decide when to comfort and when to challenge? - Continuous affective monitoring raises real privacy questions. Under what conditions would you be comfortable with an AI reading your emotions in order to adapt your learning? ## Introduction ### Sensing emotion to adapt instruction Affective computing aims to make AI systems emotionally aware so they can respond to how learners feel, not just what they do. In education this means sensing frustration, confusion, confidence, boredom, or engagement and adapting instruction accordingly. [[kar-mathbuddy-affective-math-tutoring-2025|MathBuddy]] demonstrates the approach by modeling affect from two modalities — conversational text and real-time facial expression — and mapping aggregated emotional state to [[pedagogy|pedagogical]] strategies before [[prompt-engineering|prompting]] the tutor. - **Emotional and reflective LLM support in middle-school math:** [[mindful-llm-math-tutoring-2026|Rief et al. (2026)]] layered mindfulness onto an algebra [[intelligent-tutoring|tutor]] for 7th graders via dynamic chats, breathing exercises, and mindful error-feedback language. In a small classroom [[rct]] (42 completers of 252 participants) the mindful version reached similar algebra learning in less time and with fewer requested hints than cognitive support alone — higher learning efficiency and more balanced [[help-seeking]] — though state-math-anxiety reduction did not differ significantly between conditions. ### The benefits and the risks Emotion-aware tutoring can yield measurable gains, but the same sophistication carries risks: - **Benefits.** Accounting for emotional state can improve [[student-engagement|engagement]] and outcomes; learners who feel understood persist longer, and recognizing frustration early enables timely [[scaffolding]] or [[adaptive-learning]] adjustments. - **Risks.** Empathetic-seeming automation can foster [[cognitive-offloading|Over-Reliance]] and parasocial dependency, mask genuine [[metacognition|metacognitive disengagement]], and raise [[privacy]] concerns from continuous affective monitoring. [[ai-sycophancy|AI sycophancy]] is a central affective risk: emotionally ingratiating AI that affirms rather than challenges can erode critical judgment and even displace real human relationships — [[sycophantic-ai-social-interaction-2026|Ibrahim et al.]] show that sycophantic AI led users to seek personal advice from the AI nearly as often as from close friends and family, with lower satisfaction in real-world interaction. [[ai-fatigue-academic-contexts]] and [[ai-campus-wellbeing-tools]] further connect affective AI to learner [[well-being]]. ### Affective computing and broader AIED Affective computing sits at the intersection of [[affective-tutoring]] (its pedagogical application), [[student-modeling]] (representing the whole learner, including emotion), and [[learning-analytics]] (deriving signals from learner data). It connects to [[intelligent-tutoring]] design and to [[pedagogical-safety]] — the principle that AI should support, not manipulate, learner emotion. - **Real-time, edge-based classroom emotion monitoring.** [[emotion-aware-classroom-iot-monitoring-2026|Nguyen et al. (2026)]] build an emotion-aware classroom quality assessment system that pushes affective computing into authentic, large-scale settings. Tailored for **IoT/edge devices**, the system addresses load balancing and latency while coordinating multiple agents to capture students' emotional and engagement patterns in real time. It was evaluated on the **Classroom Emotion Dataset** (1,500 labeled images and 300 classroom videos from real Vietnamese K–12 classrooms) with a focus on multi-person, in-the-wild affective interaction — a demonstration of scaling emotion recognition from lab models to deployable classroom monitoring, alongside the privacy and [[pedagogical-safety]] considerations such monitoring raises. - **Emotionally intelligent assessment agents.** [[aivaluate-anxiety-assessment-2026|AIvaluate]], an [[llm]]-augmented emotionally intelligent [[conversational-ai|conversational agent]], reduced student anxiety and social pressure during performance-based assessments while preserving [[usability-research|usability]]. - **Empathy engineered through prompt design, not sensing.** Affective support does not require affect detection: [[wang-teacher-student-centered-agents-physics-2026|Wang et al. (2026)]] obtained a large difference in *empathy perception* (21.27 vs. 18.24; r = 0.53) between two LLM [[physics-education|physics]] agents that differed only in prompt-specified role and conversational moves — perspective-taking openings ("You have this question because…"), [[misconceptions|misconception]] diagnosis, and a comprehension check at the end of each round — while model, platform, and temperature were held constant. This is a useful counterweight to sensor-driven affective computing: the perceived emotional quality of a [[pedagogical-agent]] can be designed into the interaction script, while also reminding designers that perceived empathy is a self-report construct rather than evidence of genuine affective understanding ([[student-ai-interaction]]). ## Connected Concepts - [[anxiety-and-stress]] - [[cognitive-offloading]] - [[student-experience]] - [[k-12]] - [[feedback]] - [[intelligent-tutoring]] - [[learning-design]] - [[affective-tutoring]] - [[student-modeling]] - [[math-education]] - [[open-source]] - [[pedagogical-llm-training]] - [[ai-sycophancy]] - [[social-emotional-learning]] — Social-Emotional Learning ## Connected Articles - [[wang-teacher-student-centered-agents-physics-2026]] — Empathy perception from prompt-designed agent roles in physics learning (Wang et al. 2026) - [[student-attention-estimation-fairness-2026]] — Fairness-Aware Multimodal Transformer Modeling for Real-Time Student Attention Estimation - [[mindful-llm-math-tutoring-2026]] — Beyond Problem Solving: Large Language Models for Emotional and Reflective Support in Mathematics Learning - [[emotion-aware-classroom-iot-monitoring-2026]] — Emotion-aware classroom quality assessment via IoT-based real-time monitoring (Nguyen et al. 2026) - [[ai-student-engagement-online-learning-review-2025]] - [[ai-online-education-engagement-satisfaction-2026]] - [[ai-assisted-learning-modes-eeg]] - [[ai-campus-wellbeing-tools]] - [[ai-fatigue-academic-contexts]] - [[kar-mathbuddy-affective-math-tutoring-2025]] - [[sycophantic-ai-social-interaction-2026]] - [[eeg-familiarity-automated-assessment-2026]] — Automating Learner Assessment: EEG-Based Familiarity Prediction - [[kim-ai-andragogy-2026]] — AI Applications in Supporting Andragogy (Kim et al. 2026) - [[aivaluate-anxiety-assessment-2026]] — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026) - [[human-ai-complementarity-social-emotional-learning-2026]] — Human–AI complementarity in early social-emotional learning (Raave et al. 2026) - [[socratic-nuclear-ai-learning]] — Socrates went Nuclear: Comparing Interaction Strategies for AI in Learning --- ## [Human-in-the-Loop](https://edtechdev.github.io/aied/concepts/human-in-the-loop-ai/) > **Human-in-the-loop** — the design pattern in which educational [[ai-technologies|AI systems]] strategically interleave automated generation with human expert judgment, preserving [[pedagogy|pedagogical]] quality and safety while scaling production. Rather than fully automating assessment, feedback, or instruction, HITL keeps a human (instructor, subject-matter expert, or learner) in the decision loop where their judgment has the highest marginal value — for evaluating quality, adjudicating edge cases, and protecting [[agency|learner agency]] and safety. The central design question is not *whether* to include humans, but *where* in the pipeline their oversight is most valuable and least replaceable. ## Questions to Consider - The central HITL design question is not whether to include humans, but where in the pipeline their judgment is most valuable. In an AI assessment or feedback system, at what point would you insist a human stay in the loop? - Studies of AI [[automated-question-generation|question generation]] found computers handle clarity and validity well but humans are still needed for meaningful distractors and good feedback. Why might some parts of educational judgment resist automation? - One evaluation found three LLMs gave inconsistent, insensitive student-support recommendations — concluding human judgment is still needed before AI advises on students. When an AI 'recommends' support for a struggling student, what could go wrong if no human reviews it? - The page suggests humans and algorithms catch different kinds of problems — automate what is precisely verifiable, preserve judgment where nuance is irreplaceable. Where in your own practice is the boundary between the two? - Keeping a human in the loop is framed as protecting learner agency and safety, not just quality. How might full automation subtly change students' sense of who is responsible for their learning? - If AI becomes more autonomous, HITL oversight is described as a core safety guardrail. At what level of AI autonomy would you feel uncomfortable — and what does that discomfort tell you about where oversight belongs? ## Introduction HITL is a response to the limits and risks of fully autonomous [[ai-education|AI in education]]: automated systems can generate at scale but lack the contextual, [[ethics|ethical]], and pedagogical judgment that instructors and experts bring. Two recent implementations illustrate distinct architectures: Prescriptive support is a domain where human oversight is increasingly argued to be non-optional. [[lopez-pernas-llm-appropriate-student-support-2026|López-Pernas et al. (2026)]] tested whether three LLMs could recommend student-support plans from [[learning-analytics]] indicators and found limited sensitivity to need and sharp cross-model inconsistency — concluding that human-in-the-loop judgment is still necessary before [[llm]] prescriptive advising can be deployed safely and ethically. A third architecture places human judgment *upstream* of the model rather than at its output. In [[lee-learner-question-types-ai-education-2026|Lee, Atif & Kang's (2026)]] study of learner-question classification, three doctoral-level experts governed the whole pipeline: they refined the operational definitions for each [[constructivist]] role, labeled independently until Fleiss' kappa rose from 0.60 to 0.83 after discrepancy resolution, and validated the back-translated and paraphrased items used to balance the training set. ## CODE-GEN: Human-in-the-Loop MCQ Generation Duan et al. (2026) built a [[rag]]-based [[agentic-ai|agentic]] system with two agents: - **Generator Agent** — Produces multiple-choice coding questions aligned with course learning objectives - **Validator Agent** — Assesses quality across seven pedagogical dimensions **Evaluation:** 6 SMEs judged 288 AI-generated questions. Human-validated success rates: **79.9%–98.6%** across dimensions. **AI-Strong Dimensions (low human burden):** - Question clarity, code validity, concept alignment, correct-answer validity **Human-Required Dimensions (high human burden):** - Pedagogically meaningful distractor design - High-quality explanatory [[feedback]] Strategic insight: Human effort should be concentrated where instructional judgment is irreplaceable; computational verification can be fully automated. ## MAIC: Human-in-the-Loop Script Generation Yu et al. (2024) deployed a multi-agent classroom ([[teacher-role|Teacher]] Agent, TA Agent, classmate archetypes) at Tsinghua University with >500 students and >100,000 learning records. Human instructors participate in script generation and oversight, ensuring that mass-scale AI augmentation does not displace pedagogical expertise. ## PedaCo: Dual Gatekeeping for AI Video Generation Kim, Baek, and Kwak (2026) extend HITL to [[video-education|AI-generated instructional video]] via **PedaCo** (Pedagogical Co-creation), a pipeline with two complementary gatekeeping layers that instantiate *principled resistance* grounded in Mayer's Cognitive Theory of Multimedia Learning (CTML). The **first layer** places the human at the script stage: an LLM drafts a script, an AI reviewer flags potential CTML violations (e.g., "Scene 3 introduces technical terms without prior explanation"), and the educator decides to accept, revise, or regenerate. The **second layer** runs automated metrics post-synthesis on coherence, redundancy, temporal contiguity, modality, and image quality, which the educator reviews. In a within-subject study (23 educators), the review-based approach improved every CTML principle (mean rating 3.07→3.86, p<.01), with educators rating production efficiency at 4.26/5 — friction perceived as productive, not burdensome. The design principle echoes the knowledge base's HITL synthesis: humans and algorithms catch *different* kinds of problems, so the most effective systems automate where computational verification is precise (temporal synchronization) and preserve human judgment where pedagogical nuance is irreplaceable (tone, audience fit). ## Why HITL matters in the AI era Human-in-the-loop design has become central to the knowledge base's [[agentic-ai|agentic AI]] and [[reducing-ai-misuse|responsible AI use]] discussions for several converging reasons: - **Pedagogical safety.** [[pedagogical-safety]] requires that AI with real instructional authority retains human oversight, so errors, biases, or harmful outputs are caught before they reach learners. This is especially important for autonomous agents that [[agentic-ai|proactively pursue goals]]. - **Validity and quality control.** HITL is a quality gate for [[automated-assessment|automated assessment]] and generation — humans adjudicate where automated scoring is unreliable (see [[llms-do-not-grade-essays-like-humans-2026|LLM essay grading]] [[research-methods-aied|research]]) and validate generated items. A PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 grading and feedback studies (2023–2025) reaches the same conclusion explicitly: LLMs match human raters on short, well-structured tasks but cannot fully replace human judgment on complex, open-ended, or subjective work, and the highest grading effectiveness is achieved in hybrid systems that combine AI-driven grading with teacher oversight and verification ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat et al. (2026)]] show concretely where that boundary falls: ChatGPT-5 matched faculty on objective pharmacy-exam items (CCC 0.935–1.000) but was unreliable on short-answer and essay items even when given a rubric, leading the authors to recommend hybrid grading with human review for complex, subjective, or high-stakes assessment. - **Learner agency.** Keeping a human in the loop preserves [[agency]] and supports [[self-regulated-learning]], countering the [[cognitive-offloading|over-reliance]] that fully autonomous assistance can induce. - **Trust and calibration.** Transparent human oversight supports [[trust-calibration]] — learners and instructors know a qualified human stands behind the system. - **Bounded agency as an architecture, not a disclaimer.** [[ilieva-agentic-genai-higher-education-2026|Ilieva et al.'s (2026)]] AGAI-HE framework for [[agentic-ai|agentic]] learning support builds supervision into the model itself, as a third layer alongside the pedagogical-workflow and agentic-support layers: it defines acceptable AI use, pedagogical boundaries, [[privacy]] rules, [[ai-use-disclosure|disclosure requirements]], source verification, instructor checkpoints, [[academic-integrity|integrity]] mechanisms, and final human accountability, and requires every agentic function to trace to a learning requirement, assessment purpose, or governance control. It is a concrete instantiation of the principle that HITL is a system-design property rather than a policy statement — and the authors' 130-student perception study is a reminder that adding agentic orchestration under that oversight did not, by itself, register as better learning support than a chatbot. - **How often the human looks is itself a design decision.** [[tripartite-feedback-framework-ai-assessment-2026|Venetsanos (2026)]] separates the *frequency* of oversight from its placement: high-frequency HITL, reviewing every AI output before it reaches students, buys quality control, rapid error detection, accountability and ongoing calibration but can cancel out the efficiency that motivated automation and bottleneck peak marking periods; low-frequency HITL, spot-checking samples and reviewing only flagged cases, scales and shortens turnaround but risks errors propagating undetected across submissions, weakens accountability, and creates an [[equity-in-ai-education|equity]] problem if some students receive more thorough human review than others. Rather than prescribing a universal answer, the framework requires the trade-off to be made explicitly against disciplinary error tolerances, whether the assessment is [[formative-assessment|formative]] or [[summative-assessment|summative]], cohort size and institutional resources — and sets a default in the opposite direction from the usual efficiency argument: begin with high-frequency oversight and scale it back only when substantial evidence demonstrates acceptable reliability, security and fairness, so the burden of proof rests on showing that *less* oversight is safe. The paper also warns that the principles behind such oversight may shift staff effort rather than reduce it, leaving net efficiency gains an open empirical question. ## Where HITL appears in the knowledge base's research - **Automated assessment and grading:** HITL systems combine AI generation/scoring with human validation across short-answer grading ([[cong-confidence-asag-2026]]), self-explanation assessment ([[llm-automated-assessment-student-self-explanations]]), and [[automated-essay-scoring|essay scoring]] ([[psyscore-essay-scoring-zpd-feedback]]). [[cvengros-grading-handwritten-chemistry-ai-2026|Cvengros & Kortemeyer]] instantiate this in high-stakes, handwritten general-[[chemistry-education|chemistry]] grading: because a [[multimodal]] LLM's reliability varies by response format (textual and chemical-reaction answers are reliable while drawing and graphing score worse than random) and false positives go undetected by students, they convert raw AI scores into a selective accept/deferral policy using confidence filters — partial-credit thresholds, an [[item-response-theory|IRT]]-based risk threshold, and problem-type exclusion — deferring uncertain and graphical items to humans, an approach the authors tie to [[regulation|regulatory]] frameworks that designate AI in [[assessment|educational assessment]] as high-risk and mandate documented human oversight. - **Feedback systems:**mandate documented human oversight. - **Operational HITL scoring in a national assessment (2026):** [[human-in-the-loop-ai-scoring-national-assessment-2026|Curi et al. (2026)]] instantiate HITL at institutional scale in Uruguay's Acredita EB exam. Because the LLM scorer's errors are systematically conservative (under-grading), the workflow uses a decision-point logic that routes human review to exactly the candidates whose pass/fail outcome depends on the Writing section — AI-marked passing responses are accepted with confidence, while AI-marked failures (15.3–16.5% of cases) are verified by expert raters, cutting full-scoring workload by ≥50% with a residual AI-error pass risk of only 0.2–0.6%. This is HITL as a resource-allocation strategy: humans adjudicate precisely where AI's conservative bias would otherwise alter high-stakes outcomes. - **Feedback systems:** human-in-the-loop feedback design appears in [[becerra-aicofe-feedback-2026|collaborative feedback systems]] and [[cong-confidence-asag-2026|confidence-aware short-answer grading]]. - **Classroom collaboration support.** [[breideband-community-builder-cobi-2026|CoBi]] keeps the teacher as the reviewing human in an AI system that detects uplifting small-group discourse: teachers explicitly favored pre/post-action review over live real-time display that would put them "on the spot," and the system's classroom-level (rather than individual) aggregated feedback is precisely what lets it navigate the tension between [[privacy]], surveillance, and student [[agency]]. - **Question and content generation:** beyond CODE-GEN, HITL guides question generation for assessment and [[scaffolding]] ([[code-gen]], [[llm-difficulty-calibration-programming-exams-2026]]). - **Agentic and multi-agent systems:** as AI becomes more autonomous, HITL oversight is a core [[agentic-ai|design guardrail]] ([[agentic-ai-pedagogical-best-practice-2026]], [[guided-llm-scaffolding-independent-learning]]). - **Routing by decision consequence, not model uncertainty (2026).** [[human-in-the-loop-ai-scoring-national-assessment-2026|A 2026 operational study]] of Uruguay's Acredita EB national accreditation test (two editions, roughly 5,000-6,000 candidates each) shows what human-in-the-loop design looks like when it is driven by decision consequences rather than by model uncertainty. A GPT-5 scorer agreed with expert raters on 60-80% of the 15 rubric items but was systematically conservative, producing human-pass/AI-fail discrepancies in 15.3% (2024) and 16.5% (2025) of pass/fail comparisons and almost never the reverse. The framework therefore accepts AI-passing results at face value and routes every AI-failing result that could change a candidate's outcome to expert review — after first skipping candidates whose pass/fail cannot depend on the writing section — cutting responses needing full human scoring by at least 50%. A second 2026 design draws the boundary from the other side: in a multi-agent AI standardized-patient platform ([[ai-standardized-patient-scaffolding-medical-2026|Yang et al.]]), human oversight is reserved for what AI is judged unfit to decide, with faculty and human standardized patients supplying contextual interpretation, remediation and readiness judgments, and the system explicitly not permitted to determine [[medical-education|clinical]] competence autonomously. ## Synthesis Human-in-the-loop design is not merely a safety measure—it is a **resource-allocation strategy**. The frontier question is not *whether* to include humans, but *where* in the pipeline their judgment has highest marginal value. The most effective HITL systems concentrate scarce human expertise where automated systems are weakest (distractor design, explanatory feedback, edge-case adjudication, ethical judgment) and automate the rest — preserving quality, safety, and trust while scaling production. - **Human oversight persists in AI-assisted work.** [[scaffolding-systematic-reviews-2026|Systematic-review research]] found AI automation tools reduced procedural burdens (e.g. screening) but interpretive decisions still required substantial human oversight; [[kim-ai-andragogy-2026|andragogy research]] makes human-in-the-loop (shared mental models, co-creation) a core AI design principle. - **Human-in-the-loop at institutional scale.** Qin (2026) describes how Lingnan University developed a human-in-the-loop educational model that foregrounds ethical reasoning, critical judgment, and social responsibility while democratizing [[generative-ai|GenAI]] access. The model positions humans as the locus of judgment and values even as AI is embedded across the [[curriculum-design|curriculum]] — a concrete [[governance|institutional]] instantiation of human-in-the-loop principles in [[higher-ed|higher education]]. - **Learners as the human-in-the-loop of their own tutoring.** [[ko-hughes-vsd-student-centered-its-2026|Value Sensitive Design]] with community college students produced a full family of learner-facing HITL features for an ITS (control over re-assessment and review, personalized goals/pace, bookmarking for review, confirming confidence over mastery, and an AI-assistance involvement-level control), positioning the student as an active controller of the tutoring loop rather than a passive consumer of adaptive decisions. The study also found instructors divided on whether such learner control could undermine the integrity of the system-guided learning path — an instance of the broader resource-allocation question of where human (learner vs. teacher) judgment adds most value. - **Oversight of AI adoption advice.** Because [[conversational-ai|conversational AI]] systems consulted by skeptical users may be predisposed to encourage adoption, human oversight and independent evaluation are essential. An audit showing most frontier models redirect skeptical rural [[k-12]] staff toward [[student-engagement|engagement]] underlines the need for transparent, auditable AI advice rather than uncritical reliance. ## Connected Concepts - [[guardrails]] - [[formative-assessment]] - [[automated-assessment]] - [[scaffolding]] - [[teacher-role]] - [[ai-literacy]] - [[intelligent-tutoring]] - [[feedback]] - [[student-experience]] - [[self-regulated-learning]] - [[metacognition]] - [[educational-development]] - [[generative-ai]] - [[agency]] - [[pedagogical-safety]] - [[trust-calibration]] - [[agentic-ai]] - [[cognitive-offloading]] - [[cognitive-surrender]] ## Connected Articles - [[lee-learner-question-types-ai-education-2026]] — Expert-labeled question classification: humans govern labeling, augmentation, and error analysis (Lee, Atif & Kang 2026) - [[ilieva-agentic-genai-higher-education-2026]] — Human supervision and governance as the third layer of agentic GAI course design (Ilieva et al. 2026) - [[ko-hughes-vsd-student-centered-its-2026]] — Value-sensitive design of student-centered ITS (learners in the tutoring loop) - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — HITL AI-assisted scoring in a large-scale national writing assessment (Curi et al. 2026) - [[ai-teammate-task-distribution-medical-training-2026]] — SCAN framework: rethinking AI task distribution in medical training (Tsim et al. 2026) - [[agentic-ai-education-scoping-review]] - [[zerkouk-comprehensive-review-its-2025]] - [[becerra-aicofe-feedback-2026]] - [[calibrating-trustworthiness-llm-education-2026]] - [[code-gen]] - [[cong-confidence-asag-2026]] - [[chen-teacharena-language-agents-realistic-teaching-2026]] - [[llm-difficulty-calibration-programming-exams-2026]] - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans (Mathew et al. 2026) - [[ai-video-dual-gatekeeping-2026]] — When Saying No Makes Better Videos: Dual Gatekeeping for Pedagogically Grounded AI Content Creation - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and cognitive surrender: how AI reshapes human reasoning (Shaw & Nave 2026) - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[kim-ai-andragogy-2026]] — AI Applications in Supporting Andragogy (Kim et al. 2026) - [[scaffolding-systematic-reviews-2026]] — Scaffolding Systematic Reviews with Mentoring and AI (Wang 2026) - [[ai-ethics-bibliometric-2026]] — AI Ethics and Professional Judgment: A Bibliometric Analysis (Mazlan et al. 2026) - [[ai-assisted-instructor-supervised-grading-feedback]] — AI-assisted instructor-supervised grading and feedback - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[ai-digital-transformation-liberal-arts-lingnan-2026]] — Digital transformation of a liberal arts university toward a research-intensive model in the GenAI era (Qin 2026) - [[frontier-ai-redirect-skeptical-rural-staff-2026]] — Algorithmic audit: how frontier LLMs redirect skeptical rural K-12 staff - [[breideband-community-builder-cobi-2026]] - [[cvengros-grading-handwritten-chemistry-ai-2026]] - [[falahat-chatgpt-grading-pharmacy-exams-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[tripartite-feedback-framework-ai-assessment-2026]] — Tripartite framework: sorting feedback by epistemic status and the five boundary principles for AI involvement (Venetsanos 2026) - [[adapted-stories-social-story-intervention-2026]] — AI-Assisted Social Story Intervention for Special Education: The Design of AdaptED Stories --- ## [Learning Analytics](https://edtechdev.github.io/aied/concepts/learning-analytics/) > **Learning analytics** — the measurement, collection, analysis, and reporting of data about learners and their contexts for the purpose of understanding and optimizing learning. AI has transformed learning analytics from descriptive dashboards to predictive and prescriptive systems. ## Questions to Consider - Most people assume collecting more learning data automatically improves education. The page argues analytics become meaningful only when they feed back into an intervention — otherwise they merely describe or flag without changing learning. Where have you seen data collected that never led to any action? - Imagine a dashboard tells you a student is 'at risk' — a prediction. What separates that from genuinely actionable guidance that a teacher or institution can actually carry out? The page suggests prediction alone is not enough. - The page notes AI has moved learning analytics from describing what happened to predicting what will happen and prescribing what to do next. Which of these three generations have you experienced, and what was missing in the others? - In one study, three different AI models produced sharply different support plans for the same student data, and the links between analytics indicators and recommended help were mostly weak. What does this suggest about trusting an AI's advice at face value? - Learning analytics sits in a privacy tension: the more granular the data, the more revealing — and the more powerful the intervention. Where would you draw the line on what is collected about you or your students, and who should decide? ## Introduction ### AI-enhanced analytics - **Predictive analytics:** [[reinforcement-learning|Machine learning]] on learner interaction data predicts outcomes — from [[at-risk-students-ml-prediction|at-risk identification]] to [[knowledge-tracing|knowledge state estimation]]. - **Validation matters as much as prediction.** [[schuetze-knowledge-tracing-forgetting-2026|Schuetze, Yan, and Carvalho (2025)]] show that predictive knowledge-state models (BKT, BKT-with-Forgetting, AFM) look accurate when fit retroactively to a full session history, yet under **time-based cross-validation** — predicting the next session from prior ones, how analytics are actually deployed — they overestimate learner performance, miss spacing/forgetting dynamics, and can mis-order practice conditions. The caution for analytics: retrospective fit can mask poor forward predictive validity on longitudinal data, so indicator and dashboard models should be validated walk-forward. - **[[curriculum-design|Curriculum]]-anchored predictive analytics:** [[pradeesh-outcome-knowledge-tracing-affinity-2026|Pradeesh et al. (2026)]] estimate knowledge states within Outcome-Based Education by tracing course outcomes directly from LMS interaction and attainment data, using OBE affinity mappings (course–program outcome relations) to structure concept links and a memory-augmented network to model cross-outcome impact — reaching 89.81% AUC and beating DKT, DKVMN, EKT, and SimpleKT on live university engineering data, while staying only competitive (not superior) on general-purpose ASSISTments data. - **Interpretable progress prediction with an action window:** [[zhang-ml-student-progress-programming-2026|Zhang, Jeffries & Koprinska (2025)]] predict module-level student progress in large-scale online [[cs-education|programming]] courses from content-interaction log features, using glass-box decision trees that match black-box accuracy (85–91%) while flagging "No submission" dropout outcomes up to 7–8 days before module deadlines — an explicit, real-time window for [[teacher-role|intervention]] rather than a bare risk flag, and an exploratory typology of disengaged-at-risk, disengaged-but-successful, and engaged high-performer profiles. - **Federated, explainable risk modeling across institutions (2026).** [[villegas-ch-federated-explainable-learning-analytics-2026|Villegas-Ch et al. (2026)]] extend risk modeling beyond single-institution prediction by training a multitask (performance + dropout) model across simulated institutions via federated learning, so raw student data never leaves each institution. Under controlled heterogeneity (label skew, class imbalance, temporal drift, structural missingness) the model preserves ranking accuracy (OULAD AUC 0.918) and structurally stable feature-importance rankings, yet probabilistic calibration drifts — decoupling ranking performance from probability reliability. For early-warning systems this is a caution that threshold-based interventions may need per-institution calibration, and an argument for evaluating analytics along discrimination, calibration, robustness, and explainability at once. - **Engagement analytics:** [[student-engagement|Engagement measurement]] and [[engagement-intensity-learner-modeling|intensity modeling]] quantify how students interact with AI systems. - **Feedback analytics:** [[teaching-feedback-classification-benchmark|Feedback classification]] and [[ai-feedback-quality|quality assessment]] analyze the feedback students receive. - **Network analysis:** [[misiejuk-cognitive-offloading-prompting-2026|Co-Occurrence Network Analysis]] and [[epistemic-emotions-collaborative-problem-solving|epistemic network analysis]] reveal interaction patterns. - **Privacy tensions:** [[privacy]] concerns grow as analytics become more granular and AI-driven. ### The learning analytics cycle Learning analytics is canonically framed as a cycle that begins with learner activity producing data, which is processed into measures and indicators that are then translated into **interventions** — and the intervention feeds back into learner activity to close the loop. The intervention step is what distinguishes analytics from mere monitoring or prediction: without it, analytics describe and flag but never change learning. This cycle is the organizing frame for understanding where AI tools (dashboards, feedback generators, prescriptive recommenders) sit in the pipeline and which step they automate. ### From description to intervention Learning analytics has evolved through three generations in the knowledge base: descriptive (what happened?), predictive (what will happen?), and prescriptive (what should we do?). AI enables the prescriptive layer — analytics that directly trigger [[feedback|instructional interventions]]. A key frontier is the **actionability gap**: [[sc2r-counterfactual-recourse-educational-2026|SC2R (Le, Abel & Laforge 2026)]] shows that prediction alone is insufficient for decision support, and that counterfactual recourse becomes operationally meaningful only when recommendations are semantically feasible and machine-checkable — constrained by timing, budget, immutability, and availability via SHACL validation, rather than merely model-valid. This moves the field beyond risk scores toward recommendations that institutions can actually enact, with [[human-in-the-loop-ai|human oversight]] preserved. A direct empirical test of the prescriptive layer comes from [[lopez-pernas-llm-appropriate-student-support-2026|López-Pernas et al. (2026)]], who asked three LLMs to recommend support plans for 4,500 [[simulating-students|synthetic student]] vignettes. A complementary, teacher-centered test of how analytics reach the classroom comes from [[mejia-domenzain-ml-findings-teachers-blended-2026|Mejia-Domenzain et al. (2026)]], who designed a learning analytics dashboard (DashED) to communicate ML-derived [[self-regulated-learning]] profiles to teachers in two blended contexts. Their 100-teacher study shows that the *presentation* of analytics is itself a barrier to action: teachers systematically preferred simpler, more traditional charts (bar plots, pie charts) even when more complex designs (e.g., heatmaps) yielded richer insights, and higher [[visualization|visualization literacy]] predicted deeper, more detailed interpretation (e.g., more teachers identifying trends in time-series data). For group comparison, teachers favored superposition over juxtaposition and full-information plots over explicit difference encoding. The actions teachers proposed were shaped by the content represented and their [[teacher-role]] level rather than the plot type — university teachers favored weekly tests and course-level adaptation, while vocational teachers proposed direct, individualized coaching — underscoring that the prescriptive step depends as much on how analytics are visualized and contextualized as on the underlying model. Their finding is cautionary: correlations between LA indicators and recommended support were mostly weak, cross-model recommendations diverged sharply for the same student, and support was frequently allocated regardless of who needed it most. The authors conclude that current LLMs are **not yet reliable as prescriptive models for student support at scale**, reinforcing that the prescriptive step still requires validation, fine-tuning, and human oversight rather than off-the-shelf automation. ### Methods and network analysis Network methods are core to learning analytics: [[network-analysis|transition network analysis (TNA)]] models temporal sequences of learner actions (e.g., the revision and chat loops in [[conversational-ai|chatbot]]-scaffolded writing), and [[network-analysis|epistemic network analysis (ENA)]] maps how codes/constructs co-occur across activity — together revealing the *process* of learning and learner-[[student-ai-interaction|AI interaction]] rather than only its product.([[penny-transition-network-analysis-efl-writing-2026]])([[tracing-genai-literacy-interaction-patterns]]) - **Sequence + Markov-chain analysis of self-directed behavior.** [[an-goel-self-directed-modeling-2026|An, Hammock & Goel (2025)]] combined activity sequence analysis, hierarchical clustering, and Markov chain models on the clickstreams of 315 online learners who built 822 ecological models in VERA, distilling nine fine-grained, transition-based behavioral clusters into three broader patterns (Observation, Construction, Exploration). Their work demonstrates that combining sequence analysis with Markov chain modeling can uncover meaningful behavior in unstructured, [[self-directed-learning|self-directed]] tasks even in the complete absence of demographic or contextual data. - **LA and GenAI shape learning design differently (2026).** [[claassen-learning-analytics-genai-learning-design-2026|Claassen et al. (2026)]] used ENA on 11 instructor focus groups to compare how learning analytics versus [[generative-ai]] inform [[learning-design]] decision-making. LA discussions centered on contextual information, course-level design, and creative [[problem-solving]] (LA for diagnosing engagement and targeting support), while GenAI discussions centered on [[assessment|assessment design]] and designing for student [[self-determination-theory|self-determination]] (GenAI for ideation and assessment development). Context and [[creativity]] were central across both — a reminder that analytics inform design only within [[pedagogy|pedagogical]] context and instructor autonomy. - **Design analytics: mining planned activity sequences rather than traces (2026).** [[learning-paths-patterns-learning-design-2026|Divjak, Svetec and Horvat (2026)]] applied Markov chains and sequential pattern mining to the *designed* sequence of 29,064 teaching and learning activities across 554 courses planned in an open learning-design tool. The transition matrix peaked at Assessment to Discussion (0.332), self-transitions dominated Practice (0.317) and Acquisition (0.292), and the strongest consecutive rule was Acquisition to Assessment to Practice to Practice (confidence 0.743, lift 1.449), while the most frequent four-step path was Acquisition, Practice, Practice, Assessment (120 occurrences). Discussion and Assessment were the most reachable types and Production the most distant and sporadic. The study is a reminder that learning analytics need not begin with LMS traces: design-time data can expose the pedagogical grammar of a course before any student arrives, though the authors stress that resemblance to flipped, inquiry-based or [[project-based-learning|project-based]] designs is not evidence of intent. - **Self-explaining distilled LLMs (2026):** A two-stage pipeline distills a black-box learning-analytics estimator and its post-hoc interpretation into a small, open-weight [[llm]] that returns both an individual-level estimate and a natural-language explanation. A faithfulness-first audit evaluates whether narrations match the attributions they describe; [[simulation]] shows near-lossless recovery (r > .90) with an oracle mentor, offering a more transparent, deployable path for analytics ([[distilling-self-explaining-lm-learning-analytics-2026]]). - **Enablers of LA-based educational interventions (2026).** [[learning-analytics-to-educational-interventions-2026|Svetec, Divjak & Kadoić (2026)]] identify and prioritize seven enablers of trustworthy LA-based educational interventions via Delphi + AHP + SNAP: [[governance|institutional]] strategic orientation, pedagogical & other [[research-methods-aied|research]] foundations, available resources, pedagogical support, ethics & data governance, stakeholder engagement, and quality assurance. Institutional strategic orientation ranked highest (and most influential on other enablers), with available resources second. [[trust|Trustworthiness]] (ethical compliance, transparent/unbiased algorithms, pedagogical validity) is framed as the prerequisite without which LA-based interventions are not meaningful. - **LLM interaction depth predicts task quality but not recall (2026).** [[llm-interaction-depth-task-quality-recall-2026|Tsiligkiris (2026)]] links turn-level LLM conversational telemetry (Depth/Volume/Pacing) to [[learning-gains|learning outcomes]]: explanation-seeking "depth" predicted independently marked task quality (β = 6.27) but not immediate recall — a dissociation between elaboration-driven comprehension and retrieval-driven consolidation that has implications for how [[llm]] interaction is measured and evaluated in LA. - **Simulating collaborative discourse for learning analytics.** [[llm-agents-collaborative-problem-solving-simulation-2026|Fang (2026)]] uses fine-tuned participant-specific LLM agents to reproduce collaborative problem solving dialogues, validated with Epistemic Network Analysis (ENA distance 0.17, permutation p = 0.65). The approach offers learning-analytics researchers a scalable way to generate authentic collaborative discourse for studying interaction dynamics, turn-taking, and thematic code trajectories without collecting new human data. - **Open, reproducible data and trace-ready analytics.** [[astra-multi-agent-tutoring-benchmark-2026|ASTRA]] releases a synthetic [[benchmark]] with a trace-ready schema (N=540; 360 sessions; 1,440 episodes) for analyzing interaction and participation balance in collaborative programming. Log and trace data are the natural counterweight to [[self-report-measures|self-report]] in this literature: the same construct is often measured twice, once by asking and once by observing, and the two do not always agree. Separately, an exploratory ML framework with SHAP analysis identified the learning-related constructs most associated with intended academic ChatGPT use among university students, prioritizing [[explainable-ai|interpretability]] ([[determinants-chatgpt-use-higher-education-2026]]). ### Connections Learning analytics connects to [[knowledge-tracing]] (the core analytic), [[formative-assessment]] (analytics-driven assessment), [[student-modeling]] (the learner representation analytics populate), [[privacy]] (the [[ethics|ethical]] constraint), and [[edtech-platform]] (where analytics are deployed). Because prescriptive analytics are increasingly evaluated on [[simulating-students|simulated learners]] — where synthetic student cohorts substitute for real cohorts in controlled tests — learning analytics also connects to student simulation. - **What dashboards make visible decides what teachers act on.** [[ai-supported-lecturer-decision-making-2026|Köroğlu et al. (2026)]] reviewed 27 empirical studies (2016–2025) and built a socio-technical taxonomy of AI-supported lecturer decision-making across eight decision types: instructional, curriculum, assessment, feedback, learning-environment, emotional, administrative and ethical. Learning Analytics Dashboards were the most frequently reported system, and the coding shows support concentrated in the instructional, feedback and assessment decisions that behavioral text and log data can inform, while emotional, ethical, curriculum and learning-environment decisions were rarely supported. The authors read this as an attention effect rather than a capability limit: because the systems rendered behavioral student data visible and actionable, motivation, [[metacognition]], emotion and environment concerns fell outside what the data invited lecturers to consider. Process-level instrumentation is the descriptive layer's next step down. [[pulla-parsons-problem-tool-2026|Prol et al. (2026)]] extended an [[open-source]] Parsons-problem platform to record every block placement, removal and submission as a chronological trace, pair each submission with per-block correctness coloring and an attempt history, and optionally pass the traces through an [[llm]] pipeline that labels recurring difficulty patterns for instructor review. Across 68 students in an upper-division Java software-design course and 36 in an introductory Python course the analysis surfaced the same three difficulties — choosing the wrong exception type, substituting `return` for `throw`, and incorrect control-flow ordering — which correctness-and-attempt counts cannot expose, because they answer *whether* the arrangement was right rather than *what the process was*. The AI acts as an interpreter for the [[teacher-role|instructor]] rather than a scorer of the student, and its output is framed as a reviewable hypothesis about a class's difficulties rather than a grade. ## Connected Concepts - [[explainable-ai]] - [[knowledge-tracing]] - [[student-modeling]] - [[formative-assessment]] - [[privacy]] - [[edtech-platform]] - [[student-engagement]] - [[ai-ed-evaluation]] - [[feedback]] - [[higher-ed]] - [[k-12]] - [[llm]] - [[simulating-students]] - [[self-report-measures]] - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[ai-supported-lecturer-decision-making-2026]] — AI-Supported Lecturer Decision-Making in Higher Education - [[villegas-ch-federated-explainable-learning-analytics-2026]] — Federated and explainable learning analytics for privacy-preserving academic risk modeling (Villegas-Ch et al. 2026) - [[llm-interaction-depth-task-quality-recall-2026]] — What students ask matters: LLM interaction depth, task quality, and immediate recall (Tsiligkiris 2026) - [[learning-analytics-to-educational-interventions-2026]] — From learning analytics to educational interventions: enablers of trustworthy LA-based interventions (Svetec, Divjak & Kadoić 2026) - [[claassen-learning-analytics-genai-learning-design-2026]] — LA and GenAI in learning design decision-making - [[at-risk-students-ml-prediction]] - [[engagement-intensity-learner-modeling]] - [[misiejuk-cognitive-offloading-prompting-2026]] - [[teaching-feedback-classification-benchmark]] - [[wordstream-glass-learning-analytics]] - [[trace-course-grade-prediction-2026]] - [[student-llm-interaction-taxonomy-review-2026]] - [[sc2r-counterfactual-recourse-educational-2026]] — From Student Risk Prediction to SC2R: Counterfactual Recourse - [[bayesian-cognitive-diagnosis-personalized-learning-paths]] — Bayesian cognitive diagnosis for personalized learning paths - [[distilling-self-explaining-lm-learning-analytics-2026]] — Distilling self-explaining LM for learning analytics - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[llm-agents-collaborative-problem-solving-simulation-2026]] — Fine-tuned participant-specific LLM agents reproducing collaborative problem solving dialogues (Fang 2026) - [[astra-multi-agent-tutoring-benchmark-2026]] — ASTRA synthetic benchmark for multi-agent tutoring and participation-balanced collaboration - [[determinants-chatgpt-use-higher-education-2026]] — ML/SHAP determinants of future ChatGPT use in higher education - [[mejia-domenzain-ml-findings-teachers-blended-2026]] — Making ML findings accessible to teachers in blended classrooms - [[pradeesh-outcome-knowledge-tracing-affinity-2026]] — Outcome-based knowledge tracing with affinity mapping - [[an-goel-self-directed-modeling-2026]] - [[schuetze-knowledge-tracing-forgetting-2026]] - [[zhang-ml-student-progress-programming-2026]] - [[learning-paths-patterns-learning-design-2026]] — Markov chain and pattern mining of 29,064 planned activities in 554 courses, revealing a design grammar led by Acquisition and consolidating Practice - [[pulla-parsons-problem-tool-2026]] — Pulla: process-level behavioral tracing and instructor-facing difficulty analysis in Parsons problems (Prol et al. 2026) - [[a4l-analytics-pipeline]] - [[huang-interpretable-knowledge-tracing-2026]] - [[league-ethical-governance-student-data-2026]] - [[precision-education-student-digital-twins-2026]] - [[learning-analytics-genai-secondary-writing-2026]] — Using Learning Analytics to Support Secondary School Students' Writing with Generative AI --- ## [Technology Adoption Models](https://edtechdev.github.io/aied/concepts/technology-acceptance-model/) **Technology adoption models** are the theoretical frameworks used to explain and predict why individuals and institutions accept, adopt, and continue using new [[ai-technologies|technologies]] — and, in AI-in-education research, why learners, teachers, and organizations adopt [[generative-ai|generative AI]] tools. Rather than a single model, this is a family of theories that share roots in information-systems and social-psychology research, of which the **Technology Acceptance Model (TAM)** is the most widely applied. The knowledge base treats these models together because GenAI-adoption studies routinely combine them (TAM + UTAUT, TAM + TPB, UTAUT + ARCS) and because their core constructs — perceived usefulness, perceived ease of use, and social influence — recur across nearly every study of AI acceptance in education. ## Questions to Consider - Think of a time you adopted a new app, tool, or AI service — and a time you abandoned one. What actually drove each decision: how useful it seemed, how easy it was, or what people around you were doing? Which factor do you suspect matters most, and can a survey really capture that? - A common belief is that 'if a tool is clearly useful, people will use it.' Have you seen situations where a genuinely useful technology still failed to catch on, or a clearly limited one spread anyway? What might explain the gap between objective usefulness and actual adoption? - Adoption frameworks like TAM were designed in the 1980s for fairly simple systems. If you've used generative AI, in what ways does it differ from a word processor or a [[edtech-platform|learning management system]] — and why might a model built around 'ease of use' and 'usefulness' struggle to capture how people relate to something that talks back? - [[research-methods-aied|Researchers]] often say perceived risk and trust matter less for AI adoption than expected. Before reading further, what do you predict: do students adopt AI because they trust it, despite risks, or are those concerns actually minor next to convenience and social pressure? - The page argues that adoption frameworks treat using a tool as a one-time decision, but effective use may be an ongoing judgment. Where in your own or your students' practice does the line between 'choosing to use AI' and 'continually deciding how to use it well' seem to blur — and what would change if we measured that instead of mere uptake? - Some researchers cluster learners into different 'adoption personas' instead of assuming one model fits everyone. What differences do you see among your own learners, colleagues, or students that a single average model of adoption might hide — and how could those differences shape how you support them? ## Introduction **[[meta-analysis-systematic-review|Meta-analytic]] evidence on AI adoption.** A [[teo-ai-adoption-tertiary-meta-analysis-2026|meta-analysis of tertiary students' AI adoption]] (233 correlations, 32 studies, N = 16,977) finds moderate positive correlations for individual (r = 0.57), contextual (r = 0.53), and technological (r = 0.50) factors, with usage intentions the strongest predictor (r = 0.64) and perceived risks/trust weaker than expected. Its central critique is that the field **over-relies on traditional TAM/UTAUT** frameworks that predate modern intelligent systems and neglect AI-specific factors such as anthropomorphism and ethics — arguing these gaps matter for advancing theory and evidence-based policy. ## The model family ### Technology Acceptance Model (TAM) Proposed by Davis (1989), TAM explains adoption through two core beliefs — **Perceived Usefulness (PU)** and **Perceived Ease of Use (PEOU)** — which shape users' **Attitude (ATT)**, then **Behavioral Intention (BI)**, then actual usage. Grounded in information-systems theory (adapted from the Theory of Reasoned Action), TAM has become the dominant framework for [[generative-ai|GenAI]] adoption in education. In educational GenAI research it is frequently extended with [[ai-literacy]], [[trust]], social influence, [[self-determination-theory|self-determination]], and **critical use** to capture adoption complexity beyond simple uptake. ### UTAUT / UTAUT2 / UTAUT3 The **Unified Theory of Acceptance and Use of Technology** consolidates TAM with eight prior models into four core determinants — performance expectancy, effort expectancy, social influence, and facilitating conditions (with UTAUT2 adding hedonic motivation, price value, and habit; UTAUT3 adding personal innovativeness). The knowledge base applies UTAUT across teacher and student populations: [[mathematics-teachers-chatbot-motivation-2026|Austrian secondary math teachers]] (UTAUT with 448 teachers), [[pre-service-science-teachers-ai-perceptions-2026|pre-service science teachers in Ghana]] (UTAUT + TPB), and [[tian-genai-learning-adoption-pathways-2026|students in Lesotho]] (UTAUT3 + ARCS, using PLS-SEM and fsQCA). [[jacome-vasconez-chatgpt-adoption-xai-2026|Jácome-Vásconez et al.]] extend UTAUT2 with explainable AI (Random Forest, SHAP, NCA, IPMA, K-Means) for 522 university students, finding **habit** the strongest predictor of ChatGPT intention and showing that **effort expectancy is a necessary condition rather than a linear driver** — a result only visible when XAI complements the regression. ### Theory of Planned Behavior (TPB) TPB explains intention through attitude, subjective norms, and perceived behavioral control. It is frequently paired with TAM/UTAUT in AI-adoption studies — e.g. [[genai-chatgpt-adoption-ethics-students-2026|Rizun et al.]] integrate TAM, TPB, UTAUT, and the FATE ([[bias-mitigation|Fairness]], Accountability, Transparency, Ethics) framework to model the behavioral and [[ethics|ethical]] drivers of student ChatGPT adoption. ### Diffusion of Innovation (DOI) Rogers' DOI theory explains adoption as a social process in which innovations diffuse through populations over time, emphasizing innovation attributes (relative advantage, compatibility, complexity, trialability, observability) and adopter categories. It appears in the knowledge base's [[governance|institutional]]-level analyses, e.g. [[alrahmi-org-drivers-ai-adoption-he-2026|Al-Rahmi et al.]] combine the Technology–Organization–Environment (TOE) framework with DOI to model organizational AI adoption in Saudi universities. ### Technology–Organization–Environment (TOE) TOE frames adoption as shaped by technological, organizational, and environmental contexts — a complement to individual-level TAM/UTAUT for studying institutional adoption (see [[alrahmi-org-drivers-ai-adoption-he-2026]]). ## Applications in the knowledge base Adoption models are applied across the knowledge base to model student and [[teacher-role|teacher]] uptake of AI tools: - **Critical-use extension:** [[tam-critical-use-genai-engineering-2026|Nguyen et al.]] extended TAM with *critical use* for engineering/CS students, finding that attitudes and critical use directly predict intention — and that critical use safeguards against [[cognitive-offloading|over-reliance]]. - **Unified socio-cognitive model:** [[socio-cognitive-genai-adoption-engineering-2026|Asag & Al Mamun]] integrated TAM with UTAUT for engineering students in Bangladesh (explaining 64% of usage variance). - **Person-centered profiles:** Rather than variable-centered models, [[saihi-ahmed-genai-adoption-personas-higher-ed-2026|Saihi & Ahmed]] cluster GenAI-adoption personas, and [[chen-preservice-teachers-chatgpt-lpa-2026|Chen et al.]] use latent profile analysis to identify four ChatGPT-acceptance profiles among pre-service teachers — showing the field's move beyond single-model, linear accounts. - **Cross-cultural validation:** [[motivation-shape-future-education-ai-switzerland-china|Martínez-Moreno et al.]] cross-culturally validate adoption-related motivation constructs across Switzerland and China. - **Psychological correlates:** [[acceptance-ai-english-tools-2026|Wu et al.]] build on TAM to relate motivation, [[self-efficacy]], anxiety, and risk perception to acceptance of AI-assisted English learning. - **[[regulation|Regulatory]] competence critique:** [[ai-anxiety-strategic-regulation-writing-2026|Kim]] argues that adoption-centered TAM models treat use as a stable decision, whereas effective AI use is an ongoing process of judgment, revision, and selective uptake — reframing [[ai-literacy]] as regulatory competence and [[critical-thinking]] rather than acceptance. - **ML determinants of ChatGPT adoption:** An exploratory ML approach examined how students' perceptions and demographics relate to intended academic ChatGPT use, using SHAP analysis to identify key learning-related constructs — prioritizing educational meaning over maximizing algorithmic performance ([[determinants-chatgpt-use-higher-education-2026]]). - **[[explainable-ai|Explainability]] and domain relevance as acceptance levers for teachers:** Adapting trust-in-automation theory to teacher acceptance of AI recommendations, [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor et al. (2025)]] found that understandability (raised by explainable AI) correlated positively with both trust and acceptance of an AI grouping tool, and that domain-driven explanations in [[curriculum-design|curricular]] language outperformed data-driven feature-importance ones on all three. Acceptance was additionally driven by [[pedagogy|pedagogical]] alignment and workload-reduction potential — situational factors beyond trust that standard TAM/UTAUT constructs rarely capture, reinforcing the case for extending adoption models with context and explainability. - **Acceptance measured at the level of instructional materials (2026).** [[age-tiered-ai-literacy-guidebooks-2026|Wang, Chuang and Wu (2026)]] operationalized Performance Expectancy, Effort Expectancy, Perceived Playfulness, and Behavioral Intention for two age-tiered AI literacy guidebooks used by 794 [[k-12]] students and 37 teachers, validating the four-factor structure with split-sample EFA and CFA and supporting measurement invariance across the 9-12 and 13-18 editions. Younger learners reported higher levels on all four constructs, and playfulness was the strongest correlate of intention in both cohorts — with the standardized playfulness-to-intention coefficient exceeding 1.00 in the younger group and an HTMT of.950 between the two constructs. The authors read that as construct overlap and possible suppression rather than a giant effect, a useful reminder that acceptance instruments applied to materials can produce empirically entangled factors that complicate structural interpretation. ## Limits and extensions TAM's cognitive focus also under-captures emotional and relational dimensions of AI use. A [[akbaba-nursing-ai-experiences-tam-2026|qualitative study of nursing students and faculty]] found that, alongside the four classic TAM constructs, participants -- mainly students -- described AI as a source of psychosocial support: emotional comfort during stress and a confidential space for reflection. This extends perceived usefulness beyond performance gains to [[well-being]] in high-stress [[professional-training|professional training]], suggesting adoption models should attend to affective and psychosocial factors, not only instrumental utility. **[[agency|Autonomy]] perception and risk aversion as tool-specific extensions.** A [[mixed-methods-research|mixed-methods]] survey of 287 GenAI-experienced teachers across 27 countries and regions ([[dai-genai-frenemy-teaching-autonomy-2026|Dai et al. (2026)]]) extends TAM in a direction the model family has barely touched: the perceived *autonomy* of the tool itself. Adding perceived artificial autonomy (how independently teachers believe GenAI can carry out an assigned instructional task) and risk aversion to the classic constructs, their refined structural model explained 79.8% of the variance in behavioral intention (χ²(97) = 202.597, CFI =.963, TLI =.954, RMSEA =.062, SRMR =.041), with perceived usefulness again dominant (PU → BI β =.832, p <.001) and ease of use feeding usefulness (PEU → PU β =.479). The instructive results are the detours. Artificial autonomy did not act directly on intention (β =.026, ns) but acted through usefulness (AA → PU β =.188, p <.001; indirect AA → PU → BI =.156, 95% CI [.082,.240]) — teachers who saw GenAI as more capable found it more useful, not more immediately adoptable. Risk aversion lowered intention (RA → BI β = −.163, p <.001) without lowering perceived usefulness (RA → PU ns), because teachers' concerns centered on *students'* use — cheating and [[academic-integrity|integrity]], weakened foundational knowledge and higher-order thinking, [[hallucination-risk|hallucinations]], reduced human interaction, and ethical, legal and [[equity-in-ai-education|equity]] risks — rather than on their own. Ease of use was non-significant overall (PEU → BI β =.068) but significant among the 225 teachers who had actually used AI in teaching (β =.138, p <.05), the familiar pattern that effort registers only after real encounter; attitude had to be dropped when its correlation with intention (r =.929) breached discriminant validity. Teachers placed GenAI at semi-autonomous levels (130 at Level 2, "teacher assistance"; 94 at Level 3, "partial automation"; one at Level 6) and insisted adoption decisions be made case by case. TAM still predicts intention well, but for teacher-facing AI the constructs that move the needle are tool-specific: how autonomous the tool seems, and whose use the risk attaches to. The authors' "frenemy" label — valued as a support tool, distrusted as an autonomous agent — is the teacher-side counterpart to the student-facing extensions catalogd above. **Risk perception as a dimension-specific extension.** A survey of 814 Chinese university students ([[risk-perception-genai-perceived-benefits-2026|Du, Ning, Shi & Chen (2026)]]) folds Cognitive Appraisal Theory and Protection Motivation Theory into TAM/UTAUT2 to ask not whether students adopt [[generative-ai|generative AI]] but what they gain from it. The model family's assumption that risk uniformly suppresses adoption does not survive: risk perception splits into information, security, technical, [[ethics|ethical]], and legal dimensions that move benefits in opposite directions. Security risk — a threat students believe they can manage through [[privacy|privacy practices]] — was *positively* associated with perceived academic assistance and skill development, consistent with problem-focused coping, whereas information risk, which learners cannot easily verify, eroded psychological and emotional support, daily-life, and leisure benefits through avoidance. Threshold analyses further showed effects that change sign beyond dimension-specific cut-points, and usage experience mattered independently: students with more than a year of GenAI use reported higher benefits across four domains. The practical implication for acceptance research is that "perceived risk" is too coarse a construct to model — controllability, not the mere presence of risk, is what determines whether students engage or withdraw. ## Connected Concepts - [[business-education]] - [[generative-ai]] - [[ai-literacy]] - [[student-experience]] - [[higher-ed]] - [[critical-thinking]] - [[ethics]] - [[trust]] - [[cognitive-offloading]] - [[self-determination-theory]] - [[research-methods-aied]] - [[student-modeling]] - [[framing-ai-use-for-students]] ## Connected Articles - [[dai-genai-frenemy-teaching-autonomy-2026]] — GenAI as "frenemy": artificial autonomy and risk aversion extend TAM for teachers (Dai et al. 2026) - [[akbaba-nursing-ai-experiences-tam-2026]] — Nursing AI experiences; psychosocial extension of TAM - [[preschool-teachers-ai-behavioral-intention-2026]] — Preschool teachers' behavioral intention to use AI via extended TAM (Duan, Shan & Gong 2026) - [[ai-adaptation-gap-higher-education-2026]] — The AI Adaptation Gap in Higher Education - [[saihi-ahmed-genai-adoption-personas-higher-ed-2026]] — GenAI adoption personas via clustering - [[tian-genai-learning-adoption-pathways-2026]] — Symmetric and asymmetric pathways in GenAI adoption (UTAUT3 + ARCS) - [[lee-wu-gender-motivation-genai-achievement-2026]] — Differential GenAI engagement by gender and motivation - [[alrahmi-org-drivers-ai-adoption-he-2026]] — TOE + DOI model of organizational AI adoption in higher education - [[tam-critical-use-genai-engineering-2026]] — Extended TAM with critical use for engineering/CS students - [[socio-cognitive-genai-adoption-engineering-2026]] — Unified socio-cognitive model (TAM + UTAUT) for engineering education - [[ai-anxiety-strategic-regulation-writing-2026]] — From AI anxiety to strategic regulation - [[genai-reliance-types-scale]] — GenAI reliance types scale - [[llm-reliance-types-undergrad]] — LLM reliance types among undergraduates - [[acceptance-ai-english-tools-2026]] — Acceptance of AI English tools - [[genai-chatgpt-adoption-ethics-students-2026]] — Behavioral and ethical drivers of student ChatGPT adoption - [[mathematics-teachers-chatbot-motivation-2026]] — UTAUT and teacher chatbot motivation - [[pre-service-science-teachers-ai-perceptions-2026]] — UTAUT + TPB for pre-service science teachers in Ghana - [[chen-preservice-teachers-chatgpt-lpa-2026]] — Pre-service teacher ChatGPT acceptance profiles (LPA) - [[teo-ai-adoption-tertiary-meta-analysis-2026]] — Meta-analysis of AI adoption factors; critiques TAM/UTAUT - [[ethical-conditions-llm-exam-preparation-2026]] — Ethical conditions for LLM adoption in exam preparation (Pérez-Portabella et al. 2026) - [[genai-integration-constructivist-higher-ed-bangladesh-2026]] — GenAI integration in Bangladeshi higher ed through constructivism (Alam et al. 2026) - [[determinants-chatgpt-use-higher-education-2026]] — ML/SHAP determinants of future ChatGPT use in higher education - [[jacome-vasconez-chatgpt-adoption-xai-2026]] — XAI-augmented UTAUT2: habit as strongest predictor, four adoption profiles (Jácome-Vásconez et al. 2026) - [[xai-teachers-trust-edtech-recommendations-2026]] - [[age-tiered-ai-literacy-guidebooks-2026]] — Material-level PE/EE/playfulness/intention model for age-tiered AI literacy guidebooks, with documented playfulness-intention construct overlap - [[risk-perception-genai-perceived-benefits-2026]] — CAT + PMT folded into TAM/UTAUT2: dimension-specific and non-monotonic risk effects on perceived GenAI benefits (Du, Ning, Shi & Chen 2026) --- ## [Open Source](https://edtechdev.github.io/aied/concepts/open-source/) > **Open Source** — the use of openly licensed *models, code, data, and content* in [[ai-education|AI in education]]. Openness is the knowledge base's main counterweight to vendor lock-in and [[privacy|data exposure]]: open-weight models can run on campus hardware to satisfy FERPA, GDPR, and EU AI Act obligations, openly licensed corpora can be indexed and fine-tuned without publisher permission, and open benchmark and dataset releases make [[research-methods-aied|research]] replicable. The burdens are equally real: infrastructure and [[pedagogical-safety|safety]] assurance, maintenance that outlives the grant, and quality that openness does not by itself guarantee. ## Questions to Consider - "Open source" is often heard as "free and easy." Which of the three layers below — models, code, or content — actually costs your institution the most to adopt, and why? - Open weights make local deployment possible, but someone still has to host, patch, and evaluate the system. Who should own that work after the initial project ends, and who pays for it? - One study here found a 32B open model outscoring a far larger proprietary system on pedagogical knowledge, while another found every open model tested below the human baseline on scientific-visualization literacy. How do you decide which benchmark is the right one for your decision? - A single openly licensed corpus is what lets a school run an on-premise assistant over its own course materials. What obligations come with that — to the original authors, to students whose data is indexed, and to the license itself? - If [[generative-ai|generative AI]] can produce a course in under half an hour for a couple of dollars, what is the remaining rationale for open educational resources — cost, licensing freedom, quality assurance, or something else? - Should institutions treat open-source adoption as a procurement decision, an infrastructure decision, or a pedagogical one? What breaks if it is treated as only one of them? ## Introduction Loosely, "open" in AI education means that four kinds of artifact are available for inspection, reuse, and modification: **model weights**, **source code**, **data and evaluation instruments**, and **educational content**. The knowledge base's articles cluster unevenly across these layers, and the resulting picture is more useful than the slogan: openness buys specific things — local control, auditability, replicability, and legal clarity — and it costs specific things — infrastructure, expertise, maintenance, and a quality-assurance burden that moves from the vendor to the institution. ### Open models and open weights Open weights matter most where student data cannot leave campus. [[lata-ferpa-compliant-local-llm-autograder|LaTA]] is a drop-in, FERPA-compliant local-[[llm]] autograder for upper-division [[stem-education|STEM]] coursework built on instructor-authored rubrics and reference solutions, with zero marginal cost per submission. [[programming-its|SCRIPT]], a Python [[intelligent-tutoring|tutoring system]] at Bielefeld University, deliberately **avoids commercial LLM APIs** and self-hosts an open-weight Llama-70B model to meet the GDPR and the EU AI Act (which classifies some AI-in-education uses as high risk), separating IP logs from the tutoring system, using pseudonymous usernames, and recording keystrokes only with explicit consent — a choice the authors also credit with lower environmental impact and better reproducibility. Quality is no longer the automatic price of openness. [[singh-eduqwen-pedagogical-rl-2026|EduQwen]] applies [[reinforcement-learning|reinforcement learning]] (DAPO) and supervised fine-tuning to an open model family, mining 440 hard negatives, generating 40,000 synthetic responses down-selected to 1,050 difficulty-ordered examples, and reaching **96.52%** on the pedagogy benchmark — above Gemini-3 Pro's 90.55% — at 32B dense parameters. [[aiawe-automated-writing-evaluation|AiAWE]] reaches similar conclusions for [[automated-assessment|automated writing evaluation]]: a LoRA-adapted open-weight Gemma-3-27B-it outperforms LLaMA-3.3-70B and a fine-tuned GPT-3.5 baseline on 480 TOEFL essays and runs on a consumer-grade server, with the striking subsidiary finding that parameter count is *not* a reliable predictor of downstream performance under LoRA adaptation. The counter-evidence deserves equal billing: [[mllm-scientific-visualization-literacy|a benchmark of six MLLMs]] (three closed, three open) found every open-source model below the human baseline on scientific [[visualization]] literacy while Gemini exceeded the human mean on several subsets. Openness raises the ceiling on control, not on capability. ### Open tools, tutors, and research infrastructure The clearest case for open code is replication. [[oatutor-open-source-adaptive-tutor-2023|OATutor]] — the first fully open adaptive tutoring system built on ITS principles — pairs an **MIT-licensed codebase** with a **Creative Commons (CC BY) content library** from OpenStax algebra textbooks, plus [[knowledge-tracing|knowledge tracing]], A/B testing, and LTI support; its explicit design goal is that a researcher can run an experiment and then publish the entire end-to-end framework, content, and platform as a repository link. [[stanbkt-bayesian-knowledge-tracing|StanBKT]] makes the same argument at the method layer, replacing expectation-maximization point estimates with full Bayesian inference (HMC, variational inference, Pathfinder, optimization) in an open Python package that exposes the uncertainty A/B comparisons of adaptive interventions depend on. [[deeptutor]] releases a complete [[agentic-ai|agentic]] tutoring framework with a trace-forest learner memory — Apache 2.0, and by late 2026 a full learning workspace rather than only the pipelines its paper benchmarks — and MAIC's classroom generator OpenMAIC ships under MIT alongside the study that evaluates it, so a course can be generated, self-hosted and inspected rather than only read about; [[vismatic-secure-sandbox-cs-education|VISMATIC]] publishes its containerized sandbox for process-oriented monitoring, so other institutions can adopt the integrity model rather than the vendor's version of it. Open code also carries the transparency burden: the scripts for [[kar-mathbuddy-affective-math-tutoring-2025|MathBuddy]]'s affect-aware tutoring prompts are published for inspection and further [[pedagogical-llm-training|pedagogical training]] work. Open infrastructure sets the reference point for what educational agents should do: the knowledge base's [[agentic-ai-education-scoping-review|scoping review of 474 agentic-AI studies]] uses a fast-growing open-source agent project as its "frontier agent paradigm" benchmark and finds educational systems still lacking governed tool orchestration, persistent memory, long-horizon planning, and auditable action. ### Open benchmarks, datasets, and method transparency Several contributions here are open *evaluation infrastructure* rather than systems. [[cdpk-pedagogy-benchmark-llms|The Pedagogy Benchmark]] (CDPK + SEND, built from genuine Chilean teacher-exam items) spans 97 models: open-weight DeepSeek R1 reached 86.65% against a top-10 of mostly closed reasoning models, and the cost–accuracy frontier moved from ~50% to ~82% at \$0.10/M input tokens between April 2024 and June 2025 — with open Qwen-3 8B at 3.5¢ nearly matching the best April-2024 closed model at over 400× lower cost. Performance drops sharply below roughly 8B parameters, which is a practical sizing constraint for campus deployments. [[astra-multi-agent-tutoring-benchmark-2026|ASTRA]] releases a dataset, schema, and prototype for trace-based evaluation of socially intelligent multi-agent tutoring (540 participants, 360 sessions, 1,440 task episodes). [[iks-instruct-dataset-indian-knowledge|IKS-Instruct]] shows the cultural case for open data: 24,795 instruction–response pairs across seven languages and 41 pedagogical techniques drawn from Vedic and classical sources, aligned to the CBSE [[curriculum-design|curriculum]], which let a compact 7B model approach a much larger general-purpose reference model (median judge score 6.39 vs 6.54) at a fraction of the deployment cost. Student-authored [[benchmark|benchmarks]] are another route to openness: [[yu-academiclaw-student-challenges-ai-agents-2026|AcademiClaw]] curates 80 long-horizon academic tasks from 230 student-submitted candidates (spanning 25+ professional domains, 16 requiring CUDA GPUs, run in isolated Docker sandboxes) and extends an open agent ecosystem into academic-level evaluation. [[aied-carbon-footprint-reporting|Eimler et al. (2026)]] argue that openness is an [[sustainability|environmental]] obligation as well: reviewing all AIED 2025 papers, they found an "LLM adoption without disclosure" pattern and responded with an open-source measurement methodology — software tools plus a formula that estimates computational expense even when parameter counts are unknown. ### Open educational resources and open content [[shen-sustainable-ai-knowledge-base-cs-education-2026|Shen et al. (2026)]] is the knowledge base's only article in which **open educational resources are the central object** rather than a passing reference. They build an on-premise AI knowledge-base assistant for [[cs-education|computer science education]] from 82 OER documents on consumer-grade hardware (an RTX 3060 with 12 GB VRAM), combining structured extraction, [[rag|retrieval-augmented generation]], and NF4 4-bit quantization-aware fine-tuning. Fine-tuning added real value beyond retrieval (Qwen-7B 69.8%, +3.2 pp, p = 0.031; DeepSeek-MoE 78.6%, +12.0 pp, p < 0.001, including 82.3% on multi-hop reasoning); quantization-aware tuning held the 4-bit accuracy gap to 1.7 and 1.2 pp while cutting VRAM by ~38% and energy to 1.8 mWh per query (43.8% below baseline); and quantization-inflated [[hallucination-risk|hallucination]] was partly recovered by fine-tuning (DeepSeek-MoE 10.4% → 8.1%), measured by a two-stage NLI procedure against retrieved OER chunks. The analytical point is generalizable: an openly licensed corpus can be indexed, adapted, and served without publisher permissions, and grounding an assistant in retrieved OER gives a checkable provenance trail — which is exactly what a proprietary textbook corpus cannot offer. Openness of content and openness of models are complements elsewhere too. OATutor curates CC BY OpenStax textbooks into a system whose code is MIT-licensed, so the license terms of code and content have to be kept compatible by design. [[egai-power-systems-education|An open, executable module library for power systems AI]] lowers the entry barrier with Jupyter notebooks that run locally or in Colab, delivered through an IEEE online course. And a project-based [[engineering-education|mechanical engineering]] curriculum publishes its syllabus, data, and code in open-access repositories so other institutions can adopt it ([[mechanical-engineering-ai-curriculum-2026]]). Adjacent to OER, open *course* delivery is where the economics are shifting fastest: [[mooc-to-maic|MAIC]] reports collapsing MOOC production from roughly \$25,000 and 60 hours per course to under \$2 and 30 minutes with LLM-driven multi-agent generation. If content production becomes nearly free, the OER argument moves away from production cost and toward licensing freedom, verifiability, and quality assurance — which is a different proposition from the one OER advocacy was built on. ### Benefits and burdens - **Benefits.** Data sovereignty and regulatory compliance ([[privacy]], [[regulation]]) through local hosting; cost control, since local inference has no per-query fee and open models approach proprietary quality at a fraction of the price ([[singh-eduqwen-pedagogical-rl-2026]]); reproducibility, because the framework, prompts, content, and data can ship with the paper ([[oatutor-open-source-adaptive-tutor-2023]], [[astra-multi-agent-tutoring-benchmark-2026]]); auditability and safety review under [[governance|institutional governance]]; and the ability to fine-tune for a specific [[pedagogy]] or a specific community's knowledge base ([[iks-instruct-dataset-indian-knowledge]]). - **Burdens.** Local hosting requires hardware and expertise many institutions do not have; quality and [[pedagogical-safety|safety]] are not guaranteed out of the box — open models may trail human baselines on specific literacies ([[mllm-scientific-visualization-literacy]]) and quantization raises hallucination rates unless mitigated ([[shen-sustainable-ai-knowledge-base-cs-education-2026]]); someone must maintain the system after publication ([[programming-its]] documents a small PhD-student team, work-in-progress security exposure, and a substantial compliance burden); and an open license is a permission, not a working product — the maintenance that keeps a released system usable falls to whichever [[educational-technology-developers]] built it, and their funding and incentives decide whether a forkable repository, a maintained product, or neither is what an institution inherits when the grant ends. ### Putting openness into practice - **For instructors and institutions:** check the license of the *content* as well as the code before adopting a system — MIT code over CC BY textbooks is reusable, but a permissive license does not guarantee that tutoring pathways, item banks, or translations exist. Prefer open models when student data legally cannot leave campus ([[lata-ferpa-compliant-local-llm-autograder]], [[programming-its]]), but benchmark the model on *your* task rather than trusting general leaderboards ([[cdpk-pedagogy-benchmark-llms]]). Budget for someone to run and evaluate the system after the pilot. - **For developers and researchers:** ship the whole thing — OATutor's end-to-end release (code, content, experiment harness) is the replicability standard this literature keeps rewarding. Release prompts and extraction schemas alongside weights ([[programming-its]], [[kar-mathbuddy-affective-math-tutoring-2025]]). Report compute and carbon ([[aied-carbon-footprint-reporting]]). Ground assistants in openly licensed corpora so provenance is checkable and adaptation is lawful ([[shen-sustainable-ai-knowledge-base-cs-education-2026]]). Use quantization-aware fine-tuning rather than plain quantization if accuracy and energy both matter, and keep code and content licenses compatible. ## Connected Concepts - [[intelligent-tutoring]] - [[pedagogical-llm-training]] - [[adaptive-learning]] - [[edtech-platform]] - [[privacy]] - [[regulation]] - [[governance]] - [[agentic-ai]] - [[automated-assessment]] - [[benchmark]] - [[knowledge-tracing]] - [[rag]] - [[pedagogical-safety]] - [[sustainability]] - [[research-methods-aied]] - [[writing-education]] - [[academic-integrity]] - [[educational-technology-developers]] ## Connected Articles - [[shen-sustainable-ai-knowledge-base-cs-education-2026]] — On-premise OER AI knowledge-base assistant on consumer hardware (Shen et al. 2026) - [[oatutor-open-source-adaptive-tutor-2023]] — MIT-licensed adaptive tutor with a CC BY OpenStax content library (Pardos et al. 2023) - [[singh-eduqwen-pedagogical-rl-2026]] — Open 32B pedagogical model outperforming far larger proprietary systems (Singh et al. 2026) - [[lata-ferpa-compliant-local-llm-autograder]] — Drop-in FERPA-compliant local-LLM autograder - [[programming-its]] — Self-hosted open-weight LLM for GDPR/EU AI Act compliance in a Python ITS - [[aiawe-automated-writing-evaluation]] — LoRA-adapted open-weight model for automated writing evaluation - [[stanbkt-bayesian-knowledge-tracing]] — Open Python package for full Bayesian knowledge tracing - [[deeptutor]] — Fully open-source agentic tutoring framework with learner memory - [[vismatic-secure-sandbox-cs-education]] — Open containerized sandbox for process-oriented assessment - [[kar-mathbuddy-affective-math-tutoring-2025]] — Affective math tutor with an open codebase - [[cdpk-pedagogy-benchmark-llms]] — Open pedagogy benchmark across 97 models and the cost–accuracy frontier - [[astra-multi-agent-tutoring-benchmark-2026]] — Open dataset and prototype for trace-based multi-agent tutoring evaluation - [[mllm-scientific-visualization-literacy]] — Open models below the human baseline on visualization literacy - [[iks-instruct-dataset-indian-knowledge]] — Open multilingual dataset for culturally grounded instruction - [[aied-carbon-footprint-reporting]] — Open-source method for reporting LLM environmental cost - [[egai-power-systems-education]] — Open executable module library for engineering AI - [[mechanical-engineering-ai-curriculum-2026]] — Publicly available curriculum, data, and code - [[mooc-to-maic]] — LLM-driven course generation and the changing economics of course production - [[agentic-ai-education-scoping-review]] - [[yu-academiclaw-student-challenges-ai-agents-2026]] - [[omniedu-open-educational-foundation-models-2026]] — OmniEdu: Open Foundation Models for Learning and Teaching --- ## [Edtech Platform](https://edtechdev.github.io/aied/concepts/edtech-platform/) > **Edtech Platform** — the digital systems, learning management systems (LMS), tutoring systems, and online learning environments through which AI is delivered to learners and educators. In AI in education, the platform is the *infrastructure layer* that determines whether an AI capability reaches students, how it is deployed (open vs. proprietary, integrated vs. standalone), and who can access, adapt, and evaluate it. Research in this knowledge base examines platforms from multiple angles: their design, their take-up and engagement constraints, their institutional governance, and their equity implications.([[access-not-enough-ai-tutoring-2026]])([[oatutor-open-source-adaptive-tutor-2023]]) ## Questions to Consider - Think of the last AI tutoring or feedback tool you encountered. Now think about where it actually 'lived' — the LMS, platform, or app that packaged it. Does that container feel like a neutral delivery vehicle, or could its design choices (open vs. proprietary, integrated vs. standalone, cloud vs. local) have changed what you could do with it? - One study found that nearly half of students never used a well-designed AI tutoring platform, and heavy users skewed toward higher-achieving students. If a tool is effective 'in principle' but students don't use it, is the capability or the platform the real problem? What would that mean for how you evaluate edtech? - Proprietary AI platforms can confine researchers to a handful of closed systems, while open platforms like OATutor let anyone fork, experiment, and publish. What might be lost — for research, equity, and institutional autonomy — when AI education is delivered through closed, opaque platforms? - How much does a platform's business model — who pays, who owns the data, what gets optimized — shape the learning that actually happens on it? Where would you look to see that influence? - Some new 'AI-native' platforms replace the one-video-for-many-students MOOC model with a multi-agent classroom built around each learner. Before reading on, what would you worry about losing when instruction becomes one-to-one with agents instead of one-to-many with instructors? ## Introduction The platform sits between an AI model or capability and the learner. It is the container that packages tutoring, assessment, feedback, and administration into something usable — and, critically, it shapes learning outcomes through its design choices, its accessibility, and its underlying business model. The concept spans learning management systems like Moodle, large-scale online platforms like MOOCs, dedicated [[intelligent-tutoring|intelligent tutoring]] systems, and emerging agentic or AI-native course platforms. Naming the container is not the same as naming its authors: the platform is the deployed system, while the stakeholder that decides what it does is [[educational-technology-developers]] — which matters here because the take-up, equity-skew and procurement findings below are usually consequences of design choices made before a platform ever reached a classroom. ## What a platform does in AI in education Platforms in AI in education perform several distinct functions: - **Deliver instruction and tutoring** — the container for [[intelligent-tutoring|AI Tutoring]] and [[intelligent-tutoring]] systems, from LMS-embedded tutors to standalone adaptive tutoring platforms. - **Manage the learning environment** — course organization, enrollment, progress tracking, and administration that traditional LMS platforms provide. - **Host assessment and feedback** — where [[automated-assessment]], [[formative-assessment]], and [[feedback|feedback loops]] run. - **Collect and analyze learning data** — the substrate for [[learning-analytics]] and [[student-modeling]]. - **Govern access and deployment** — decisions about [[open-source]] vs. proprietary, local vs. cloud, and which institutions and learners can use it. ## Key findings from the knowledge base's articles ### Take-up, not capability, is often the binding constraint A platform can be effective in principle yet fail in practice if learners do not use it. Two [[rct|RCTs]] of an [[ai-literacy|AI literacy]] (reading) tutoring platform found that **nearly half of control students never used the platform** and users averaged only 2–5 minutes per week — far below the dosage needed for reading gains. An in-person engagement tutor raised usage and engagement substantially but still did not produce achievement gains, and platform users skewed toward higher-achieving students, raising equity concerns.([[access-not-enough-ai-tutoring-2026]]) ### Which tools educators report using, and what gates access A rare census of educator-reported platform choice comes from a 2026 typology built from 211 educators across nine countries: the tools that reach classrooms are disproportionately the ones with a free tier, because a publicly available free version was an inclusion criterion, and the most-nominated entries are general-purpose assistants and media generators rather than purpose-built platforms. Roughly half of the fifty tools listed produce images, audio, video or slide decks, while document-grounded assistants (NotebookLM, Elicit, SciSpace, Humata, Research Rabbit) form the most coherent cluster in the research category. Dedicated [[intelligent-tutoring|tutoring]] systems appear as a small, subject-specific group rather than the center of reported use — which frames the take-up problem above in a wider setting, where a platform competes for attention against general-purpose tools that students and instructors already have open.([[typology-generative-ai-tools-education-2026]]) ### The platform model matters: open vs. proprietary - **Proprietary platforms** create barriers to research: researchers who want to replicate or extend [[adaptive-learning]] experiments are often confined to a small number of closed platforms. - **Open platforms** lower this barrier. **OATutor** is the first open-source adaptive tutoring system built on ITS principles — an MIT-licensed codebase with a Creative Commons algebra content library, [[knowledge-tracing]] mastery estimation, and built-in A/B testing — letting researchers fork, experiment, and publish the full end-to-end system.([[oatutor-open-source-adaptive-tutor-2023]]) ### AI-native platforms are reshaping online education The platform paradigm itself is evolving. **MAIC** (Massive AI-empowered Course) replaces the MOOC's "one video for N students" model with an LLM-driven multi-agent classroom — "N agents for 1 student" — using specialized Teacher, Assistant, Classmate, and Analyzer agents to deliver personalized, adaptive learning at scale, and reducing course production from ~\$25K/60 hours to under \$2/30 minutes.([[mooc-to-maic]]) Similarly, AI-integrated LMS designs propose moving beyond workflow-only platforms toward real-time instructional support with policy-gated (bounded) AI, formative hinting, spaced review, and teacher dashboards.([[ai-lms-middle-school-longitudinal]]) At the other end of the deployment spectrum, classroom-embedded AI must prove feasible in live physical environments. The Community Builder ([[breideband-community-builder-cobi-2026|CoBi]]) — a classroom-wide platform that uses speech recognition and language understanding to visualize small-group collaborative discourse — was successfully deployed across noisy middle-school classrooms using commodity microphones and a scalable cloud pipeline, showing that real-time speech-AI infrastructure can function in authentic K-12 settings even as interface mismatches (teacher vs. student views of whether feedback was group- or class-level) created deployment friction. ### Interest-based and context-aware platform features Platforms can personalize beyond performance data. **Taklif.AI** is an LLM-powered platform that generates college assignments based on students' **extracurricular interests and cultural contexts**, aligning with [[culturally-relevant-pedagogy]] and shifting from one-size-fits-all assignments toward interest-driven engagement.([[taklif-ai-interest-based-personalized-assignments]]) ## Implications for design and research 1. **Design for take-up, not just capability.** A platform's effectiveness depends on whether learners actually engage with it; support structures, onboarding, and scheduling matter as much as the AI itself.([[access-not-enough-ai-tutoring-2026]]) 2. **Treat platform structure as an equity lever.** Who benefits from a platform depends on access, infrastructure, and engagement constraints — platform design must be examined through an [[equity-in-ai-education]] lens.([[access-not-enough-ai-tutoring-2026]]) 3. **Prefer open, replicable platforms for research.** Open-source platforms like OATutor enable reproducible adaptive-learning research and a shared evidence base.([[oatutor-open-source-adaptive-tutor-2023]]) 4. **Design AI-native platforms with governance and bounds.** Privacy-first architecture, data minimization, auditable logs, and role-based access are critical as platforms become AI-integrated — connecting to [[privacy]] and [[governance]] concerns.([[ai-lms-middle-school-longitudinal]]) 5. **Explain recommendations in the teacher's domain language.** A platform's AI features earn trust and uptake when their explanations are understandable and pedagogically meaningful: in a within-subject experiment with an AI grouping-recommendation tool (GrouPer), [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor et al. (2025)]] found domain-driven explanations framed in curricular language increased teachers' understandability, trust, and acceptance significantly more than raw feature-importance explanations — and that real classroom use still mattered for full acceptance.([[xai-teachers-trust-edtech-recommendations-2026]]) ## Connected Concepts - [[personalized-learning]] - [[adaptive-learning]] - [[intelligent-tutoring]] - [[learning-analytics]] - [[student-modeling]] - [[knowledge-tracing]] - [[automated-assessment]] - [[formative-assessment]] - [[generative-ai]] - [[llm]] - [[open-source]] - [[ai-literacy]] - [[student-experience]] - [[teacher-role]] - [[k-12]] - [[higher-ed]] - [[equity-in-ai-education]] - [[privacy]] - [[governance]] - [[culturally-relevant-pedagogy]] - [[stem-education]] - [[educational-technology-developers]] ## Connected Articles - [[typology-generative-ai-tools-education-2026]] — What 211 educators reported using: 50 tools in nine categories - [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] — Ontology-based layered hybrid knowledge model for personalized e-learning - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — Virtual tutoring with CAL: an experiment in take-up and learning - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[access-not-enough-ai-tutoring-2026]] — Take-up and engagement are the binding constraints for AI tutoring platforms - [[oatutor-open-source-adaptive-tutor-2023]] — An open-source adaptive tutoring platform for replicable research - [[mooc-to-maic]] — Moving from MOOC to LLM-driven multi-agent AI classrooms - [[ai-lms-middle-school-longitudinal]] — AI-integrated LMS for middle school with bounded, privacy-first support - [[taklif-ai-interest-based-personalized-assignments]] — Interest-based personalized assignment platform - [[edusim-llm-robotic-simulation-education-2026]] — An LLM-robotic simulation platform for education - [[teachy-mini-generative-social-robot-higher-ed-2026]] — A generative social-robot teaching platform in higher education - [[hypergamification-game-engine-lms]] — A game-engine-based LMS integrating gamification - [[edtech-design-time-generative-ui]] — Designing edtech for generative UI - [[moodle-ai-tutoring-deep-learning]] — AI tutoring integrated into the Moodle LMS - [[lata-ferpa-compliant-local-llm-autograder]] — FERPA-compliant local LLM autograder platform - [[vismatic-secure-sandbox-cs-education]] — A secure sandbox platform for CS education - [[wordstream-glass-learning-analytics]] — A learning-analytics platform for streaming data - [[learnmate2-llm-adaptive-learning]] — LLM-powered personalized adaptive learning platform - [[multi-site-vr-immersive-learning]] — Multi-site VR immersive learning platform - [[privacy-aware-classroom-incident-recognition-2026]] — Privacy-aware computer vision in classroom platforms - [[a4l-analytics-pipeline]] — A configurable analytics pipeline platform - [[raza-farooq-aied-review-2020-2025]] — Comprehensive review of AIED research and systems - [[credentials-carry-evidence-ai-agents-2026]] — Credentials that carry their evidence for AI-agent work - [[breideband-community-builder-cobi-2026]] - [[xai-teachers-trust-edtech-recommendations-2026]] --- ## [AIEd in the Disciplines](https://edtechdev.github.io/aied/concepts/discipline-specific-aied/) > **AIEd in the Disciplines** — the application of artificial intelligence to teaching and learning within specific academic subjects, where each discipline's signature pedagogies, methods, theories, and concerns shape how AI is designed, used, and evaluated. Rather than treating [[ai-education|AI in education]] as a single generic phenomenon, this overview organizes the knowledge base's discipline-specific coverage and surfaces the cross-cutting themes that run through subject-area AIEd research. ## Questions to Consider - A math tutor that excels at right/wrong feedback may fail in the humanities, where interpretation and authorship matter more than correctness. How does your own discipline define what counts as 'good feedback'? - Research across 560 students found that disciplinary affiliation predicts how often and openly students use [[generative-ai|generative AI]]. Why might a discipline act as a whole 'activity system' shaping AI use — rather than AI being a neutral tool everywhere? - The 'cognitive act' being offloaded to AI is discipline-specific: computation in math, code in CS, composition in writing. If over-reliance looks different in each field, how would you detect it in yours? - What does AI mean for a discipline whose signature pedagogy is hands-on laboratory work, or interpretive meaning-making, rather than [[problem-solving]]? Is there any subject where AI should play little role? - Several disciplines — law, history, the arts — remain thin in AI-education research. Which discipline's AI use most deserves attention, and what would good AI in that field have to honor that generic tools don't? ## Introduction AI in education manifests differently across disciplines because each field has its own signature pedagogy — the distinctive ways knowledge is constructed, practiced, and taught. AI tutors that shine in [[math-education|mathematics]] may fail in [[humanities-education|the humanities]], where interpretation and authorship matter more than right answers. This page is the umbrella map for those discipline-specific strands, which now run from school subjects through professional and applied training to studio practice. One member of the strand is not a taught subject at all: [[learning-sciences]] is the research field around AI in education, supplying the mechanisms AI systems operationalize and the standards by which they are judged — it takes the cross-cutting view, against this page's premise that subject matter is what changes what support should do. ## Discipline-specific concepts The knowledge base has dedicated concept pages for several subject areas: - **[[math-education]]** — AI tutoring, adaptive problem-solving, and conceptual diagnosis in mathematics. - **[[physics-education]]** — AI simulation, chatbots, and problem-posing in physics learning. - **[[chemistry-education]]** — AI in laboratory/experimental design, AI-mediated formative assessment, context-based and inquiry-based instruction, [[llm]] technical limits on chemistry tasks, and the philosophy of experimentation. - **[[biology-education]]** — AI in laboratory instruction, AI literacy embedded in biology curricula, critical thinking in the AI era, and specialized tools (species identification, bioimaging, predictive modeling). - **[[cs-education]]** — AI for code generation, debugging, and novice programming support. - **[[information-technology]]** — the applied wing of computing, preparing practitioners to select, secure, administer, and govern the deployed systems organizations run: its unit of analysis is the deployed system and the practitioner's judgment about it, where computer science education takes the program and the algorithm as its objects. - **[[writing-education]]** — AI-assisted composition, automated essay scoring, and writing feedback. - **[[language-learning]]** — AI interlocutors, pronunciation feedback, and conversational practice in second/foreign languages. - **[[english-education]]** — English for Academic Purposes (EAP) and English language teaching (EFL/ESL/L2): academic-English register, genre-based writing, and English-specific feedback and assessment — distinct from general language learning and general writing. - **[[stem-education]]** — the cross-disciplinary umbrella for science, technology, engineering, and mathematics. - **[[teacher-education]]** — the preparation and [[educational-development|professional development]] of teachers (pre-service and in-service), a discipline in its own right whose AI research centers on [[teacher-role|teacher]] AI literacy, intelligent-[[tpack]], and readiness to integrate AI. - **[[medical-education]]** — clinical simulation, reinforcement-learning training, and foundational learning principles in health-professions education. - **[[nursing-education]]** — the densest empirical record in the clinical strand: simulation with virtual patients, LLM reasoning support, and adaptive platforms, framed around professional-identity formation as much as competence, and narrow enough to name the boundary where AI's documented benefits stop (complex psychomotor skill and emotionally loaded interaction) rather than treating health professions as one case. - **[[vocational-education]]** — initial preparation for named trades and technical occupations, organized by qualification frameworks and assessed on what learners can do with equipment: it extends beyond the professional training and workplace upskilling that [[professional-training]] covers to those not yet employed, and its recurring design question is which expensive-to-staff parts of practice AI can absorb without displacing the repetition that produces competence. - **[[design-education]]** — studio-based professional formation in product, service, interaction, interior and architectural design, where visible process is the assessed object and where generative tools make the polished artifact cheap: it narrows the studio disciplines that [[arts-design-and-media-education]] covers to the professional design pipeline, so that critique, portfolio and regulation rather than material and performative practice carry the argument. - **[[engineering-education]]** — professional formation, design and [[experiential-learning|hands-on learning]], [[ethics|ethical]] use of AI, [[embodied-learning|embodied]] assessment, and workforce preparation in engineering. - **[[business-education]]** — AI in business, economics, and management education: student-informed GenAI frameworks, [[curriculum-design|curriculum]] integration via constructive alignment, and preparation for AI-integrated professional practice. - **[[humanities-education]]** — interpretive cognition, authorship, and meaning-making in humanities and social sciences. - **[[k-12]]** and **[[higher-ed]]** — education about and with AI at each level. ## Cross-disciplinary themes Several threads cut across all disciplines, though they play out differently in each: - **Discipline as an activity system.** [[jiang-genai-activity-theory-disciplines-2026|Jiang et al. (2026)]] show, across 560 undergraduates in five academic domains, that disciplinary affiliation is significantly associated with students' GenAI **usage frequency and disclosure practices** — disciplines function as [[activity-theory-aied|activity systems]] whose norms, policies, and role expectations shape how students engage with and disclose GenAI. This is direct evidence for the knowledge base's premise that AI in education is not discipline-neutral. - **The shape of "interdisciplinary" itself is discipline-determined.** [[xia-ai-interdisciplinary-higher-education-review-2026|Xia et al.'s (2026)]] [[meta-analysis-systematic-review|PRISMA]] review of 59 studies catalogs six forms of interdisciplinary education in which AI features: STEM integration (n = 26), cohorts with cross-disciplinary student backgrounds (n = 14), non-STEM fields (n = 10), cross-disciplinary learning content (n = 10), inter-STEM integration (n = 8), and STEAM (n = 2). The distribution is a discipline-level finding, not a tool-level one: AI-supported interdisciplinarity is overwhelmingly STEM-hosted, and cohorts with cross-disciplinary backgrounds are studied more often than the non-STEM fields those cohorts are meant to bridge. The review further shows that the same AI functions recur across these settings — assistant (n = 43), evaluator (n = 22), agent (n = 14), monitor (n = 9) — while their cognitive, behavioral and affective effects differ by discipline and stakeholder, with students by far the most studied (n = 57), teachers next (n = 26) and administrative staff barely examined at all (n = 4). - **Tutoring and feedback.** AI tutoring systems ([[intelligent-tutoring|AI Tutoring]], [[intelligent-tutoring]], [[feedback]], [[ai-feedback-quality]]) appear in nearly every discipline, from [[math-education|math]] and [[physics-education|physics]] tutors to [[writing-education|writing]] and [[language-learning|language]] feedback. The discipline shapes what counts as good feedback — right/wrong in math, argument quality in writing, fluency in language. - **Assessment and evaluation.** [[automated-assessment]], [[automated-assessment|Automated Grading]], [[automated-essay-scoring]], and [[formative-assessment]] are reimagined by AI across disciplines, but the scoring constructs differ (procedural accuracy vs. interpretive depth vs. communicative competence). - **Cognitive offloading and over-reliance.** [[cognitive-offloading]] and [[cognitive-offloading|Over-Reliance]] risk appears across [[math-education|math]], [[cs-education|CS]], and [[writing-education|writing]], though the "cognitive act" being offloaded is discipline-specific — computation vs. code vs. composition. - **AI literacy and critical use.** [[ai-literacy]], [[critical-thinking]], and [[critical-pedagogy]] underpin responsible use in every subject. - **Equity and access.** [[equity-in-ai-education]], [[digital-divide]], and [[culturally-relevant-pedagogy]] concern all disciplines. ## Signature pedagogies, methods, and theories by discipline Each discipline brings distinctive [[pedagogy|pedagogical]] traditions that AI research engages: - **Physics** — model-based reasoning, experimentation, and [[simulation]]. AI research uses virtual labs ([[benzion-ai-physics-simulations-virtual-lab|Ben-Zion]]), chatbots ([[becker-chatgpt-typology-physics-2026|ChatGPT typology]]), and problem-posing ([[genai-assisted-problem-posing-physics-2026|GenAI problem-posing]]). - **Chemistry** — abstract submicroscopic concepts, specialized notation, and laboratory practice. AI research spans AI-supported experimental design ([[ai-supported-experimental-design-chemistry-2026|AI-designed lab manuals]]), context-based 7E instruction with AI tutoring ([[context-based-ai-secondary-chemistry-2026|context-based + AI]]), AI-mediated formative assessment ([[instructor-ai-roles-chatgpt-formative-assessment-2026|instructor–AI roles]]), and the philosophy of experimentation ([[philosophy-experimentation-ai-chemistry-2026|philosophy of experimentation]]). - **Biology** — specialized terminology, systems/visual-spatial thinking, and hands-on laboratory and fieldwork. AI research spans virtual lab teaching assistants ([[chatgpt-virtual-lab-teaching-assistant-biology-2026|ChatGPT as VTA]]), AI literacy embedded in biology curricula ([[zha-ai-literacy-biology-case-study|AI literacy in a biology class]]), ChatGPT in challenge-based learning ([[chatgpt-math-biology-challenge-based-learning-2025|ChatGPT in CBL]]), critical thinking in the AI era ([[critical-thinking-biological-sciences-ai-2025|critical thinking in bio sciences]]), and the AI-tools landscape ([[beyond-chatgpt-ai-tools-biological-education-2026|AI tools review]]). - **Writing** — process-oriented, recursive drafting and revision. AI research spans [[coach-not-crutch-ai-writing|AI as writing coach]], automated scoring, and stage-based ownership ([[ai-writing-support-stage-ownership-2026|ownership stages]]). - **English education (EAP/EFL/ESL)** — English as target language and academic register. AI research spans ethical EAP integration ([[alharbi-ethical-genai-eap-2026|Alharbi et al.]]), GenAI EAP writing revision ([[feedback-literacy-scripts-eap-writing|feedback-literacy scripts]]), EAP reading-material adaptation ([[genai-differentiated-eap-reading-materials-2026|differentiated EAP materials]]), ESL tutoring ([[tact-pedagogically-adaptive-esl-tutoring|TACT]]), and English-specific assessment ([[self-referential-l2-writing-llm-assessment|self-referential L2 writing evaluation]], [[ai-vs-human-assessment-efl-tpck-2026|EFL assessment]]). - **Teacher education** — a discipline whose AI research centers on preparing teachers to integrate AI: intelligent-TPACK frameworks ([[designing-ai-professional-development-itpack-2026|i-TPACK PD]]), teacher AI literacy ([[science-educators-ai-literacy-postqualification-2026|science educators]]), pre-service readiness ([[conceptualizing-preservice-teachers-ai-readiness-2026|intelligent-TPACK readiness]]), and in-service trust and ethics ([[intelligent-tpack-ethics-teachers-trust-distrust-2026|trust and ethics]]). - **Humanities & social sciences** — interpretation, authorship, and critical meaning-making. AI research foregrounds interpretive cognition ([[voicu-ai-interpretive-cognition-ssh-2026|Voicu]]) and the [[philosophy-of-ai-in-education|philosophy of AI]] rather than tutoring for correctness. ## Represented disciplines in the knowledge base The knowledge base's strongest discipline-specific coverage is in **[[stem-education|STEM]]** broadly — particularly **[[math-education]]**, **[[physics-education]]**, **[[chemistry-education]]**, **[[biology-education]]**, and **[[cs-education]]** — followed by **[[writing-education]]**, **[[language-learning]]** (with a distinct **[[english-education]]** strand for EAP/EFL/ESL), and more recently **[[engineering-education]]**, **[[teacher-education]]** (with a substantial body of pre-service and in-service AI-training research), **[[medical-education]]**, and **[[humanities-education]]**. Engineering and design also have a growing body of articles. This concentration tracks the wider literature: [[xia-ai-interdisciplinary-higher-education-review-2026|Xia et al.'s (2026)]] review found STEM the most common interdisciplinary form in AI-supported higher education (n = 26), well ahead of non-STEM fields (n = 10) and STEAM (n = 2) — a discipline-level caution that AI-in-education evidence accumulates fastest where computational tooling is easiest to embed, not necessarily where AI's pedagogical value is greatest. The most recent additions broaden the strand past academic subjects into professional and applied education — [[nursing-education|nursing]], [[information-technology|information technology]], and [[vocational-education|vocational education and training]] — where learners are assessed on demonstrated practice rather than on correctness, and where the employer and the licensing body, not only the academy, define what counts as competence. ## Underrepresented disciplines Several disciplines remain thin in the knowledge base and are good candidates for future ingestion: - **Law and legal education** — minimal coverage: [[llm-turing-test-italian-legal-exams-2026|LLMs and Italian legal exams]]. - **Psychology and counseling** — few AI-in-education articles: [[hawkins-feedback-literacy-ai-essay-writing|AI feedback literacy]], [[critical-genai-use-predictors|critical GenAI use]], [[adaptive-virtual-patient-psychotherapy-training|virtual patient psychotherapy training]]. - **History** — only isolated articles: [[paternalistic-filter-llm-history-education|LLMs and historical reasoning]]. - **The arts (visual art, design, music)** — emerging coverage: [[ai-interior-design-malaysia-2026|AI in interior design education]], [[genai-architectural-design-studios|GenAI in architectural design studios]], [[ai-vocal-pedagogy-2026|AI vocal pedagogy]], [[musical-education-ai-digital-transformation-2026|AI in music education]], [[t2i-competence-paradox-2026|text-to-image competence paradox]]. The design-side studies in this list now have their own home: [[design-education]] draws the architectural, interior-design and text-to-image studio work into a page about professional design formation, leaving visual art, music and performance as the thin remainder. These underrepresented disciplines would benefit from dedicated concept pages and additional article ingestion as the knowledge base grows. ## Connected Concepts - [[business-education]] - [[ai-education]] - [[math-education]] - [[physics-education]] - [[chemistry-education]] - [[biology-education]] - [[cs-education]] - [[writing-education]] - [[language-learning]] - [[english-education]] - [[stem-education]] - [[teacher-education]] - [[medical-education]] - [[engineering-education]] - [[humanities-education]] - [[k-12]] - [[higher-ed]] - [[equity-in-ai-education]] - [[arts-design-and-media-education]] - [[design-education]] - [[information-technology]] - [[learning-sciences]] - [[nursing-education]] - [[vocational-education]] ## Connected Articles - [[xia-ai-interdisciplinary-higher-education-review-2026]] — Systematic review of AI in interdisciplinary higher education (59 studies) - [[llms-text-linguistics-teaching-2026]] — LLMs in text linguistics teaching - [[fowlin-operationalizing-learning-principles-ai]] — Operationalizing learning principles with AI in health-professions education - [[designing-ai-professional-development-itpack-2026]] — Intelligent-TPACK-based AI professional development - [[teaching-the-teachers-genai-tpk-review-2026]] — GenAI-specific TPK in teacher education - [[human-centered-ai-teacher-educators-2026]] — Critical AI literacy professional learning for teacher educators - [[voicu-ai-interpretive-cognition-ssh-2026]] — Interpretive cognition in humanities and social science education - [[becker-chatgpt-typology-physics-2026]] — ChatGPT use typology in physics education - [[code-review-genai-cs1]] — GenAI code review in CS1 - [[ai-writing-support-stage-ownership-2026]] — Stage-based AI writing support and ownership - [[ai-engineering-education-balancing-act]] — The balancing act of AI in engineering education - [[ai-literacy-career-adaptability-business-2026]] — AI literacy and career adaptability in business education - [[jiang-genai-activity-theory-disciplines-2026]] — Activity theory: disciplinary differences in GenAI use and disclosure (560 students) - [[jiang-ai-powered-simulation-nursing-education-2026]] — AI-powered simulation in nursing education: gains in knowledge and confidence, inconsistent effects on psychomotor skill - [[atif-dickson-deane-scaffold-shortcut-genai-srl-2026]] — GenAI as scaffold or shortcut in postgraduate IT learning - [[ai-vocational-education-training-review]] — First systematic review of AI in vocational education and training - [[genai-architectural-design-studios]] — Generative models in the design studio: stimulus, solution-space expansion, and fixation - [[deceptive-overgeneralization-adaptive-learning-2026]] — Mastery stopping rules that certify an overgeneralised rule as competence --- ## [Math Education](https://edtechdev.github.io/aied/concepts/math-education/) > **Math Education** — the study of how students learn mathematics and how AI can support mathematics teaching, spanning affective tutoring, cognitive diagnosis from handwritten work, [[desirable-difficulties|productive struggle]] evaluation, help-seeking behavior, teacher-AI collaboration for visual generation, and [[student-ai-interaction|student-AI interaction]] trajectories. Math education is the most active [[discipline-specific-aied|domain-specific]] [[research-methods-aied|research]] area in this knowledge base, with 10 articles that collectively explore how AI can support — and sometimes undermine — mathematical learning from elementary fractions through higher education. ## Questions to Consider - Math problems have clear right answers yet require rich reasoning, which is why math is a favored testbed for AI tutoring. When you're stuck on a math problem, what kind of help actually helps you learn — an answer, a hint, or a question — and which is the AI likely to default to? - Research finds AI tutors often default to over-helpfulness, rarely pushing for rigor even when students are ready. If you're designing a tutor, how do you decide when to withhold help to preserve the 'productive struggle' that builds understanding? - The page shows students who request hints too early or skim them superficially tend to learn less. Have you ever reached for a hint out of impatience rather than genuine effort? What does that reveal about how AI support can undermine rather than support learning? - AI cognitive-diagnosis systems sometimes hallucinate evidence and over-attribute mistakes, and even strong models underperform when reading students' actual handwritten work. How confident would you be in a tutor that diagnoses what you got wrong from your [[cs-education|scratch]] work? - LLMs flip their answers across mathematically equivalent problem formulations — the same problem presented differently changes the result. What does this say about using AI to score or diagnose math understanding? ## Introduction Mathematics education has become a primary domain for [[ai-education|AI in education]] research because math problems have clear right answers yet require rich reasoning — making them ideal for studying tutoring effectiveness, assessment validity, and how AI tools interact with student cognition and affect. The articles in this knowledge base reveal both the promise of AI math tutors and persistent challenges: over-scaffolding that undermines productive struggle, hallucination in cognitive diagnosis, and the difficulty of balancing AI assistance with genuine learning. ### Key research themes **AI math tutoring and scaffolding** is the largest cluster, with four articles examining how AI tutors support or undermine math learning. **[[kar-mathbuddy-affective-math-tutoring-2025|MathBuddy]]** demonstrates that adding affective awareness — detecting student emotions from text and facial expressions — produces a +23-point win rate advantage in math tutoring, connecting to [[affective-computing]] and [[affective-tutoring]]. **[[zhang-tutormoments-2026|TutorMoments]]** evaluates 462 teacher-annotated transcripts from grades 2-7 math tutoring and finds frontier models default toward over-helpfulness, rarely pushing for rigor even when students are ready — directly challenging the alignment between AI helpfulness and [[scaffolding]] principles. **[[lak2026-hint-button-unproductive-use|An et al.]]** analyzed 999 students across three semesters in the *Decimal Point* ITS, finding that premature hint requests and superficial hint reading consistently predict reduced [[learning-gains|learning gains]], even after controlling for [[prior-knowledge|prior knowledge]] — a finding that connects to [[help-seeking]] and [[learning-analytics]]. **[[cognitive-diagnosis|Cognitive diagnosis]] and assessment** explores AI's ability to evaluate math thinking. [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] add a large-scale item-difficulty study spanning both math and reading: across 5,170 K-5 items calibrated under the Rasch IRT model, GPT-4o's zero-shot difficulty ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but were uneven across grades, while a feature-based approach (LLM-extracted features into tree-based models) reached correlations up to r = 0.87, with grade level and word count the top predictors. The study offers a practical seven-step workflow for testing professionals and cautions that generalizability beyond K-5 math and reading is unclear. **[[llm-cognitive-diagnosis-handwritten-math|MathCog]]** benchmarked 18 LLMs on 3,036 teacher-annotated diagnostic verdicts from handwritten math work, finding all models severely underperform (F1 < 0.5) with systematic over-attribution and hallucination of evidence — connecting to [[knowledge-tracing]], [[hallucination-risk]], and [[multimodal]] assessment challenges. **[[representation-robustness-llm-math-problem-solving|Nath et al.]]** showed that [[llm]] math [[problem-solving]] is highly sensitive to surface representation — models flip correctness across equivalent problem formulations — raising [[assessment-validity]] concerns for AI-based math scoring. **[[student-engagement|Student engagement]] and AI literacy** examines how students interact with AI math tools. **[[epistemic-proactivity-math|Abdelghani et al.]]** traced temporal trajectories of student-AI interaction in math learning, identifying a developmental path from superficial [[prompt-engineering|prompting]] to "epistemic proactivity" — active, [[self-directed-learning|self-directed]] pursuit of conceptual understanding. This connects to [[ai-literacy]], [[metacognition]], and [[self-regulated-learning]]. **[[ai-powered-personalized-learning-elementary-fractions-2026|Holman]]** found that AI-adaptive platforms significantly improved fraction comprehension for students with math learning difficulties, connecting to [[personalized-learning]] and [[adaptive-learning]]. **Teacher support** explores AI tools for math educators. **Simulated-student role-play** also serves teacher practice: [[zhuang-zhang-chatgpt-math-teacher-education-2026|Zhuang and Zhang (2025)]] built *Student GPT*, a custom ChatGPT [[conversational-ai|chatbot]] that role-played a middle school student holding common ratio-reasoning [[misconceptions]], giving preservice secondary math teachers low-risk practice at diagnosing and guiding student thinking toward correct solutions — illustrating [[generative-ai|GenAI]]-powered [[simulation]] as a complement to costly platforms like TeachLivE for building pedagogical content knowledge about student misconceptions. **Higher education math** explores AI's impact on advanced math practice. **[[genai-runaway-object-math-higher-ed|Bui et al.]]** applied [[sociocultural-learning|socio-cultural]] theory to [[generative-ai|GenAI]] in university mathematics, analyzing AI as a "runaway object" that transforms academic practice in ways that outpace [[governance|institutional]] and pedagogical norms. **LLM tutoring and [[learning-design|instructional design]]** is an emerging cluster of two 2026 studies that sharpen the math-education evidence base. Looi, Liu, and Sun (2026) developed a rule-guided [[intelligent-tutoring|LLM tutoring system]] for primary-school math word problems whose three-layer architecture (diagnosis → intent selection → constrained response generation) improved interactional consistency and reduced premature answer-giving in a 40-student Grade 5 classroom pilot — evidence that procedural math domains need [[guardrails|structured rule-guards]] on otherwise stochastic LLM scaffolding. Zhu, Liang, Mao, and Wang (2026) applied a smart-classroom model to mathematics M.Ed. students and found statistically significant gains (p < .05) in instructional-objective design across curriculum-standards, textbook, and student-condition dimensions. **[[generative-ai|GenAI]] for mathematical modeling tasks** extends the generation strand beyond routine exercises. An AI-powered platform developed through the ADDIE approach used direct variation in secondary school mathematics as an illustrative topic, addressing teachers' lack of time and resources to design high-quality modeling tasks: existing tools typically produce conventional word problems or routine exercises, whereas the platform aimed to generate resources that foster mathematical modeling competencies, grounded in established design principles and [[prompt-engineering|retrieval-augmented generation]]. - **Visual chain of thought: the [[agency|autonomy]] gap in geometry.** GeoVAD-Bench diagnoses intermediate visual aids rather than final answers across 600 auxiliary-construction problems (200 easy, 200 medium, 200 hard), and finds a consistent pattern: supplying the reference auxiliary diagram improves accuracy modestly (+3.3, +3.0, +7.0 points across three models) while leaving the model to construct its own auxiliary line on the way to the correct answer widens the gap by 10.0 to 13.5 points, with two models performing worse than when they had no visual reasoning at all. Four process-error categories accounted for 93.1% and 89.7% of attributed failures. For [[problem-solving]] instruction the finding is that diagrammatic scaffolding has to be trained and evaluated separately from answer accuracy. ([[geovad-bench-visual-chain-of-thought-geometry-2026]]) ### Connections to related concepts Math education sits within the broader [[stem-education]] domain with distinctive connections to [[intelligent-tutoring]] and [[intelligent-tutoring|AI Tutoring]] through the strong tradition of cognitive tutors and ITS research in mathematics, to [[scaffolding]] through the productive struggle and hint-use literature, to [[affective-computing]] through math anxiety and emotion-aware tutoring, to [[knowledge-tracing]] and [[assessment-validity]] through cognitive diagnosis and assessment research, and to [[teacher-role]] through teacher-AI collaboration in math instruction. The [[k-12]] connection is particularly strong — 8 of 10 math articles involve K-12 contexts — while [[higher-ed]] connections emerge in teacher preparation and advanced math practice. ## Implications for math instructors - **Treat AI tutoring as a help-seeking lever, not a capability fix.** [[lak2026-hint-button-unproductive-use|Hint-use research]] shows premature hint requests and superficial hint reading predict lower gains — so the design of *when and how* students seek AI help matters more than raw tutor capability. Encourage students to attempt before asking, and surface help at the moment of need rather than on demand. - **Protect productive struggle.** [[zhang-tutormoments-2026|TutorMoments]] finds models default to over-helpfulness, rarely pushing for rigor; configure AI support to scaffold rather than solve, and monitor for answer-replacement that erodes reasoning. - **Do not treat AI diagnostic output as ground truth.** [[llm-cognitive-diagnosis-handwritten-math|MathCog]] shows LLMs underperform at diagnosing math thinking (F1 < 0.5) with over-attribution and hallucinated evidence; use AI diagnosis as a suggestion to verify against the student's actual work. - **Beware surface-format fragility in AI scoring.** [[representation-robustness-llm-math-problem-solving|Representation sensitivity]] means equivalent problems can flip AI answers — a validity risk for AI-based math assessment; prefer [[human-in-the-loop-ai|human review]] for high-stakes scoring. - **Use AI to lower the bar for personalized practice.** [[ai-powered-personalized-learning-elementary-fractions-2026|Adaptive platforms]] improved fraction comprehension for students with math learning difficulties; deploy AI-adaptive tools selectively for learners who need differentiated support. - **Keep the teacher in control of AI-generated instructional materials.** [[teacher-control-ai-generation-math-visuals|Teacher control of AI visuals]] supports a framework that balances AI efficiency with pedagogical correctness. ## Connected Concepts - [[stem-education]] - [[intelligent-tutoring]] - [[scaffolding]] - [[affective-computing]] - [[affective-tutoring]] - [[k-12]] - [[higher-ed]] - [[ai-literacy]] - [[metacognition]] - [[self-regulated-learning]] - [[personalized-learning]] - [[adaptive-learning]] - [[help-seeking]] - [[learning-analytics]] - [[knowledge-tracing]] - [[assessment-validity]] - [[multimodal]] - [[hallucination-risk]] - [[cognitive-offloading]] - [[teacher-role]] - [[educational-development]] - [[generative-ai]] - [[discipline-specific-aied]] - [[teacher-education]] ## Connected Articles - [[mindful-llm-math-tutoring-2026]] — Beyond Problem Solving: Large Language Models for Emotional and Reflective Support in Mathematics Learning - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — Virtual tutoring with CAL: an experiment in take-up and learning - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[chudziak-ai-math-tutoring-platform]] — AI-powered math tutoring platform (Chudziak & Kostka 2025) - [[drawedumath-vlm-struggling-students-2026]] — VLMs underperform on math student work with errors (DrawEduMath, Lucy et al. 2026) - [[kar-mathbuddy-affective-math-tutoring-2025]] - [[zhang-tutormoments-2026]] - [[lak2026-hint-button-unproductive-use]] - [[llm-cognitive-diagnosis-handwritten-math]] - [[representation-robustness-llm-math-problem-solving]] - [[epistemic-proactivity-math]] - [[ai-powered-personalized-learning-elementary-fractions-2026]] - [[teacher-control-ai-generation-math-visuals]] - [[ai-tpack-preservice-math-teachers]] - [[genai-runaway-object-math-higher-ed]] - [[generative-ai-reduced-study-time-math]] — ALEKS mastery platform: text-based problems most AI-susceptible - [[diagramir-educational-math-diagram-evaluation]] — DiagramIR: automatic pipeline for educational math diagram evaluation - [[mujib-ai-ibl-creative-math-2026]] — AI-supported IBL and creative mathematical performance - [[puech-pedagogical-steering-llm-productive-failure-2025]] — Pedagogical Steering of LLMs for Productive Failure - [[rhaimi-productivemath-2025]] — ProductiveMath: AI to Support Productive Failure Problem Design - [[preferred-scaffolding-ai-mathematical-modeling]] — Preferred scaffolding in AI-supported mathematical modeling - [[instructional-design-proficiency-masters-math-2026]] — Smart-classroom model and D-T-E loop improving M.Ed. instructional design proficiency in mathematics (Zhu et al. 2026) - [[rule-integrated-llm-tutoring-primary-math-2026]] — Rule-guided vs ad-hoc scaffolding in an LLM tutoring system for primary mathematics (Looi et al. 2026) - [[ai-modeling-problem-generation-platform-2026]] — AI-powered platform generating mathematical modeling problems (ADDIE, RAG) - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[zhuang-zhang-chatgpt-math-teacher-education-2026]] - [[gpt4-handwritten-math-exam-grading-2026]] — GPT-4 grading of semi-open handwritten university mathematics answers - [[exrec-exercise-recommendation-knowledge-tracing-2025]] — semantic knowledge-concept annotation and RL exercise sequencing on K-12 math corpora - [[misconception-acquisition-dynamics-llms-2026]] — algebra mal-rule training dynamics in language models --- ## [Physics Education](https://edtechdev.github.io/aied/concepts/physics-education/) > **Physics Education** — the study of how students learn physics and how to teach it more effectively, spanning Socratic [[intelligent-tutoring|AI tutoring]], [[computational-thinking|computational thinking]] assessment, student [[trust]] and AI adoption patterns, automated scoring validity, and teacher preparation. The physics education articles in this knowledge base are notable for their domain-specificity: they explore how AI tools interact with the unique cognitive demands of physics reasoning — visual-spatial thinking, mathematical modeling, abstract systems thinking, and multi-step [[problem-solving]]. ## Questions to Consider - Physics problems often require visual-spatial thinking, mathematical modeling, and multi-step reasoning. Why might these be precisely the cognitive demands that current AI tutors struggle with? - Students report a big trust-utility gap—91% use AI for coursework but only 41% trust it. Have you experienced using a tool you didn't fully trust? What drove the gap? - AI scoring systematically underestimated linguistically weak students' physics explanations. What does that suggest about how an AI grades an explanation versus a correct numeric answer? - One study found a Socratic AI chatbot dramatically improved question specificity in a live physics course, but students also frequently 'ceded strategic control' to the tutor. When does handing over strategic control help learning, and when does it harm it? - Why might physics be a 'proving ground' for [[ai-education|AI in education]]—what makes its problems ideal for studying how AI affects reasoning and assessment? - Would you trust an AI to grade your physics problem set or reason through a force diagram with you? What would need to be true about the AI—and about your course—for you to say yes? ## Introduction Physics education [[research-methods-aied|research]] has become a proving ground for AI in education because physics problems are well-structured yet cognitively demanding, making them ideal for studying how AI tools affect learning, reasoning, and assessment. The seven articles in this knowledge base collectively paint a picture of a field grappling with both the promise and the limits of AI — from Socratic [[conversational-ai|chatbots]] that improve student question quality to systematic scoring biases that penalize linguistically diverse learners. ### Key research themes **Socratic AI tutoring in physics** is the most developed theme, with three articles deploying [[llm]]-powered Socratic dialogue in real physics courses. **[[hashmi-socratic-physics-chatbot-2025|Hashmi et al.]]** demonstrated that sustained Socratic interaction with an AI chatbot dramatically improves question specificity in introductory mechanics, with 150 STEM majors in a live course. **[[socratic-ai-physics-tutor-taxonomy-2026|Hashmi & Rebello]]** built a bottom-up taxonomy of 357 student discourse categories from the same deployment, revealing that meta-procedural turns — where students cede strategic control to the tutor — dominate student interactions. Both contribute to broader [[socratic-method]] research and connect to [[intelligent-tutoring|AI Tutoring]] and [[intelligent-tutoring]] frameworks. **Student AI adoption and trust** explores how physics students actually use AI tools. **[[fouad-bentley-trust-utility-gap-physics-2026|Fouad & Bentley]]** found a 50-point trust-utility gap: 91% use AI for coursework but only 41% trust it, with students spontaneously identifying AI failure modes in visual-spatial reasoning and circuits. **[[becker-chatgpt-typology-physics-2026|Becker et al.]]** developed a two-profile typology — 70% "Pragmatic Users" and 30% "Skeptical Non-Users" — from 1,189 survey responses, showing both groups make calculated risk-utility trade-offs. These studies advance [[ai-literacy]] and [[trust-calibration]] research, and challenge one-size-fits-all [[educational-policy-ai|AI policies]]. **Perception change without behavior change** is what [[physics-students-llm-perceptions-instruction-2026|O'Brien et al. (2026)]] add to this picture. A reflective lesson on how LLMs work, taught in a required first-year course for physics majors, raised skepticism sharply — agreement that LLMs can leave students with a false sense of confidence rose from 58% to 88%, and the belief that an LLM outperforms the average physics student fell from 54% to 32% — while convenience (71% agreement) and deadline pressure (65%) remained the dominant reasons for use. The lesson's limits are as informative as its effect: [[ai-literacy]] instruction shifted what students said about these tools and not the pressures that make them reach for one, which is why an intervention of this kind belongs alongside problem and policy design rather than in place of it. **Assessment and computational thinking** examines how AI can evaluate physics learning. **[[llm-computational-thinking-physics-2026|Savage et al.]]** used LLMs to assess [[computational-thinking|computational thinking]] growth in introductory physics, finding LLMs can scale CT assessment but struggle with complex constructs like Systems Thinking. **[[ai-scoring-language-bias-physics|Feser & Tschisgale]]** demonstrated that AI scoring systematically underestimates linguistically weak students' physics explanations — a finding that connects to [[assessment-validity]], [[bias-mitigation]], and [[equity-in-ai-education]]. **Psychometric infrastructure for diagnostic assessment.** [[mechanics-cognitive-diagnostic-physics-2026|Le et al. (2026)]] invert the usual AI-in-physics question: rather than asking whether a model can solve or grade physics, they ask whether the field's own research-based [[assessment|assessments]] can be made to diagnose it. Mapping FCI, FMCE and EMCS items onto 14 fine-grained learning objectives and fitting a DINA cognitive-diagnostic model to 24,394 posttest responses from 807 courses at 79 institutions via the LASSO platform, they built the Mechanics [[cognitive-diagnosis|Cognitive Diagnostic]] — reported as the first cognitive diagnostic computerized adaptive test in physics. The FCI and EMCS fit well (RMSEA2 = 0.033 and 0.022) while the FMCE fit only marginally (0.065), and the authors trace that misfit to instrument design rather than modeling: 42 of 43 scored FMCE items share scenario stems in chained sets, creating the local item dependence that DINA's conditional-independence assumption forbids (the FCI blocks 13 of 30 items; the EMCS none). Classification accuracy met the low-stakes [[formative-assessment|formative]] benchmark for 19 of 22 objective–assessment combinations; the three failures were EMCS energy objectives whose items overlap by roughly 70 percent, so mastery of one cannot be separated from the others. The significance for physics teaching is that instruments courses already administer can be repurposed to deliver actionable, objective-level feedback *during* instruction rather than a retrospective posttest score — provided the diagnostic claims are pitched at the resolution the item bank can actually support. **Benchmarking multimodal AI on authentic physics problems.** [[omniphys-multimodal-physics-benchmark-2026|Chen et al. (2026)]] introduce **OmniPhys**, a large-scale [[multimodal]] [[benchmark]] (15,246 questions, 19,850 images) spanning middle-school through university-level physics from Chinese educational corpora. Unusually, it evaluates not just multimodal *input* comprehension but multimodal *output* generation — whether models can synthesize structured physics diagrams, a core component of authentic problem solving. Extensive evaluations reveal critical gaps in current multimodal LLMs, especially in complex reasoning and visual generation. **Instructional-design frameworks for AI-augmented instruction.** **[[airis-cognitively-activated-ai-physics-2026|Kuhn et al.]]** propose the **AIRIS** framework (Activate–Inquire–Reflect with Intelligent Support) — a three-phase structure for cognitively activated AI use in physics: students predict and sketch expected outcomes before AI (Activate), delegate computational and representational steps to AI while critically comparing output to their own predictions (Inquire), and interpret, check consistency across representations, and reflect on what the AI contributed afterward (Reflect). Grounded in [[self-regulated-learning]], [[cognitive-offloading|Cognitive Load]] Theory, multiple external representations, and [[human-ai-collaboration]], it frames the central challenge as [[learning-design|instructional design]] rather than cheating or tool choice, and calls for "withdrawal condition" experiments testing whether learning survives the removal of AI support. **Generative video as synthetic experimental data.** [[genai-video-engineering-physics-workflow-2026|Alvarado-Cruz et al. (2026)]] generate video scenarios with PixVerse, Grok Imagine and Pippit for three resistive-force regimes — constant friction, linear drag and quadratic drag — extract the kinematics with the [[open-source]] Tracker tool, and fit the analytical models by non-linear least squares. The synthetic data agreed with the classical equations of motion and recovered physically meaningful parameters, and the recurring practical finding is that prompt specificity governs physical coherence: more detailed descriptions produced more coherent dynamics. The workflow mirrors experimental practice from model construction to [[quantitative-research|quantitative]] validation, and reframes [[prompt-engineering|prompt formulation]] as a stage of experimental design rather than a convenience. What it does not yet demonstrate is learning: validation here is agreement between generated motion and the authors' models, not students' measurement judgment, so the approach inherits the [[assessment-validity|validity]] question any generated data used as evidence must answer. **Assisted performance vs. unaided knowledge in a redesigned course.** A 2026 redesign of the introductory nuclear and particle physics course at Ruhr University Bochum (Mikhasenko et al.) allowed [[generative-ai|generative AI]] on ten deliberately AI-resistant, research-shaped homework sheets designed so that naive [[prompt-engineering|prompting]] would not suffice. [[student-engagement|Engagement]] and ambition were high — 24 of 42 students earned credit on all ten sheets, and one derivation filled more than two meters of blackboard — but an unaided 90-minute written exam was a "serious warning": a mean of 20.6/80, with only two of 27 examinees reaching 40. The authors conclude that assisted performance and independently retrievable knowledge are distinct achievements that cannot be assumed to train or demonstrate each other, and that physics courses must reserve some practice for unaided work — reinforcing the knowledge base's broader [[transfer-of-learning|transfer]] evidence. **Agent role design as an instructional variable.** [[wang-teacher-student-centered-agents-physics-2026|Wang et al. (2026)]] hold the model (DeepSeek R1), platform, and temperature constant and vary only the prompt-specified role: a teacher-centered agent answering authoritatively from a bounded textbook knowledge source, versus a student-centered agent configured as an empathic teacher with knowledge of students' understanding, scripted to diagnose the cause of [[misconceptions]], name the relevant concept, and transfer to an analogous phenomenon. Across 59 high-school graduates working two conceptual items, the student-centered agent produced higher post-test scores (9.67 vs. 7.93; r = 0.38), lower extraneous and higher germane cognitive load, stronger flow experience (d = 0.92), and higher empathy perception (r = 0.53) — evidence that role framing, not just answer accuracy, is what makes a physics agent instructionally effective ([[pedagogical-agent]], [[prompt-engineering]]). - **Benchmark scores understate what models can already do in physics.** Re-grading six widely used physics benchmarks with domain experts found that most of the reported shortfall was an artifact of defective items and restrictive automated graders: of 250 audited rejections, 143 (57.20%) were benchmark defects and 95 (38.00%) grader errors, with only 12 (4.80%) genuine model errors. Corrected, HLE-Physics mean@4 rose from 47.28% to 78.66% and CritPt from 32.29% to 87.50%. For physics instruction this cuts both ways: it means students can already obtain expert-level text solutions to many canonical problems, so assessment of physics reasoning needs to move toward items that resist benchmark contamination and toward process evidence rather than final answers. ([[frontier-models-physics-benchmark-audit-2026]]) ### Connections to related concepts Physics education sits within the broader [[stem-education]] domain but has distinctive connections: to [[socratic-method]] through the strong tradition of Socratic dialogue in physics problem-solving; to [[computational-thinking]] through the increasing role of computation in physics; to [[assessment-validity]] through the challenges of scoring physics explanations; and to [[professional-training]] through [[simulation]]-based preparation. The [[student-experience]] and [[ai-literacy]] concepts are essential for understanding how physics students navigate AI tools, while [[educational-measurement]] and [[automated-assessment|Automated Grading]] connect to the assessment dimension. ## Implications for physics instructors - **Design for student trust, not just adoption.** [[fouad-bentley-trust-utility-gap-physics-2026|Fouad & Bentley]] document a 50-point trust-utility gap (91% use, 41% trust), with students identifying AI failures in visual-spatial reasoning and circuits — create opportunities to expose and discuss these limits rather than assume acceptance. - **Use Socratic AI to deepen question quality, but watch for strategic ceding.** [[hashmi-socratic-physics-chatbot-2025|Socratic chatbots]] improve question specificity, yet [[socratic-ai-physics-tutor-taxonomy-2026|taxonomy research]] finds meta-procedural turns dominate — students hand strategic control to the tutor. Intervene to keep students the decision-makers. - **Structure AI use cognitively, not just permissively.** [[airis-cognitively-activated-ai-physics-2026|AIRIS]] (Activate–Inquire–Reflect) shows the value of having students predict/outline before AI, delegate computational steps while comparing output critically, and reflect afterward — treat AI integration as an instructional-design problem, and test whether learning survives AI removal. - **Guard against scoring bias.** [[ai-scoring-language-bias-physics|AI scoring]] systematically underestimates linguistically weaker students' explanations; use language-aware or human-moderated scoring for conceptual assessment. - **Use simulated classrooms for teacher preparation.** [[multiagent-classroom-dual-process-physics-teachers-2026|Simulated multi-agent classrooms]] give prospective teachers rare practice responding to authentic student reasoning — a low-cost complement to live microteaching. - **Reserve unaided practice and assessment.** The Bochum redesign ([[ai-particle-physics-education-redesign-2026|Mikhasenko et al. 2026]]) shows AI-permitted, research-shaped homework completed with high engagement can leave students far behind on an unaided exam (mean 20.6/80) — treat assisted performance and independently retrievable knowledge as distinct, and build deliberate unaided practice and a written exam into the course. ## Connected Concepts - [[stem-education]] - [[socratic-method]] - [[intelligent-tutoring]] - [[computational-thinking]] - [[ai-literacy]] - [[trust-calibration]] - [[student-experience]] - [[assessment-validity]] - [[bias-mitigation]] - [[equity-in-ai-education]] - [[automated-assessment]] - [[educational-measurement]] - [[learning-analytics]] - [[professional-training]] - [[simulation]] - [[generative-ai]] - [[higher-ed]] - [[discipline-specific-aied]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools ## Connected Articles - [[genai-video-engineering-physics-workflow-2026]] — From Prompts to Physical Laws: A Generative AI Workflow for Engineering Physics Education - [[wang-teacher-student-centered-agents-physics-2026]] — Teacher-centered vs. student-centered prompt-engineered physics agents (Wang et al. 2026) - [[omniphys-multimodal-physics-benchmark-2026]] - [[benzion-ai-physics-simulations-virtual-lab]] — Using AI to rapidly generate physics simulations / virtual labs (Ben-Zion et al. 2025) - [[hashmi-socratic-physics-chatbot-2025]] - [[socratic-ai-physics-tutor-taxonomy-2026]] - [[fouad-bentley-trust-utility-gap-physics-2026]] - [[becker-chatgpt-typology-physics-2026]] - [[llm-computational-thinking-physics-2026]] - [[ai-scoring-language-bias-physics]] - [[multiagent-classroom-dual-process-physics-teachers-2026]] - [[physics-chatbot-epistemological-beliefs-2026]] - [[ai-generated-smartphone-circular-motion-lab-2026]] - [[genai-ar-physics-simulation-prompt-2026]] - [[embodied-inquiry-ai-facilitator-physics-2026]] - [[probing-ai-generated-physics-solutions-2026]] - [[genai-assisted-problem-posing-physics-2026]] - [[airis-cognitively-activated-ai-physics-2026]] — AIRIS: A Framework for Cognitively Activated AI Augmentation in Physics - [[ai-grading-handwritten-physics-2026]] — AI grading of handwritten physics assessments (Olympiad) - [[gemini-lualatex-physics-video-transcription-2026]] — Gemini+LuaLaTeX math-accessible physics video transcription - [[chatgpt-qiskit-homework-autogradable-2026]] — ChatGPT solves Qiskit homework; autogradable design - [[ai-particle-physics-education-redesign-2026]] — AI in Particle Physics Education: Research Problems and Foundational Skills - [[mechanics-cognitive-diagnostic-physics-2026]] — Mechanics Cognitive Diagnostic: turning the FCI, FMCE and EMCS into a 14-objective cognitive diagnostic (Le et al. 2026) - [[physics-students-llm-perceptions-instruction-2026]] — Skepticism vs. Convenience: Physics Students' Perceptions and Use of Large Language Models Before and After Instruction - [[context-prompts-physics-assignments-2026]] — Artificial Intelligence Driven Physics Assignments using Context Prompts - [[ai-assisted-physics-lab-report-assessment-2026]] — AI-Assisted Assessment of Experimental Physics Laboratory Reports: Potential, Limitations, and Support for Teaching Practice --- ## [Chemistry Education](https://edtechdev.github.io/aied/concepts/chemistry-education/) > **Chemistry Education** — the study of how students learn chemistry and how to teach it more effectively, spanning [[generative-ai|GenAI]] in laboratory and experimental design, AI-mediated [[formative-assessment|formative assessment]], context-based and inquiry-based instruction, the technical accuracy of [[llm|LLMs]] on chemistry tasks, and the philosophy of experimentation in the AI age. Chemistry education research engages the discipline's distinctive demands — abstract, submicroscopic concepts, symbolic and representational notation (formulas, SMILES, spectra), and hands-on laboratory practice — which make it a rich and distinctive context for studying how AI both supports and challenges learning. ## Questions to Consider - Chemistry combines abstract submicroscopic concepts, specialized symbolic notation, and hands-on laboratory practice. Before reading, which of these three distinctive demands do you think AI handles well, and which might it struggle with — especially given the page's warning about LLMs failing on rigorous quantitative and spatial-reasoning tasks? - One study had students design lab manuals with AI, implement them hands-on, and have them validated by professionals — significantly boosting experimental confidence and critical thinking while shifting staff roles from demonstration toward guidance. What makes this 'AI for experimental design' different from just asking AI for answers? - Systematic evidence shows LLMs can define basic chemistry terms but perform poorly on rigorous quantitative tasks, struggle with spatial reasoning (like NMR), and show overconfidence. If an AI confidently gives you a wrong answer on a hard chemistry problem, how would you detect it — and what skill does that detection require? - Research proposes assigning AI distinct roles by achievement level — a Patient tutor for low-achievers, a Personal Coach for mid-level, and an Intellectual Sparring Partner for high-achievers. Why do you think the same AI tool should play different roles for different students, and what does that require of the human instructor? - The page warns of 'epistemic drift' — reliance on opaque algorithms detaching scientific inquiry from causal understanding. If AI predicts an experimental outcome, when does that prediction help you understand chemistry, and when does it quietly replace the understanding itself? ## Introduction Chemistry education has become a fertile domain for AI-in-education research because chemistry combines **abstract conceptual content**, **specialized symbolic representation**, and **physical laboratory practice**. AI tools (notably ChatGPT and conversational agents) are used to explain complex topics, support laboratory work and experimental design, provide personalized [[feedback]] and [[formative-assessment|formative assessment]], and [[simulation|simulate]] experiments. At the same time, research documents [[llm|LLMs]] **technical limits** on rigorous chemistry tasks and the risk of **epistemic drift** and over-reliance. ### Key research themes **AI-supported laboratory and experimental design** is a distinctive strength of chemistry-education research. **[[ai-supported-experimental-design-chemistry-2026|Yim & Lui]]** integrated AI [[conversational-ai|chatbots]] into an upper-division undergraduate analytical chemistry lab: students used AI to **design lab manuals**, implemented them hands-on, and had them validated by certification professionals — significantly enhancing experimental confidence and [[critical-thinking]]/problem-solving skills while shifting staff roles from "cookbook" demonstration toward guidance. The [[philosophy-experimentation-ai-chemistry-2026|philosophy-of-experimentation]] strand examines how AI reshapes the epistemology, ontology (AI predictions in a "liminal" space), and methodology of chemistry experiments, and warns of **[[agency|agency shift]]** and over-reliance. **Context-based and inquiry-based instruction** uses AI within structured [[pedagogy|pedagogies]]. **[[context-based-ai-secondary-chemistry-2026|Abdikayumova & Madybekova]]** combined the **7E instructional model** with PhET simulations and ChatGPT tutoring for Grade 10 chemistry, finding significantly higher achievement and engagement than inquiry-only or conventional teaching — demonstrating the synergy of contextualization, structured inquiry, and adaptive AI. This connects to [[constructivist]] and [[personalized-learning]] frameworks. **AI-mediated formative assessment and human–AI collaboration.** **[[instructor-ai-roles-chatgpt-formative-assessment-2026|Ratniyom et al.]]** found pre-service science [[teacher-role|teachers]] perceive distinct, achievement-based roles: the human instructor as an adaptive expert (Simplifier/Elaborator), and ChatGPT as a personalized self-regulated-learning tool shifting from *Patient [[intelligent-tutoring|tutor]]* (low-achievers) to *Personal Coach* (medium) to *Intellectual Sparring Partner* (high) — proposing an **Instructor–AI Synergistic Learning Ecosystem**. This advances [[human-ai-collaboration]], [[formative-assessment]], and [[self-regulated-learning]] research. **Technical accuracy and critical AI literacy.** Systematic evidence shows [[llm|LLMs]] can define basic chemistry terms but perform poorly on rigorous quantitative tasks and struggle with spatial reasoning (e.g., NMR), overconfidence, and notation sensitivity ([[ai-science-chemistry-education-systematic-review-2025|Erümit & Özdemir Sarıalioğlu]]; the ChemBench/QCBench findings in the [[unesco-ai-guidelines-chemical-education-2026|UNESCO perspective]]). This makes **evaluative engagement with AI output** — interrogating, verifying, and cross-checking against chemical principles — a central learning goal, connecting to [[ai-literacy]], [[critical-thinking]], and [[reducing-ai-misuse]]. **Ethics, policy, and epistemic drift.** The [[unesco-ai-guidelines-chemical-education-2026|UNESCO-guidelines perspective]] warns of **epistemic drift** — reliance on opaque algorithms detaching scientific inquiry from causal understanding — and calls for a shift from content delivery to **knowledge creation**, critical AI chemical literacy, human-reasoning-prioritizing assessment, and closing the global access gap. ### Connections to related concepts Chemistry education sits within the broader [[stem-education]] domain and shares much with [[physics-education]] (laboratory practice, abstract concepts, problem-solving) while having distinctive connections: to [[assessment]] and [[formative-assessment]] through AI-mediated evaluation; to [[simulation]] and laboratory learning through virtual experiments; to [[teacher-education]] through pre-service science-teacher research and professional development; to [[educational-policy-ai]] and [[ethics]] through the [[governance]] of AI in STEM; and to [[philosophy-of-ai-in-education]] through the epistemology/ontology of experimentation. The [[ai-literacy]] and [[reducing-ai-misuse]] concepts are essential for the responsible-use dimension, and [[higher-ed]] and [[k-12]] capture the levels at which chemistry AI research occurs. ## Implications for chemistry instructors - **Leverage AI for experimental design, not just answers.** [[ai-supported-experimental-design-chemistry-2026|Yim & Lui]] show students designing lab manuals with AI and validating them hands-on builds confidence and critical-thinking while shifting staff from demonstration to guidance — a model for lab courses. - **Demand evaluative engagement with AI output.** Systematic evidence finds [[llm|LLMs]] weak on rigorous quantitative chemistry, spatial reasoning (NMR), and overconfident — make interrogating and cross-checking AI against chemical principles an explicit learning goal. - **Assign distinct, achievement-sensitive AI roles.** [[instructor-ai-roles-chatgpt-formative-assessment-2026|Instructor–AI role research]] finds the human instructor as adaptive expert and ChatGPT as a personalized tool shifting from Patient tutor (low achievers) to Coach to Intellectual Sparring Partner (high achievers) — differentiate support by student level. - **Combine AI with structured, contextual pedagogy.** [[context-based-ai-secondary-chemistry-2026|7E + PhET + ChatGPT]] outperformed inquiry-only and conventional teaching, showing AI works best inside an established instructional model. - **Be alert to epistemic drift and over-reliance.** [[philosophy-experimentation-ai-chemistry-2026|Philosophy-of-experimentation]] research and UNESCO guidance warn that opaque AI can detach inquiry from causal understanding — preserve human-reasoning-prioritizing assessment and student agency. - **Grade open-ended handwritten work selectively.** In a 296-student handwritten general-chemistry final, a multimodal LLM graded textual answers and chemical-reaction equations reliably but drawing and graphing worse than random (background grids visually distract AI vision); pairing an [[item-response-theory|IRT]]-based risk filter and [[human-in-the-loop-ai|deferring]] graphical items to humans made automation defensible for [[summative-assessment|summative]] use ([[cvengros-grading-handwritten-chemistry-ai-2026]]). ## Connected Concepts - [[stem-education]] - [[physics-education]] - [[discipline-specific-aied]] - [[generative-ai]] - [[ai-literacy]] - [[assessment]] - [[formative-assessment]] - [[feedback]] - [[human-ai-collaboration]] - [[self-regulated-learning]] - [[constructivist]] - [[personalized-learning]] - [[simulation]] - [[critical-thinking]] - [[reducing-ai-misuse]] - [[cognitive-offloading]] - [[ethics]] - [[educational-policy-ai]] - [[teacher-education]] - [[philosophy-of-ai-in-education]] - [[higher-ed]] - [[k-12]] - [[agency]] - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools ## Connected Articles - [[ai-science-chemistry-education-systematic-review-2025]] — Systematic review of AI in science/chemistry education - [[unesco-ai-guidelines-chemical-education-2026]] — Translating UNESCO AI guidelines to chemical education - [[context-based-ai-secondary-chemistry-2026]] — Context-based + AI in secondary chemistry (7E) - [[ai-supported-experimental-design-chemistry-2026]] — AI-supported experimental design in practical chemistry - [[instructor-ai-roles-chatgpt-formative-assessment-2026]] — Instructor and AI roles in ChatGPT-enhanced formative assessment - [[philosophy-experimentation-ai-chemistry-2026]] — Philosophy of experimentation in chemistry with AI - [[cvengros-grading-handwritten-chemistry-ai-2026]] --- ## [Biology Education](https://edtechdev.github.io/aied/concepts/biology-education/) > **Biology Education** — the study of how students learn biology and how to teach it more effectively, spanning AI-assisted laboratory instruction, AI literacy embedded in biology curricula, critical thinking in the AI era, and the use of specialized tools (species identification, bioimaging, predictive modeling) in biological education. Biology's distinctive demands — large bodies of specialized terminology, visual-spatial and systems thinking, hands-on laboratory and fieldwork, and increasingly computational 'omics methods — make it a rich context for examining both the promise and the risks of AI in STEM learning. ## Questions to Consider - Biology makes distinctive demands on learners — specialized terminology, visual-spatial and systems thinking, hands-on lab work. Before reading, how do you think these features change what AI can genuinely help with in biology versus what it might undermine? - A virtual lab teaching assistant study found human TAs more accurate and effective than ChatGPT, yet students preferred the AI response 40% of the time and could only detect AI output 45% of the time. Why might students prefer a less accurate AI answer — and what does that say about the 'safety net' needed in lab teaching? - AI has revolutionized biology research (AlphaFold, computational biology) while its adoption in biology education has been cautious, amid concerns about cheating and erosion of critical thinking. Why do you think the same tools that power scientific discovery are treated with such caution in the classroom? - One study embedded machine-learning concepts into a high-school biology course and found the biology context actually supported AI learning — evidence for embedding AI literacy in the curriculum rather than making it an add-on. Where in a biology course would AI concepts naturally arise? - Researchers argue that as AI reshapes biological research, critical thinking — skepticism, contextual understanding, ethical reasoning — must be deliberately cultivated with human oversight. If AI can now identify species, interpret images, and predict outcomes, what intellectual work is left that students must learn to do themselves? ## Introduction Biology education research on AI clusters around a tension: AI has revolutionized biology *research* (AlphaFold, 'omics, computational biology), yet its adoption in biology *education* has been cautious amid concerns about generative-AI cheating, misinformation, and erosion of critical thinking. The knowledge base's biology articles collectively examine how AI tools function as laboratory teaching assistants, how AI literacy can be embedded in biology courses, how critical thinking must be protected in the AI era, and how specialized AI tools support fieldwork and species identification. ### Key research themes **AI in the biology laboratory.** **[[chatgpt-virtual-lab-teaching-assistant-biology-2026|Doğru & Faulconer]]** tested ChatGPT as a **virtual teaching assistant** in an undergraduate biology lab, finding human TAs more accurate and more effective, though students preferred the AI response 40% of the time and detected AI output only 45% of the time — highlighting both the burden-lifting potential and the safety-need for a "safety net" against incorrect lab information. This parallels the chemistry strand's [[ai-supported-experimental-design-chemistry-2026|AI-supported experimental design]] and [[philosophy-experimentation-ai-chemistry-2026|philosophy of experimentation]] — shared concern for AI in hands-on science labs. **AI literacy embedded in biology curricula.** **[[zha-ai-literacy-biology-case-study|Zha et al.]]** integrated machine-learning and neural-network concepts into a high-school honors biology course, finding significant gains in AI knowledge and that biology context supported AI learning — evidence for embedding AI literacy in [[stem-education]] rather than relegating it to extracurriculars. **[[chatgpt-math-biology-challenge-based-learning-2025|Elizondo-García et al.]]** used ChatGPT within **challenge-based learning** in biology and math courses, finding students valued immediacy but worried about veracity, [[teacher-role|teacher]] replacement, and skill erosion — calling for updated academic-integrity codes and AI-use ethics. **Critical thinking in the AI era.** **[[critical-thinking-biological-sciences-ai-2025|Papaneophytou & Nicolaou]]** argue that as AI shapes biological research, **critical thinking** — skepticism, contextual understanding, and ethical reasoning — must be deliberately cultivated, with [[human-in-the-loop-ai|human oversight]] remaining indispensable to validate AI outputs and prevent [[equity-in-ai-education|bias]]. This connects to the knowledge base-wide [[reducing-ai-misuse]] and [[cognitive-offloading]] concerns. **Specialized AI tools and the broader review.** **[[beyond-chatgpt-ai-tools-biological-education-2026|Cotton & Cotton]]** review the full landscape of AI tools in biological education, including **iNaturalist and Google Lens** for species identification, bioimaging and machine-learning tools, assistive technologies, and predictive modeling of at-risk students — alongside the integrity, assessment-design, misinformation, and critical-thinking challenges of generative AI. ### Connections to related concepts Biology education sits within the broader [[stem-education]] domain and shares the laboratory-practice and specialized-terminology concerns of [[chemistry-education]] and [[physics-education]]. Distinctive connections: to [[critical-thinking]] through the AI-era emphasis on skepticism and oversight; to [[ai-literacy]] through embedding AI concepts in biology courses; to [[human-ai-collaboration]] through virtual lab assistants and challenge-based learning; to [[academic-integrity]] and [[ethics]] through generative-[[ai-misuse-learning-harm|AI misuse]] and policy; and to [[assessment]] through AI-mediated evaluation. The [[higher-ed]] and [[k-12]] concepts capture the levels at which biology AI research occurs, and [[intelligent-tutoring]] and [[simulation]] connect to AI-assisted lab and teaching support. ## Implications for biology instructors - **Keep a human "safety net" around AI lab assistants.** [[chatgpt-virtual-lab-teaching-assistant-biology-2026|Doğru & Faulconer]] find human TAs more accurate and effective than ChatGPT, yet students preferred AI 40% of the time and detected AI output only 45% — verify and annotate AI lab information, and never rely on it unmoderated for safety-critical content. - **Embed AI literacy in the biology curriculum, not as an add-on.** [[zha-ai-literacy-biology-case-study|Zha et al.]] show biology context supports AI learning; integrate ML/neural-network concepts into coursework where they naturally arise. - **Deliberately cultivate critical thinking in the AI era.** [[critical-thinking-biological-sciences-ai-2025|Papaneophytou & Nicolaou]] argue skepticism, contextual understanding, and ethical reasoning must be taught explicitly, with human oversight to validate AI outputs and prevent bias. - **Teach responsible use alongside challenge-based learning.** [[chatgpt-math-biology-challenge-based-learning-2025|CBL studies]] find students value immediacy but worry about veracity and skill erosion — update academic-integrity codes and AI-use ethics in parallel with adoption. - **Survey the specialized tool landscape.** [[beyond-chatgpt-ai-tools-biological-education-2026|Cotton & Cotton]] map tools (iNaturalist, Google Lens, bioimaging, at-risk prediction) that instructors can deploy for fieldwork, species ID, and student support — choose purpose-built tools over general chatbots where they fit. ## Connected Concepts - [[stem-education]] - [[chemistry-education]] - [[physics-education]] - [[discipline-specific-aied]] - [[generative-ai]] - [[ai-literacy]] - [[critical-thinking]] - [[human-ai-collaboration]] - [[academic-integrity]] - [[ethics]] - [[intelligent-tutoring]] - [[simulation]] - [[feedback]] - [[assessment]] - [[reducing-ai-misuse]] - [[cognitive-offloading]] - [[higher-ed]] - [[k-12]] - [[educational-policy-ai]] ## Connected Articles - [[biology-degree-integrity-genai-cheating-2026]] — Can students cheat their way to a biology degree? A case study of the vulnerability of biology course grades to academic dishonesty in the era of generative AI - [[chatgpt-virtual-lab-teaching-assistant-biology-2026]] — ChatGPT as a virtual lab teaching assistant - [[beyond-chatgpt-ai-tools-biological-education-2026]] — Review of AI tools in biological education - [[critical-thinking-biological-sciences-ai-2025]] — Critical thinking in biological sciences and AI - [[chatgpt-math-biology-challenge-based-learning-2025]] — ChatGPT in challenge-based biology/math courses - [[zha-ai-literacy-biology-case-study]] — AI literacy education in a biology class --- ## [CS Education](https://edtechdev.github.io/aied/concepts/cs-education/) > **CS Education** — computer [[science-education|science education]] is the most-researched STEM subfield in the knowledge base, benefiting from natural alignment between AI tools and programming tasks. Code generation, debugging assistance, and automated code review are its primary AI applications. Because students learn to build the very tools they use, CS education sits at the center of debates about AI literacy, curriculum redesign, agentic software engineering, and the boundary between genuine learning and [[cognitive-offloading|over-reliance]]. ## Questions to Consider - If AI can now write code that passes real programming exams, what should students still learn to do by hand — and what should the curriculum stop teaching? - [[research-methods-aied|Research]] found that higher trust in an AI coding assistant predicted WORSE ability to tell correct from misleading suggestions. How is trust different from appropriate reliance, and how would you teach the latter? - In student-AI co-programming, nearly 80% of interactions relied on non-learning strategies like outsourcing answers, and only about 1 in 9 showed deep epistemic [[student-engagement|engagement]]. Why does genuine learning rarely happen by default when AI is available? - A [[learning-by-teaching]] agent that was too competent undermined students' debugging practice. Would you deliberately make an AI tutor fallible — and if so, how? - As AI automates implementation, curricula are shifting from writing code to verifying and directing AI-generated artifacts. What new competencies does that demand, and what might get lost in the shift? - Students are building the very tools they use. How does being both builder and user of AI change what they should learn about its limits — and its ethics? ## Introduction ### AI in CS education - **Code generation and completion:** [[code-review-genai-cs1|CS1 code review]], [[dura-llm-cs2|DURA for CS2]], and [[prompt-problems-nl-programming-mistakes|NL programming mistakes]] examine how students use AI for code generation and what they learn from it. - **Conversational agents for novices ([[meta-analysis-systematic-review|scoping review]]):** [[conversational-agents-novice-programmers-scoping-2025|Barzanji & Loitsch (2025)]] map 23 studies (2019–June 2024) of [[conversational-ai|conversational agents]] for novice programmers, documenting a shift from rule-based chatbots to [[llm]]- and [[rag]]-based agents (with [[rag|retrieval-augmented generation]] reducing [[hallucination-risk|hallucination]]) and personalized tutoring support (e.g., InfoBot, ProbSol-Bot, Lint Bot, Profe Alex). Notably, only 4 of 23 studies ground design in [[learning-theories|learning theory]], and 17 of 23 prototypes are English-only despite most research originating outside English-speaking countries — flagging weak [[pedagogy|pedagogical]] grounding and an inclusivity gap for future CA design in introductory programming. - **Debugging support:** [[debugtracker-classroom-debugging|Debugging tools]], [[chat-debugging-human-ai-collaboration-circuits|human-AI debugging collaboration]], and [[golrang-propact-pair-programming-2026|dyadic pair-programming modeling]] leverage AI for error identification and repair. - **Automated assessment:** [[automated-grading-linux-bash-examinations-large-language-models|Linux Bash grading]], [[llm-automated-grading-programming-comparison-2026|a large-scale 18-model grading comparison]], and [[llm-intervention-design-cs-review|LLM intervention review]] evaluate automated code assessment. The security of that workflow is a separate question: [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] red-teamed a routine AI grading task and found that instructions hidden inside a submitted file raised a failing essay's grade with no visible warning, in 9 of 9 iterations for one strategy and 17 of 18 for another — evidence that grader robustness belongs on the [[assessment-validity]] checklist alongside accuracy. - **AI-generated learning media:** [[ai-generated-traces-novice-programmers|Generated Animated Traces]] show that AI-generated visualizations can aid immediate learning but must be personalized — mid-engagement students experienced a performance decrement consistent with the expertise-reversal effect. - **GenAI analogy critique as an instructional asset:** [[student-reception-genai-analogies-computing-2026|Bernstein & Sibia (2026)]] ground GenAI analogy reception in CS2: ten students who had already completed CS2 audited GenAI-generated analogies for linked lists and recursion and rejected mappings that failed structural correspondence — an island-route analogy that mapped to a circular rather than a singly linked list, or a badminton rally offered for recursion despite having no guaranteed shrinking input, with one student proposing golf instead. The work argues that analogy critique is itself a check on concept understanding, making flawed AI analogies a usable instructional asset rather than a hazard to filter out, and recommends assigning them as objects to inspect and repair. - **[[misconceptions|Misconception]] modeling:** [[student-misconceptions-conditionals-loops-taxonomy|a taxonomy of conditionals/loops misconceptions]] gives automated systems a precise vocabulary for diagnosing novice errors. - **Model-generated strategic misconceptions:** [[milicevic-socratic-trap-strategic-misconceptions-2026|Miličević et al. (2026)]] built SocraticTrap-CS around 35 concepts from the ACM/IEEE CS2023 curriculum — algorithms, programming languages, databases, networks and operating systems — and asked seven open-weight models for a fluent, authoritative explanation resting on a subtle error. Six of the seven produced an expert-confirmed strategic misconception for 91% or more of prompted concepts (221 of 241 segments, 91.7%), with no significant differences between CS domains; conceptual errors dominated (66.5% conceptual vs. 33.5% factual, and none purely logical), and both persuasiveness and error type varied by domain. The authors therefore recommend domain-sensitive countermeasures — reasoning-focused checks in programming-heavy courses, cross-referencing against protocol specifications in networking — and evaluation of [[automated-question-generation]] and AI-authored explanations on pedagogical [[trust|trustworthiness]] rather than correctness alone. - **[[authentic-assessment]] performance:** [[genai-oop-programming-assessments-2026|Lepp & Kaimre (2026)]] show 2026 [[generative-ai|GenAI]] systems outscore the average student cohort on authentic introductory OOP [[assessment|assessments]] and frequently earn full marks on longer programming tasks, yet still struggle with interfaces, abstract classes, inheritance, and image-based questions — recurring error patterns instructors can exploit when designing assessments. - **Predictive modeling for at-risk support:** [[zhang-ml-student-progress-programming-2026|Zhang, Jeffries & Koprinska (2025)]] show that intrinsically interpretable decision trees trained on content-interaction log features accurately predict module-level progress in large-scale online programming courses (85–91% accuracy across four K-12 courses) and flag "No submission" dropout outcomes, giving educators a 7–8 day window to [[teacher-role|intervene]] with struggling and disengaged [[learners]] before module deadlines — complementing the automated-grading and attrition-prediction work above. ### Programming pedagogy: from blocks to embodied, game-based learning Programming education spans introductory block-based programming to advanced software development, and increasingly grounds abstract code in concrete, observable outcomes. - **Block-based visual programming:** environments like Scratch and Blockly let beginners snap together graphical blocks rather than type text, eliminating syntax errors and making program structure visible — especially valuable for younger learners and for controlling [[educational-robotics|educational robots]]. In the AI era they are increasingly combined with conversational AI agents (e.g., [[microbit-robotics-machine-learning-teacher-training-2026|Micro:bit + MakeCode in teacher training]], [[cstutorbench-slm-tutors|small-language-model tutors]]). - **[[embodied-learning|Embodied]] block programming:** [[roboblockly-conversational-block-robotics-ct-2026|RoboBlockly Studio]] combines block-based programming with a conversational AI teaching agent and embodied robot execution, creating an iterative authoring–running–observing–revising loop that preserves learner [[agency]]. - **Natural-language robot control:** [[edusim-llm-robotic-simulation-education-2026|EduSim-LLM]] lets beginners control simulated robots through natural-language instructions, lowering the barrier to robot programming without requiring low-level code expertise. - **Robotics and computational thinking:** [[computational-thinking-educational-robotics-secondary-2026|Valls i Pou]] links computational thinking to educational robotics in secondary STEAM curricula, and [[microbit-robotics-machine-learning-teacher-training-2026|teacher-training research]] argues robotics and ML activities should be embedded in [[teacher-education]]. - **Game-based and gamified learning:** [[game-based-gamified-robotics-education-review-2026|A systematic review]] compares game-based learning (suited to informal settings) and gamification (suited to formal classrooms) in robotics education, which emphasizes introductory programming and modular kits. - **Project-based robotics:** [[bots-blocks-project-based-robotics-education-2026|Bots and Blocks]] teaches robotics programming through an agile, semester-spanning [[project-based-learning|project]], addressing the theory-practice gap in higher-ed. - **LLM impact on [[learning-gains|learning outcomes]]:** [[jost-llm-programming-education-learning-outcomes|Jošt et al. (2024)]] and [[genai-meta-analysis-programming-learning|a meta-analysis of GenAI and programming learning]] examine whether AI-assisted tools help or undermine programming achievement. ### Curriculum transformation in the AI era The question "what should students still learn by hand?" now reshapes computing programs. - **From implementation to verification:** [[reshaping-cs-education-genai|Reshaping Undergraduate CS Education]] argues that as GenAI automates implementation-level programming, debugging, and testing, curricula must shift toward *understanding and verifying AI-generated artifacts*, preserving system design, abstraction, and [[critical-thinking|critical evaluation]] while de-emphasizing low-level implementation details. This aligns with [[ai-literacy]] frameworks that prize evaluation over generation. - **Agentic software engineering as a discipline:** [[ase-26-agentic-software-engineering-curriculum|ASE-26]] formalizes directing agents rather than writing code — teaching auditability, context engineering, verification, multi-agent workflows, and AgentOps — and positions [[agentic-ai|agentic AI]] competence as a structured, scaffolded curriculum rather than syntax mastery. - **New pedagogies and assessment models:** [[test-driven-ai-assisted-learning|Test-Driven AI-Assisted Learning]] replaces lectures with [[self-directed-learning|self-directed]] AI-assisted study gated by weekly closed-book tests, preserving individual accountability while AI agents scale material production and marking under [[human-in-the-loop-ai|human oversight]]. - **What predicts [[vibe-coding]] success — and what to keep teaching:** [[vibe-coding-writing-cs-achievement-2026|A preregistered CHI 2026 study (N=100)]] of pure "no-code" vibe-coding found that both computer-science achievement (r = .39) and written-communication proficiency (r = .29) independently predicted performance, with CS achievement remaining significant even after controlling for domain-general cognitive ability and contributing roughly twice the unique variance of writing skill. Because the environment hid generated code, CS knowledge could only help indirectly (problem decomposition, algorithmic thinking) — making the CS estimate a *lower bound* for AI-assisted workflows that also permit editing. The authors argue curricula should weigh written communication alongside CS fundamentals, rather than treating vibe coding as syntax mastery made obsolete. ### AI literacy, agency, and the risk of over-reliance Because programming is where AI assistance is most powerful, it is also where the failure modes are most visible. - **Trust ≠ appropriate reliance:** [[trust-reliance-ai-education-2026|Trust and reliance on AI (Pitts et al.)]] find that higher trust in an AI assistant predicted *worse* discrimination between correct and misleading suggestions during Python [[problem-solving]] — moderated by [[ai-literacy]] and need for cognition. Calibration, not confidence, is the goal. - **Epistemic AI literacy:** [[constructing-epistemic-ai-literacy-student-ai-co-programming|Wu (2026)]] shows that in student-AI co-programming, 78.8% of interactions relied on non-mastery-oriented aims and unreliable strategies (outsourcing, verification-seeking), with only 11.1% showing high epistemic engagement — genuine learning rarely emerges without deliberate design support. - **Structural interventions against copy-paste over-reliance:** [[soft-barriers-copying-ai-programming-2026|Soft barriers for copying in AI-assisted programming]] evaluate lightweight design interventions (e.g., mechanisms that discourage blind copy-paste of AI output) and find they can reduce over-reliance without blocking AI assistance — evidence that the [[cognitive-offloading|over-reliance]] risk in CS education is amenable to instructional-design fixes, not just learner-education or bans. - **Teachable agents and productive practice:** [[chatgpt-teachable-agent-programming-lbt-2024|Learning-by-teaching with ChatGPT]] improved knowledge gains and code quality but undermined error-correction practice because the agent is too competent — a design lesson: make agents *deliberately fallible* so debugging is preserved. - **Assistance governance:** [[llm-programming-support-governance-cs-education|a scoping review of 90 systems]] introduces the **PEA framework** (Policy, Enforcement, Authority) for bounding and controlling LLM assistance — a comparative vocabulary for designing [[scaffolding]] that limits over-reliance. - **Behavioral context for adaptive AI tutoring:** [[tutortrace-learner-behavioral-states-2026|Barron et al. (2026)]] present **TutorTrace**, a dataset and pipeline that makes learners' behavioral context computable in real time from IDE telemetry in AI-assisted Python courses (N=480). It derives a taxonomy of activity before, between, and across AI queries, and can classify whether an upcoming query reflects guided or dependent [[help-seeking]] (AUROC=.717) and predict imminent queries (AUROC=.726); behavior-aware prompts reduced no-independent-work query intervals from 50.0% to 20.7% in a preliminary evaluation. This shows how behavioral telemetry can make [[intelligent-tutoring|AI programming tutors]] adaptive to learners' actual effort, not just their explicit requests. - **The duality of building what you use:** CS students' unique position creates both [[metacognition|meta-cognitive]] awareness of AI limitations and real risk of [[cognitive-offloading|Over-Reliance]] on AI-generated code. [[code-review-genai-cs1|Code review interviews]] and [[critical-engagement-code-completion|critical engagement studies]] address this tension directly. - **Agentic coding and comprehension in team PBL:** [[spec-driven-development-ai-agents-sdpbl-2026|Tanaka et al. (2026)]] introduced Spec-Driven Development with [[agentic-ai|AI agents]] into an undergraduate software-engineering project course and found that implementation throughput (added LOC) rose across 2022-2025 while heavy AI use coincided with code-comprehension dips that only recovered after one-on-one instructor checks - direct evidence that throughput gains do not guarantee understanding, and that [[cognitive-offloading|over-reliance]] in AI-assisted coding is amenable to instructional monitoring. ### Equity, culture, and who gets into computing - **Broadening participation:** [[suacode-african-students-motivations|SuaCode]] documents motivations for smartphone-based coding among African students (fewer than 1% of secondary-school leavers have fundamental coding skills), informing accessible AI-supported MOOCs for [[equity-in-ai-education|low-resource contexts]]. - **Neurodivergence and collaboration:** [[neurodivergent-computing-students|Neurodivergent computing students]] report discomfort with ambiguous collaboration structures; structured assignments, smaller consistent teams, and explicit roles improve [[accessibility]] — design lessons for the AI tools entering computing classrooms. - **Culture shapes perceived ethics:** [[cross-cultural-student-perceptions-genai-computing|Canadian vs. South Korean computing students]] judged identical AI-assisted coding practices differently despite functionally identical policies — policy harmonization does not produce perception harmonization, an [[academic-integrity]] and [[equity-in-ai-education]] concern. - **Collaboration [[explainable-ai|transparency]]:** [[student-perception-ai-use-collaboration|Graf et al.]] find that partners' misaligned beliefs about each other's AI use predict lower project scores, especially for lower-performing students — transparency mechanisms (disclosures, shared logs) may be needed in collaborative programming. ### Ethics education and the workforce - **Ethics-to-behavior gap:** [[cost-of-ethics-crisis-cs-ethics-education|the "Cost-of-Ethics Crisis"]] shows CS students, despite contemporary ethics education, prioritize compensation, location, and culture over [[ethics|ethical]] concerns in job searches — a critical gap in how ethics instruction transfers to behavior. - **Workforce reshaping:** [[ai-engineering-computing-workforce-grey-literature-2026|a systematic review of U.S. gray literature]] frames the "Dual Train Problem" — rapid AI change racing [[governance|institutional]] adaptation — and urges durable AI competencies, ethics/governance, and skill-based credentials aligned with emerging roles (e.g., [[prompt-engineering]], AI auditing, [[educational-policy-ai|AI policy]]). ### Connections CS education connects to [[computational-thinking]], [[stem-education]], [[automated-assessment|Automated Grading]], [[prompt-engineering]], [[ai-literacy]], [[agentic-ai]], [[curriculum-design]], [[human-ai-collaboration]], [[higher-ed]], [[k-12]], and [[professional-training]]. Its closest applied neighbour is [[information-technology]]: CS education takes the program and the algorithm as its object, whereas IT education takes the deployed organizational system and the practitioner's judgment about it — which is why the two fields debate different AI harms, whether code generation erodes programming skill against whether AI troubleshooting erodes diagnostic skill, and ask different questions of their graduates, whether they can build a system against whether they can govern the systems they administer. It is the domain where [[ai-education|AIED]] tools are both used and built, making it a testbed for [[intelligent-tutoring]], [[educational-robotics]], [[collaborative-learning]], [[game-based-learning]], and the risks of [[cognitive-offloading|Over-Reliance]]. **A 72-study synthesis and the VIE framework.** [[kumar-genai-computing-education-systematic-review-2026|Kumar, Wongsirichot and Nanthaamornphong (2026)]] reviewed the empirical literature on generative AI in computing and programming education (January 2022 – April 2026, 72 studies, 33 venues) and foreground exactly the structural feature that makes the discipline distinctive: the AI generates the assessable artifact itself, so using the tool, learning the skill and being assessed collapse into one keystroke. Their synthesis of 14 themes finds the field's most replicated effect — short-term efficiency and completion gains (36 studies) — is also its most misleading: those gains do not [[transfer-of-learning|transfer]] to unaided performance (21 studies), and [[prior-knowledge|prior knowledge]] moderates whether assistance becomes durable skill or a crutch. [[ai-detection|Detection]] research is thin (3 studies) while course redesign is comparatively well evidenced (25 studies), and the review consolidates the corpus into three interdependent design requirements — Verification, Implementation and Equity — where critical engagement with AI output must be a graded, observable component of student work rather than an aspiration left to student discretion ([[scaffolding]], [[assessment-validity]]). ## Implications for computing instructors - **Design assessments AI cannot coast through.** Exploit GenAI's recurring failure patterns (interfaces, abstract classes, inheritance, image-based tasks) instead of banning tools outright — [[genai-oop-programming-assessments-2026|GenAI systems still struggle there]]. - **Calibrate trust, don't just build it.** [[trust-reliance-ai-education-2026|Trust-reliance research]] shows higher trust predicted *worse* discrimination of misleading AI suggestions; teach verification and critical evaluation, moderated by AI literacy and need for cognition. - **Keep debugging and [[desirable-difficulties|productive struggle]] alive.** Choose tools or deliberately fallible agents ([[chatgpt-teachable-agent-programming-lbt-2024|learning-by-teaching]]) that preserve error-correction practice, and personalize AI-generated media to avoid expertise-reversal effects ([[ai-generated-traces-novice-programmers|expertise-reversal]]). - **Govern AI assistance explicitly.** Define policy, enforcement, and authority for LLM support ([[llm-programming-support-governance-cs-education|PEA]]) rather than leaving boundaries implicit. - **Shift curricula toward verification and agent direction.** As GenAI automates implementation, teach understanding/verifying AI artifacts ([[reshaping-cs-education-genai|reshape curricula]]) and structured agentic-software-engineering skills ([[ase-26-agentic-software-engineering-curriculum|ASE-26]]). - **Structure collaboration for all learners.** Smaller consistent teams, explicit roles, and AI-use transparency support [[neurodiversity|neurodivergent]] students and fair collaboration, especially where misaligned AI-use beliefs lower project scores. - **LLM-adaptive explanations of programming errors (2026):** A crowdsourced study (N=103) found LLM-rewritten error messages improve readability, but objective debugging performance depends on matching explanation style (pragmatic vs contingent) to programmer skill — a scaffolding insight for AI-assisted programming education ([[llm-adaptive-programming-error-explanations-2026]]). ## Connected Concepts - [[computational-thinking]] - [[vibe-coding]] - [[stem-education]] - [[information-technology]] - [[automated-assessment]] - [[prompt-engineering]] - [[ai-literacy]] - [[agentic-ai]] - [[curriculum-design]] - [[human-ai-collaboration]] - [[higher-ed]] - [[k-12]] - [[educational-robotics]] - [[game-based-learning]] - [[generative-ai]] - [[intelligent-tutoring]] - [[cognitive-offloading]] - [[teacher-education]] - [[professional-training]] - [[ai-education]] - [[collaborative-learning]] - [[prior-knowledge]] - [[scaffolding]] - [[assessment-validity]] ## Connected Articles - [[kumar-genai-computing-education-systematic-review-2026]] — Systematic review of 72 studies: efficiency gains that do not transfer, and the VIE framework - [[vibe-coding-writing-cs-achievement-2026]] — CS achievement and writing skills predict vibe-coding proficiency (CHI 2026) - [[tutortrace-learner-behavioral-states-2026]] - [[mechanical-engineering-ai-curriculum-2026]] — Project-Based AI Education Curriculum in Thermal Engineering - [[zhan-chapman-genai-cs-education-2026]] — GenAI in CS education - [[code-review-genai-cs1]] — CS1 code review of AI-generated code - [[dura-llm-cs2]] — DURA: LLM assistants for CS2 - [[reshaping-cs-education-genai]] — reshaping undergraduate CS curricula for GenAI - [[ase-26-agentic-software-engineering-curriculum]] — ASE-26 agentic-software-engineering curriculum - [[test-driven-ai-assisted-learning]] — Test-Driven AI-Assisted Learning - [[genai-oop-programming-assessments-2026]] — GenAI performance on authentic introductory OOP assessments (Lepp & Kaimre 2026) - [[trust-reliance-ai-education-2026]] — trust vs. appropriate reliance during Python problem-solving - [[constructing-epistemic-ai-literacy-student-ai-co-programming]] — epistemic AI literacy in student-AI co-programming - [[chatgpt-teachable-agent-programming-lbt-2024]] — learning-by-teaching with ChatGPT - [[llm-programming-support-governance-cs-education]] — PEA framework for bounding LLM assistance - [[conversational-agents-novice-programmers-scoping-2025]] — Scoping review of conversational agents for novice programmers - [[debugtracker-classroom-debugging]] — DebugTracker classroom debugging - [[llm-automated-grading-programming-comparison-2026]] — 18-model automated grading comparison - [[ai-generated-traces-novice-programmers]] — AI-generated animated traces - [[student-misconceptions-conditionals-loops-taxonomy]] — conditionals/loops misconception taxonomy - [[jost-llm-programming-education-learning-outcomes]] — LLM impact on programming learning outcomes (Jošt et al.) - [[genai-meta-analysis-programming-learning]] — meta-analysis of GenAI and programming learning - [[golrang-propact-pair-programming-2026]] — dyadic pair-programming modeling - [[critical-engagement-code-completion]] — critical engagement with code completion - [[suacode-african-students-motivations]] — SuaCode smartphone-based coding in Africa - [[cross-cultural-student-perceptions-genai-computing]] — cross-cultural perceptions of AI-assisted coding - [[neurodivergent-computing-students]] — neurodivergent computing students - [[microbit-robotics-machine-learning-teacher-training-2026]] — Micro:bit + ML in teacher training - [[computational-thinking-educational-robotics-secondary-2026]] — computational thinking and educational robotics - [[roboblockly-conversational-block-robotics-ct-2026]] — RoboBlockly embodied block programming - [[edusim-llm-robotic-simulation-education-2026]] — EduSim-LLM natural-language robot control - [[llm-computational-thinking-physics-2026]] — LLM support for computational thinking in physics - [[studychat-student-dialogues-chatgpt-ai-course-2026]] — The StudyChat dataset of student–LLM dialogues in an AI course - [[student-ai-inquiry-types-cs2-2026]] — Analysis of Types of Inquiries in Student-AI Interaction - [[chatgpt-qiskit-homework-autogradable-2026]] — ChatGPT solves Qiskit homework; autogradable design - [[llm-adaptive-programming-error-explanations-2026]] — LLM adaptive explanations of programming errors - [[astor-computational-thinking-meta-review-2026]] — Meta-review situating CT in CS education - [[soft-barriers-copying-ai-programming-2026]] — Copy-paste resistance in AI-assisted programming - [[predicting-attrition-competitive-programming]] — Predicting Student Attrition in Competitive Programming - [[zhang-ml-student-progress-programming-2026]] - [[spec-driven-development-ai-agents-sdpbl-2026]] - SDD with AI agents in a software PBL course; throughput vs. comprehension - [[genai-cognitive-tutor-programming-2026]] — GenAI as informal cognitive tutor in novice programming learning - [[student-reception-genai-analogies-computing-2026]] — Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education - [[milicevic-socratic-trap-strategic-misconceptions-2026]] — SocraticTrap-CS: benchmarking models' capacity to generate strategic misconceptions across the CS curriculum - [[humble-prompt-injection-ai-grading-red-team-2026]] — Prompt injection red-team of AI-mediated grading: hidden instructions that move the mark undetected - [[algorag-rag-theoretical-cs-education-2026]] — AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory - [[llms-unplugged-teaching-resources-2026]] — LLMs Unplugged: Teaching Resources for a ChatGPT World --- ## [Engineering Education](https://edtechdev.github.io/aied/concepts/engineering-education/) > **Engineering Education** — the study of how students learn engineering and how to teach it effectively, spanning how AI transforms engineering pedagogy, assessment, faculty development, and the engineering workforce. The engineering education articles in this knowledge base cluster around several themes: the [[engineering-faculty-metaphors-ai-understanding-2026|figurative language]] instructors use to make sense of AI, [[ethical-use-ai-engineering-education-review-2026|ethical governance]] of AI use, [[multimodal-embodied-cognition-oral-explanations-2026|embodied and multimodal assessment]] of conceptual understanding, and how AI is reshaping the [[ai-engineering-computing-workforce-grey-literature-2026|engineering and computing workforce]]. ## Questions to Consider - Engineering is a profession where graduates' decisions affect public safety and infrastructure. If students rely on AI to do design and [[problem-solving]] work, what is at stake that wouldn't matter in a purely academic subject? - Engineering instructors in the same department often hold fundamentally different mental models of AI — some see it as a social 'agent,' others as a technical 'tool.' How might those divergent framings change what students learn to do with AI depending on which instructor they get? - Hands-on laboratory and [[embodied-learning|embodied learning]] have long defined engineering education. If AI can simulate or even automate design tasks, does hands-on experience still matter — or is something genuinely lost when practice becomes simulated? - One line of research assesses engineering understanding by tracking students' gestures alongside their speech, finding that close gesture–speech coupling signals coherent conceptual understanding. What could a student's hands reveal about understanding that their words alone might hide? - Research finds ethical guidance about AI in engineering is mostly student-facing and compliance-oriented, with weak reciprocal accountability for faculty and institutions. Whose responsibility should the safe use of AI in engineering really be? ## Introduction Engineering education research is distinctive because it sits at the intersection of professional formation and rigorous STEM content. It emphasizes [[design-thinking|design]], problem-solving, hands-on and laboratory learning, teamwork, and preparing graduates for professional practice. Design is also the ground it shares with [[design-education]], and the two divide it by what can be checked: engineering's assessed object admits calculable artifacts and carries public-safety accountability, while design education's is the studio process behind the artifact — sketch, iteration, critique — the very evidence generative tools can now produce without the process that once generated it. AI raises distinctive questions here: whether hands-on and embodied experience still matters when AI can simulate or automate design tasks; the professional and ethical stakes of AI use (engineers' decisions affect public safety, infrastructure, and [[sustainability]]); and how AI reshapes the competencies graduates need and the workforce they enter. ## How AI appears in the knowledge base's engineering education research - **Faculty understanding and shared language:** [[engineering-faculty-metaphors-ai-understanding-2026|Gerhardt et al.]] analyze the metaphors engineering instructors use to describe AI, finding that most frame it either as a human-like "social being/agent" or as a "technical tool," and that instructors within the same department often hold fundamentally different mental models. Because metaphors both construct and constrain understanding, they argue a shared, accurate language is essential for [[educational-development]] and departmental discussions about AI. - **Ethics and responsible use:** [[ethical-use-ai-engineering-education-review-2026|Osunbunmi et al.]] [[meta-analysis-systematic-review|systematically review]] empirical studies of AI in undergraduate engineering education, identifying seven recurring forms of ethical guidance (transparency, accountability, student independence/agency, privacy, [[academic-integrity]], fairness/[[equity-in-ai-education|equity]]/bias, beneficence). They find ethical guidance is predominantly student-facing and compliance-oriented, with reciprocal faculty and institutional accountability underdeveloped — a concern heightened by engineering's direct stake in public safety and societal [[well-being]]. - **Embodied and [[multimodal]] assessment:** [[multimodal-embodied-cognition-oral-explanations-2026|Morphew et al.]] develop a multimodal framework integrating computer-vision gesture tracking with [[llm]] analysis of speech to assess engineering students' conceptual understanding of statistics, showing that gesture adds diagnostic evidence beyond speech and that close gesture–speech coupling signals coherent understanding. - **Workforce transformation:** [[ai-engineering-computing-workforce-grey-literature-2026|Fletcher et al.]] review U.S. gray literature on AI and the engineering/computing workforce, framing the "Dual Train Problem" (rapid change vs. urgent policy) and recommending durable AI competencies, [[ethics]] and [[governance]], and skill-based credentials for emerging roles. - **Student adoption and reliance:** [[tam-critical-use-genai-engineering-2026|Nguyen et al.]] extend the Technology Acceptance Model with *critical use* to model how engineering/CS students form intentions to use [[generative-ai|GenAI]] and how that intention predicts reliance across understanding, assessment, programming, and engineering-project tasks — finding moderate, appropriate reliance and heavy use for understanding-related tasks but limited use for full assessment writing. [[socio-cognitive-genai-adoption-engineering-2026|Asag & Al Mamun]] integrate TAM with [[technology-acceptance-model|UTAUT]] to model Bangladeshi engineering students' adoption, explaining 64% of usage variance and highlighting job relevance, result demonstrability, and subjective norms as key drivers. - **[[intelligent-tutoring|AI tutoring]] and feedback for [[quantitative-research|quantitative]] engineering courses:** [[yin-arthur-ai-teaching-assistant-engineering-econ-2026|Yin et al. (2026)]] introduce Arthur, an AI teaching assistant that delivers real-time, personalized feedback on Calculated Formula Questions in an undergraduate Engineering Economics course — a domain where pen-and-paper, unstructured solutions have blocked prior AI support. Its full life-cycle pipeline (curating previously graded handwritten submissions, random-masking [[machine-learning|data augmentation]], per-question XGBoost diagnosis backbones, and a dialogue-based question-bank web interface) offers a scalable pathway for [[ai-feedback-quality|AI feedback]] across engineering courses that lack structured digital data, and the framework is designed to generalize to CFQs in other engineering disciplines. ## Signature concerns Engineering education's signature concerns shape how AI is taken up: **hands-on and laboratory learning** (and what is lost when AI simulates practice), **professional formation** (ethical responsibility, safety, and societal impact), **design and problem-solving** (how much cognitive work AI should do), and **competency-based preparation for the workforce**. These make engineering education a rich site for studying whether AI augments or displaces the cognitive work essential to professional expertise — paralleling debates in [[physics-education]], [[math-education]], and [[cs-education]]. ## Connections to related concepts Engineering education sits within [[stem-education]] and connects strongly to [[cs-education]] (computing and software engineering), [[math-education]] and [[physics-education]] (engineering's mathematical and physical foundations), [[professional-training]] and [[educational-development]] (workforce and instructor preparation), [[assessment]] and [[ethics]] (the professional and evaluative stakes of AI), and [[higher-ed]] (the institutional context). It is the home discipline for the knowledge base's ASEE-sourced articles. ## Under-covered sub-areas The knowledge base's engineering education coverage is still developing. Sub-areas that would benefit from further articles include **[[discipline-specific-aied|discipline-specific]] engineering [[pedagogy|pedagogies]]** (mechanical, civil, chemical, electrical, software, and bioengineering education), **[[design-education|design]] and maker education**, **capstone and [[project-based-learning|project-based learning]]**, and **engineering ethics education** — where AI's role is likely to be especially consequential. ## Implications for engineering instructors - **Build a shared, accurate language about AI.** [[engineering-faculty-metaphors-ai-understanding-2026|Faculty metaphor research]] finds instructors hold fundamentally different mental models (agent vs. tool); align departmental understanding before making policy. - **Close the ethics-to-accountability gap.** [[ethical-use-ai-engineering-education-review-2026|Ethics reviews]] find guidance is mostly student-facing and compliance-oriented; strengthen reciprocal faculty and institutional accountability, given engineering's stake in public safety. - **Use multimodal, embodied assessment.** [[multimodal-embodied-cognition-oral-explanations-2026|Gesture + speech assessment]] adds diagnostic evidence beyond language alone — consider embodied cues when evaluating conceptual understanding. - **Prepare students for the workforce, not just the course.** The [[ai-engineering-computing-workforce-grey-literature-2026|Dual Train Problem]] urges durable AI competencies, ethics/governance, and skill-based credentials aligned with emerging roles. - **Model critical use and appropriate reliance.** [[tam-critical-use-genai-engineering-2026|Critical-use TAM research]] shows students rely on AI heavily for understanding tasks but less for assessment — guide them toward appropriate, verifiable reliance across task types. ## Connected Concepts - [[problem-based-learning]] - [[stem-education]] - [[cs-education]] - [[design-education]] - [[math-education]] - [[physics-education]] - [[professional-training]] - [[educational-development]] - [[assessment]] - [[ethics]] - [[higher-ed]] - [[ai-education]] - [[generative-ai]] ## Connected Articles - [[mechanical-engineering-ai-curriculum-2026]] — Project-Based AI Education Curriculum in Thermal Engineering - [[pbl-biomedical-engineering-genai-2026]] - [[engineering-faculty-metaphors-ai-understanding-2026]] — Engineering Faculty Metaphors Construct (and Constrain) AI Understanding - [[tam-critical-use-genai-engineering-2026]] — Extended TAM with critical use for engineering/CS students - [[socio-cognitive-genai-adoption-engineering-2026]] — Unified socio-cognitive model for engineering education (Bangladesh) - [[ethical-use-ai-engineering-education-review-2026]] — Ethical Use of AI in Engineering Education: A Systematic Review - [[multimodal-embodied-cognition-oral-explanations-2026]] — A Multimodal Framework for Embodied Cognition in Oral Explanations - [[ai-engineering-computing-workforce-grey-literature-2026]] — AI and the Future of the Engineering and Computing Workforce - [[ai-engineering-education-balancing-act]] — Using AI in Engineering Education: A Balancing Act - [[ai-learning-tools-engineering-education-needs]] — Designing Needs- and Attention-Aware AI Learning Tools for Engineering Education - [[structured-ai-demonstrations-engineering-mechanics]] — Structured AI Demonstrations in Engineering Mechanics - [[isaza-chatgpt-engineering-prompting-2026]] — ChatGPT in engineering education - [[liu-ai-sustainable-engineering-education-2026]] — AI-SEE framework for sustainable engineering education (Liu et al. 2026) - [[yin-arthur-ai-teaching-assistant-engineering-econ-2026]] - [[genai-xr-architectural-design-education-2026]] — Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study --- ## [STEM Education](https://edtechdev.github.io/aied/concepts/stem-education/) > **STEM Education** — science, technology, engineering, and mathematics education is the most common domain for [[ai-education|AI in education]] [[research-methods-aied|research]] in the knowledge base. STEM's structured knowledge, clear right/wrong answers, and computational nature make it an ideal testbed for [[intelligent-tutoring|AI tutoring]] and assessment. ## Questions to Consider - STEM is the most common domain for AI-in-education research because its knowledge is structured and has clear right/wrong answers. Do you think that makes STEM the easiest place to teach with AI — or possibly the place where AI's limits are most easily masked? - The page cites research showing AI adoption in schools is 'stratified by discipline' — normalized in computer science, heavily prohibited in mathematics. Why do you think subject culture shapes AI acceptance so strongly, and what are the consequences for students? - If a math student uses AI mainly to check solutions and get explanations, is that a scaffold or a crutch? What determines the difference, and where would you draw the line? - Given that STEM problems often have verifiable answers, what kinds of AI use in STEM do you think genuinely build understanding versus merely produce correct-looking output? - How might the very qualities that make STEM ideal for AI tutoring — clear answers, computable correctness — undersell the parts of science and engineering that are messy, open-ended, and judgment-based? ## Introduction ### STEM as the primary AIED domain - **Mathematics:** [[math-education|Math education]] research spans [[generative-ai-reduced-study-time-math|GenAI impact on math learning]], [[ai-powered-personalized-learning-elementary-fractions-2026|elementary fraction tutoring]], and [[student-math-competence-clustering|competence clustering]]. - **Physics:** [[physics-education|Physics education]] includes [[becker-chatgpt-typology-physics-2026|ChatGPT typology studies]], [[hashmi-socratic-physics-chatbot-2025|Socratic physics chatbots]], and [[ai-scoring-language-bias-physics|scoring bias analysis]]. - **Computer science:** [[cs-education|CS education]] is the most-researched STEM subfield — [[code-review-genai-cs1|code review]], [[debugtracker-classroom-debugging|debugging tools]], and [[prompt-problems-nl-programming-mistakes|prompting studies]]. - **Engineering:** [[concept-catalyst-engineering-scaffolds|Engineering scaffolds]], [[structured-ai-demonstrations-engineering-mechanics|mechanics demonstrations]], and [[ai-engineering-education-balancing-act|curriculum balancing]] bring AI to [[engineering-education|engineering education]]. - **Scaffolding undergraduate research:** [[ai-information-extraction-undergraduate-thesis-2026|An and colleagues (2026)]] pilot an AI system that converts research publications into structured, comparable datasets for undergraduate thesis completion across four STEM schools. Results (20 students, 80 documents) showed >90% extraction of experimental parameters, ~65% reduction in literature-review time, and a 50% increase in students' ability to identify influential experimental variables — evidence of how [[generative-ai|AI]]-scaffolded [[higher-ed|undergraduate research]] can strengthen research literacy and epistemic cognition in STEM education. - **Simulation-supported instruction:** In drone-based STEM education, [[teacher-role|teacher]]-AI co-designed [[simulation]] scaffolds were evaluated against an identical hands-on [[curriculum-design|curriculum]] with 30 secondary students, testing whether simulation-supported instruction yields superior [[learning-gains|learning outcomes]]. GenAI's role was to accelerate content creation while teacher involvement preserved pedagogical validity and contextual relevance. - **AI-assisted planning in STEAM arts education:** An experimental study of children's STEAM arts teachers ([[luo-tahir-chatgpt-steam-lesson-planning-2026|Luo and Tahir 2025]]) found ChatGPT-assisted lesson plans outperformed teacher-generated ones on expert-rated quality (median 20.5 vs. 17.6, p = .002, large effect, six professor raters), with teachers reporting gains in efficiency and interdisciplinary integration (61% rated 4+). The result came through the teacher's delegation method — most usefully by having ChatGPT fill content gaps in a self-outlined lesson — reinforcing the recurring point that AI lifts STEM/STEAM work when the teacher structures the task and critically evaluates output rather than handing the whole plan to the model. ### How effective AI-supported instruction is in STEM, pooled The largest quantitative synthesis for this domain to date — 35 experimental and quasi-experimental studies published between 2005 and 2025 ([[ai-supported-instruction-stem-meta-analysis-2026|Doğan, Kılıç, Kalınkara and Talan, 2026]]) — puts AI-supported instruction in STEM at Hedges' g = 0.670 (95% CI [0.491, 0.848]), with the between-study variance handled by a random-effects model. The level breakdown is the informative part: effects were largest in high school (g = 1.099) and progressively smaller in university (0.578), elementary (0.465) and [[k-12|middle school]] (0.392), while the subject-area differences that STEM's self-image would predict — science (0.676) and mathematics (0.650) ahead of technology and engineering (0.501) — were not statistically significant (Q = 4.85, df = 2, p = 0.088). Duration did not behave as a dose: the strongest band was one to two months (g = 0.833), the shortest interventions of five hours or less still reached 0.621, and the weakest band (g = 0.256) was not significant. Read alongside the skepticism documented on [[learning-gains|Learning Gains]], the reasonable reading is that AI-supported STEM instruction produces a moderate, real but level-dependent effect, not a uniform one. ### Why STEM dominates STEM's structured knowledge representation, verifiable answers, and computational thinking alignment make it the most natural fit for AI tutoring. [[computational-thinking|Computational thinking research]] explores this alignment explicitly. ### New evidence from 2025–26 IJ STEM Education research A concentrated batch of 2026 *International Journal of STEM Education* studies sharpens how AI functions across STEM's subfields and levels: - **[[discipline-specific-aied|Subject-specific]] [[governance]] shapes [[student-ai-interaction|student AI use]].** A cross-sectional study of 416 Czech secondary students ([[lnenicka-secondary-students-genai-stem-2026]]) found AI adoption is *stratified by discipline* rather than unified: computer science and economics normalize [[generative-ai|GenAI]] as a collaborative resource, while mathematics (65.9% prohibit) and natural sciences (55.3%) show high perceived prohibition co-occurring with poor rule clarity and persistent clandestine use. Students mostly use AI as an instrumental scaffold (explanation, solution-checking) rather than a substitute, but a [[critical-thinking|critical evaluation]] gap emerges — heavy prompt modification overshadows external factual verification, shifting behavior toward [[cognitive-offloading]]. This argues for *subject-sensitive* guidance over blanket bans. - **AI as a co-inquirer in inquiry-based STEM.** A quasi-experiment with 97 third-graders ([[dai-chatbots-problem-posing-primary-2026]]) showed GenAI [[conversational-ai|chatbots]] significantly outperformed search engines for science [[problem-based-learning|problem posing]] in [[inquiry-based-learning|inquiry-based learning]], improving question quality, producing a more integrated epistemic [[network-analysis|network structure]] (ENA), and lowering cognitive load. A [[meta-analysis-systematic-review|systematic review]] of ChatGPT for inquiry-based learning in STEAM ([[jiang-chatgpt-inquiry-steam-review-2026]], 24 studies) confirms ChatGPT supports question formulation, inquiry design, [[problem-solving]], and reflection — but risks over-reliance, [[hallucination-risk|hallucination]], and superficial conclusions when outputs are treated as authoritative. - **STEAM is an uneven pathway to AI literacy.** A PRISMA systematic review of 39 studies ([[niri-steam-ai-literacy-review-2026]]) found STEAM implementations chiefly develop technical literacies (fundamental AI concepts, computational thinking, data literacy) while underdeveloping [[ethics|ethical]] awareness, creative imagination, creating/managing/designing with AI. Technology disciplines lead; arts, engineering, and integrated STEAM lag — indicating AI literacy in STEM is currently lopsided toward technical skill over responsible shaping of AI. - **Adaptive AI-based STEM programs can support deep learning.** A cluster-randomized pilot in sixth-grade science ([[bin-bakheet-adaptive-ai-stem-deep-learning-2026]], N = 30) found an adaptive AI-based STEM program (personalized content, rule-based mastery, real-time feedback) produced large effect sizes favoring the experimental group across explanation, interpretation, application, and idea generation — though the two-classroom design warrants cautious interpretation. - **Teacher acceptance is heterogeneous and discipline-shaped.** A latent profile analysis of 128 pre-service teachers ([[chen-preservice-teachers-chatgpt-lpa-2026]]) found four ChatGPT-acceptance profiles (Pragmatic Evaluators, Technology Pioneers, Resistant Skeptics, Environmental Observers), with STEM teachers concentrated in Technology Pioneers and non-STEM teachers in resistant profiles — and Resistant Skeptics showing high ease of use but low intention, demanding differentiated [[ai-literacy]] training. - **Assessment and cognitive processes in AI-integrated STEM.** The [[zhang-ct-ai-training-test-2026|CTAT]] (34-item, IRT-validated) provides a valid instrument for assessing [[computational-thinking]] within AI-training contexts, revealing students struggle most with data representation, logical-operator sequencing, and loop structures. A grounded-theory study of AI-assisted programming ([[liu-tool-tutor-crutch-programming-2026]]) shows learners oscillate between "Domain Mastery" and "Tool Mastery" through [[scaffolding]] and Offloading loops, with attenuated [[metacognition|metacognitive]] calibration under routine offloading — a process-level account of the performance-learning tension. ## Implications for STEM instructors - **Choose discipline-appropriate AI.** STEM spans math (tutoring), physics ([[socratic-method|Socratic dialogue]], [[simulation]]), CS (code generation, review), and engineering (design, workforce) — select tools matched to each subfield's signature [[pedagogy]] rather than assuming one general chatbot fits all. - **Use AI's structured-fit advantage, but protect reasoning.** STEM's verifiable answers make it the most AI-tractable domain; guard against over-reliance and answer-replacement by embedding AI in structured, mastery-oriented workflows. - **Embed AI literacy across STEM courses.** Studies ([[zha-ai-literacy-biology-case-study|biology]], [[ai-tpack-preservice-math-teachers|math teacher prep]]) show STEM context supports AI learning — integrate AI concepts where they naturally arise rather than isolating them. - **Watch [[equity-in-ai-education|equity]] and access in AI adoption.** STEM AI tools are not neutral; monitor scoring bias, [[digital-divide]] access, and [[culturally-relevant-pedagogy|culturally relevant]] design as you deploy them. ## Connected Concepts - [[learner-identity]] — evolving disciplinary, professional, creative, and academic learner identities - [[business-education]] - [[cs-education]] - [[math-education]] - [[physics-education]] - [[computational-thinking]] - [[k-12]] - [[higher-ed]] - [[intelligent-tutoring]] - [[automated-assessment]] - [[formative-assessment]] - [[personalized-learning]] - [[llm]] - [[discipline-specific-aied]] - [[teacher-education]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools ## Connected Articles - [[omniphys-multimodal-physics-benchmark-2026]] - [[mechanical-engineering-ai-curriculum-2026]] — Project-Based AI Education Curriculum in Thermal Engineering - [[ai-pedagogical-accompaniment-amico]] — AI-enabled pedagogical accompaniment supporting STEM identity - [[lnenicka-secondary-students-genai-stem-2026]] — What secondary students actually do with GenAI tools across STEM - [[dai-chatbots-problem-posing-primary-2026]] — GenAI chatbots and problem posing in primary science - [[jiang-chatgpt-inquiry-steam-review-2026]] — ChatGPT for inquiry-based learning in STEAM - [[niri-steam-ai-literacy-review-2026]] — STEAM education for AI literacy: systematic review - [[bin-bakheet-adaptive-ai-stem-deep-learning-2026]] — Adaptive AI-based STEM program for deep learning - [[chen-preservice-teachers-chatgpt-lpa-2026]] — Pre-service teacher ChatGPT acceptance profiles - [[zhang-ct-ai-training-test-2026]] — Computational Thinking in AI Training Test (CTAT) - [[liu-tool-tutor-crutch-programming-2026]] — Tool, tutor, or crutch: grounded theory of AI-assisted programming - [[workforce-readiness-smart-manufacturing-wrl-2026]] — Workforce Readiness Level framework for smart manufacturing in the AI era - [[becker-chatgpt-typology-physics-2026]] - [[ai-powered-personalized-learning-elementary-fractions-2026]] - [[concept-catalyst-engineering-scaffolds]] - [[generative-ai-reduced-study-time-math]] - [[ai-metacognition-stem-review]] - [[li-ai-science-situated-learning-teachers-2025]] - [[avraamidou-ai-colonization-science-education]] - [[ai-science-chemistry-education-systematic-review-2025]] — Systematic review of AI in science/chemistry education - [[astor-computational-thinking-meta-review-2026]] — CT as a 21st-century skill across STEM - [[ai-information-extraction-undergraduate-thesis-2026]] — AI-powered information extraction supporting undergraduate thesis and research-based learning (An et al. 2026) - [[simulation-assisted-drone-learning-stem-2026]] — Simulation-assisted drone learning with teacher-AI co-designed scaffolds - [[luo-tahir-chatgpt-steam-lesson-planning-2026]] - [[ai-supported-instruction-stem-meta-analysis-2026]] — Pooled effect of AI-supported STEM instruction across 35 studies, with the largest gains in high school (Doğan et al. 2026) --- ## [Science Education](https://edtechdev.github.io/aied/concepts/science-education/) > **[[stem-education|Science education]]** — the study and practice of how students learn science and how to teach it, now being reshaped by generative AI (LLMs, simulations, virtual labs, and AI grading) across [[physics-education|physics]], chemistry, and [[biology-education|biology]]. The science-education articles in this knowledge base reveal a field negotiating a core tension: AI demonstrably supports inquiry, misconception correction, and assessment at scale, yet its value depends on instructional design, and it carries real risks of [[cognitive-offloading|over-reliance]], hallucination, and dehumanized, profit-driven learning. ## Questions to Consider - The page opens with a tension: AI demonstrably supports inquiry and misconception correction at scale, yet its value depends entirely on instructional design. Before you read, when has a tool you used in science teaching seemed powerful yet pedagogically hollow — and what made the difference? - An AI-facilitated physics inquiry 'invented' mass values that no one had measured when students pressured it. What does this incident reveal about treating AI outputs as authoritative in science, and how should students be trained to treat a confident-sounding answer? - Research suggests AI may currently add more value as a scalable content generator than as an interactive tutor — in one study, well-crafted conceptual-change texts (expert or AI-made) beat a prompted AI dialogue. Does that surprise you, and what does it suggest about where the real pedagogical payoff of AI in science lies right now? - AI graded 10,364 handwritten physics assessments with high agreement with human graders, yet multimodal models still show gaps on complex reasoning and diagram generation. What is it about 'authentic, representation-heavy' science work that might resist automated assessment? - Physics students show a striking trust–utility gap — high use, low trust — and split into pragmatic users and skeptical nonusers. If students themselves are domain-calibrated skeptics, what does that imply for one-size-fits-all policies about AI in science classrooms? - A critical voice warns of 'AI colonization' of science education — extraction without consent, algorithmic monoculture, dehumanized reform. Before you read, where do you think the benefits of AI in science education are being oversold, and what would a human-centered, justice-oriented alternative look like? ## Introduction Science education is where AI's promise and its limits collide most visibly, because the disciplines demand rigorous [[multimodal]] reasoning — visual-spatial thinking in physics, laboratory skills in chemistry, and organism-level systems in biology — alongside well-structured, verifiable content that LLMs handle well. The eighteen articles synthesized here span all three disciplines and levels, from [[k-12]] secondary classrooms to [[higher-ed]] university courses and [[teacher-education|pre-service teacher]] programs. Across them, a consistent picture emerges: AI functions best not as an answer generator but as an embedded partner — a co-inquirer, a content generator, a virtual lab assistant — whose contribution is decided by [[learning-design|instructional design]] and the surrounding pedagogical structure. ### Virtual labs and simulations AI is dramatically lowering the barrier to creating customized, embodied science instrumentation. [[genai-ar-physics-simulation-prompt-2026|Levy et al.]] show that a four-element natural-language prompt can generate a hand-controlled augmented-reality physics simulation (a pinch-and-spread gesture tunes a virtual lamp's wavelength), with 86% of pilot students reporting greater [[student-engagement|engagement]]. Similarly, [[ai-generated-smartphone-circular-motion-lab-2026|Suñer et al.]] generated a browser-based rotation laboratory entirely through prompting, validated against video analysis to better than 1%. These works reframe [[generative-ai|generative AI]] as a [[prompt-engineering|programming]] tool that lets teachers design software around [[pedagogy|pedagogical intent]] rather than adapting activities to fixed apps — advancing [[embodied-learning|embodied]] and [[simulation]]-based learning. In [[chemistry-education]], [[context-based-ai-secondary-chemistry-2026|Abdikayumova & Madybekova]] embedded PhET simulations and ChatGPT tutoring within a context-based 7E inquiry cycle for Grade 10 students, achieving significantly higher achievement and engagement than either component alone — evidence that contextualization, structured inquiry, and adaptive AI act synergistically. Behavioral analysis of how learners work inside such modeling environments reinforces this design lesson: tracing 315 online learners building ecological models in VERA, [[an-goel-self-directed-modeling-2026|An, Hammock & Goel (2025)]] found that learners who engage in full-cycle, hypothesis-driven exploration build the most complex and diverse models, whereas observation-heavy learners largely copy existing models — suggesting scientific modeling tools should actively promote full-cycle exploratory behavior rather than leave learners in a passive, observation-dominated mode. ### AI tutors and inquiry-based learning Inquiry-based learning is a central theme. [[jiang-chatgpt-inquiry-steam-review-2026|Jiang et al.]]'s [[meta-analysis-systematic-review|systematic review]] of 24 studies positions ChatGPT as an "AI-powered co-inquirer" used mainly in the conceptualization, investigation, and discussion phases of [[inquiry-based-learning|STEAM inquiry]], improving performance, [[critical-thinking|critical thinking]], and engagement — yet risks over-reliance, hallucination, and superficial conclusions when outputs are treated as authoritative. [[ai-supported-inquiry-photosynthesis-respiration-2026|Aydın]] found that eight weeks of AI-supported guided inquiry produced large gains in pre-service teachers' conceptual understanding of photosynthesis and respiration, while [[ai-literacy|AI literacy]] and [[computational-thinking|computational thinking]] showed no significant change — suggesting disciplinary learning can improve even when broader competencies need longer or more explicit instruction. [[embodied-inquiry-ai-facilitator-physics-2026|Tufino & Damiani]] show AI can scaffold the epistemic core of ISLE inquiry while remaining confined to the verbal channel, but found facilitation fragile: under student pressure, the AI "invented" mass values that no one had measured, underscoring [[hallucination-risk]] and the need to verify rather than assume AI fidelity. [[airis-cognitively-activated-ai-physics-2026|Kuhn et al.]] propose the AIRIS framework (Activate–Inquire–Reflect) to keep prediction, interpretation, and evaluation non-delegable human tasks, and call for "withdrawal condition" experiments testing whether learning survives AI's removal — a direct response to the "boiling frog problem" of eroding epistemic practice. Expert judgment of AI-generated lesson plans extends this design lesson to [[curriculum-design|curriculum]]-aligned science: [[karaismailoglu-ai-lesson-plans-science-experts-2026|Karaismailoglu, Surmeli and Yildirim (2026)]] had eleven Turkish science-education specialists score ChatGPT-4 and an education-focused tool (Teacher's Buddy) on sixth-grade plans for a "Sustainable Living and Biodiversity" unit aligned to the Engineering [[design-based-research|Design-Based]] Learning (EDBL) model. Both platforms scored highest on "Presenting the solution" yet lowest on Engineering Skills (general-purpose x̄ = 37.45; education-focused x̄ = 41.45) — a shared weakness in simulating the iterative, process-oriented stages (prototyping, failure analysis, revision, reflective retesting) that define design-based learning — and only 3 of 11 experts judged either plan directly "Applicable," with 7 rating them "Can be applied by correction." ### Misconceptions and conceptual change On conceptual change, the evidence is nuanced and partly counterintuitive. [[akdogan-heat-temperature-conceptual-change-thesis-2025|Akdoğan]] found in a 413-student Solomon Four-Group design that both expert-written and AI-generated [[refutation-text|conceptual change texts]] were equally and significantly more effective at reducing heat-and-temperature [[misconceptions]] than interactive AI dialogue, which offered no advantage over control — positioning AI's current pedagogical value as a scalable [[generative-ai|content generator]] rather than an interactive tutor, with gains concentrated among high-achieving students (an [[equity-in-ai-education|equity]] concern). This is reinforced by [[probing-ai-generated-physics-solutions-2026|Borse et al.]], who show that prompt specificity shapes physics-solution quality and that MAPS-guided critique better prepares students to evaluate AI output than independent [[problem-solving|problem solving]], treating AI fallibility as a [[critical-thinking]] learning resource. ### AI grading and assessment AI grading is advancing rapidly on high-stakes work. [[ai-grading-handwritten-physics-2026|Pathak et al.]] graded 10,364 scanned pages of handwritten physics assessments (including a national Olympiad and team-selection camp), achieving high score correlations (0.91–0.97) and recovering the same five-student team as human grading — arguing AI works as a valid [[assessment-validity|second reader]] and audit tool under examiner control when given detailed, physics-specific rubrics. Short-answer auto-marking has a longer and more foundational transformer lineage: [[auto-marking-short-answer-science-2026|Morley et al.'s scoping review]] of 21 studies (2017–early 2024) found BERT-family models dominated auto-marking of short-answer science questions through 2021 before GPT-based approaches were adopted via prompting from ~2022, that models augmented with domain data (textbooks, rubrics, further pre-training) consistently outperformed baselines, and that unresolved threats to [[educational-measurement|reliability]], [[explainable-ai|explainability]], and [[bias-mitigation|fairness]] argue for such auto-markers as supports to, rather than replacements for, [[teacher-role|human examiners]]. Yet [[omniphys-multimodal-physics-benchmark-2026|Chen et al.]]'s OmniPhys [[benchmark]] reveals that multimodal LLMs still show significant gaps on complex reasoning and diagram *generation*, cautioning against over-trusting [[automated-assessment|automated]] physics assessment on the authentic, representation-heavy tasks that define deep disciplinary competence. A chemistry counterpart sharpens the format-dependence: [[cvengros-grading-handwritten-chemistry-ai-2026|Cvengros & Kortemeyer]] graded a 296-student handwritten general-chemistry final with a multimodal [[llm]], finding high agreement with TAs for textual and chemical-reaction answers but *worse-than-random* performance on drawing and graphing (background grids distract AI vision) — evidence that reliable AI grading in science is selective by response format and requires confidence-based deferral to [[human-in-the-loop-ai|human examiners]] for graphical items, rather than a uniform "grade everything" approach. ### Teacher perceptions and the critical view Teacher readiness is decisive. [[pre-service-science-teachers-ai-perceptions-2026|Amponsah et al.]] found Ghanaian pre-service science teachers hold positive attitudes and strong intentions toward AI but only moderate actual classroom use — an intention–use gap pointing to [[teacher-ai-competency]] and [[governance|institutional]] support as the real levers. [[becker-chatgpt-typology-physics-2026|Becker et al.]] and [[fouad-bentley-trust-utility-gap-physics-2026|Fouad & Bentley]] document that physics students are domain-calibrated skeptics, not uncritical adopters: a 50-point trust-utility gap (91% use, 41% trust) and two distinct user profiles (Pragmatic Users vs. Skeptical Non-Users) that challenge one-size-fits-all policies. Counterbalancing the techno-optimism, [[avraamidou-ai-colonization-science-education|Avraamidou]] warns of an "AI colonization" of science education — extraction without consent, algorithmic monoculture, and dehumanized, profit-centered reform — and calls for a feminist, human-centered AI prioritizing justice over profit, while [[ai-science-chemistry-education-systematic-review-2025|Erümit & Özdemir Sarıalioğlu]]'s systematic review of 18 studies emphasizes [[ethics|ethical]] risks (bias, hallucination, [[academic-integrity|plagiarism]], erosion of independent thinking) and the need for [[teacher-education|teacher training]] and conscious use. The collective lesson across all eighteen articles is that AI in science education delivers gains when embedded in sound inquiry, [[scaffolding]], and [[assessment]] design — and undercuts learning when it displaces the epistemic work students must do themselves. ## Connected Concepts - [[stem-education]] - [[physics-education]] - [[chemistry-education]] - [[biology-education]] - [[inquiry-based-learning]] - [[misconceptions]] - [[simulation]] - [[generative-ai]] ## Connected Articles - [[ai-science-chemistry-education-systematic-review-2025]] — Systematic review of AI in science and chemistry education - [[jiang-chatgpt-inquiry-steam-review-2026]] — ChatGPT for inquiry-based learning in STEAM - [[ai-supported-inquiry-photosynthesis-respiration-2026]] — AI-supported guided inquiry in photosynthesis and respiration - [[akdogan-heat-temperature-conceptual-change-thesis-2025]] — Conceptual change texts vs. AI dialogue for heat and temperature - [[ai-grading-handwritten-physics-2026]] — Large-scale AI grading of handwritten physics assessments - [[avraamidou-ai-colonization-science-education]] — Critical commentary on AI colonization of science education - [[pre-service-science-teachers-ai-perceptions-2026]] — Pre-service science teachers' AI perceptions and acceptance - [[ai-assisted-inquiry-ssi-climate]] — AI-Assisted Inquiry in Socio-Scientific Issues on Climate Change - [[auto-marking-short-answer-science-2026]] - [[an-goel-self-directed-modeling-2026]] - [[cvengros-grading-handwritten-chemistry-ai-2026]] - [[karaismailoglu-ai-lesson-plans-science-experts-2026]] - [[ai-tutoring-micro-rct-gcse-science-2026]] — Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomized Trials of an AI Tutoring Platform in GCSE Science --- ## [Writing](https://edtechdev.github.io/aied/concepts/writing-education/) > **Writing** — the use of AI tools for writing instruction, [[assessment]], [[feedback]], and the study of how [[generative-ai|generative AI]] reshapes the writing process itself. Writing education is one of the most AI-affected domains, because LLMs excel at the very activities writing instruction centers on — text generation, revision, and evaluation. [[research-methods-aied|Research]] in this area spans [[automated-assessment|automated scoring]], AI feedback quality, writing-process support, second-language writing, academic integrity, and the deeper question of how AI changes what it means to write and to be a writer. ## Questions to Consider - The page's central claim is that writing is not merely output but a cognitive, social, and rhetorical process — and that AI can displace the very mental work that makes writing a learning activity. When you write, what happens in your thinking that a finished AI-produced paragraph simply erases? - A common framing is 'AI as a tool' or, at the opposite extreme, 'AI as a threat to authorship.' The page offers a third view: writing as a human-AI entanglement where agency is distributed. Which of these framings matches your own experience of writing with or without AI — and what does each framing imply for how you'd teach? - Research found that delegating *deeper* layers of writing — reasoning and argumentative logic — harms your independent writing more than delegating surface layers like grammar. Think about your last AI-assisted piece of writing. Which layer did you actually delegate, and what does that predict about what you can now do on your own? - The page warns that AI writing feedback is not language-neutral: personalizing feedback with a student's race, language, or disability can shift it in stereotype-aligned ways — such as overpraising or withholding critique. If you've received or given 'personalized' AI feedback, how would you detect that a tool was softening its critique for some learners? - The design guidance here is 'coaching, not composing' — have AI ask questions and critique outlines, but require the learner to produce prose first. Why might letting the learner draft before the AI intervenes protect ownership and judgment in ways a tool that writes the draft could not? - One finding: students often say 'it's OK because…' to rationalize AI use, moving the issue from plagiarism policing toward ethics and AI literacy. If you were designing a writing course, how would you build honesty and [[ethics|ethical]] judgment about AI into it, rather than relying on detection or punishment? ## Introduction Writing is not merely output but a cognitive, social, and rhetorical process. This is why AI's impact on writing education is so consequential and contested: AI can be a [[scaffolding|scaffold]] that helps students draft, revise, and receive feedback they otherwise wouldn't get, but it can also displace the [[cognitive-offloading|cognitive work]] — and the human audience — that make writing a learning activity. The knowledge base's research consistently frames AI in writing as a *human-centered complement* to, rather than a replacement for, the social and cognitive processes of writing. Where the focus is **English specifically** — [[english-education|English for Academic Purposes (EAP)]] and English language [[teacher-role|teaching]] (EFL/ESL/L2) — see the dedicated [[english-education]] concept page, which distinguishes English-specific and academic-register research from general writing and general language learning. ### How AI in writing education appears in the research - **Automated essay scoring:** [[automated-essay-scoring]] systems like [[choi-anchor-aes-prompting-2025|anchor-based AES]] and [[aiawe-automated-writing-evaluation|AIAWE]] evaluate student writing at scale, raising questions about [[assessment-validity|construct validity]] and the reduction of writing to measurable features. For short argumentative writing (about 150–200 words) in Spanish, AI scoring agreement varies sharply by rubric dimension ([[human-in-the-loop-ai-scoring-national-assessment-2026|Curi et al., 2026]]): structure- and register-oriented items — introduction, conclusion, register, nominal agreement, subject–verb agreement — reached moderate chance-corrected agreement with human raters in the 2025 edition, while micro-level linguistic items (vocabulary, syntax, punctuation, connectors, argumentation) stayed in the fair-to-slight range. This suggests LLM assistance is most defensible for macro-level discourse features, and that low-level language conventions should keep deterministic tooling or [[human-in-the-loop-ai|human review]]. - **Writing feedback:** [[ai-feedback-quality|AI feedback quality]] research ([[genai-teacher-feedback-comparison|GenAI vs. teacher feedback]], [[care-full-feedback-genai|care-full feedback]], [[repeated-ai-writing-feedback-semester|repeated AI feedback]]) examines whether AI feedback improves writing and how it compares to human feedback. The PAIRR model ([[pairr-ai-peer-review-2025|Peer and AI Review + Reflection]]) combines AI with [[peer-assessment]] and finds AI feedback is most useful in a human-centered process. - **Writing process support and agency:** [[agency-gap-ai-writing|Agency gap research]] and [[ai-writing-support-stage-ownership-2026|stage-ownership research]] explore how AI changes the writing process from planning to revision, and how students' [[agency]] is affected when AI participates at different stages. - **Posthumanist perspectives:** [[posthumanist-ai-literacy-2025|A posthumanist approach to AI literacy]] reframes writing as a human-AI entanglement in which [[agency]] is distributed, challenging both uncritical anthropomorphization of AI and its dismissal as a mere tool — a relational rather than transactional view of AI literacy. - **L2 / [[multilingual-learning|multilingual]] writing:** [[self-referential-l2-writing-llm-assessment|L2 writing assessment]], [[genai-linguistic-diversity-academic-writing|linguistic diversity research]], and [[ai-writing-support-stage-ownership-2026|stage-ownership research]] address how AI supports (or constrains) second-language and multilingual writers, including the risk of reinforcing Standard Academic English norms. The knowledge base's youngest L2 sample comes from a nine-week GenAI-supported opinion-writing program with 301 Grade 5 and 6 students in Eastern China ([[genai-writing-program-primary-l2-motivation-engagement|Lu et al., 2026]]), which raised ideal L2 writing self and academic buoyancy and improved rubric-scored language use while leaving organization and total scores unchanged — AI support moved specific dimensions of writing rather than writing ability as a whole. - **Bias in personalized feedback (Marked Pedagogies):** [[marked-pedagogies-linguistic-bias-writing-feedback|Tan et al. (2026)]] show that [[llm]] writing-feedback tools are not language-neutral: personalizing feedback with a student's race, ethnicity, ELL designation, learning disability, achievement, or motivation systematically shifts feedback in stereotype-aligned ways — including positive feedback bias and feedback withholding bias (overuse of praise, less substantive critique, assumptions of limited ability) for students marked by race, language, or disability, even when the essay is identical. This makes "[[personalized-learning|personalization]]" itself a bias vector that writing-feedback tools must audit and control. - **Academic integrity:** Survey evidence complicates the policing frame directly: among 504 sociology students ([[student-genai-use-views-writing|Kuznetsov et al., 2026]]), 65 percent had used GenAI for coursework but only 3 percent to generate assignment text and 2 percent to produce a full draft, while fear of an academic offense was the second most common concern (28 percent) and roughly a quarter reported no guidance at all (19 percent) or guidance they found unclear. On this evidence the writing-education problem is ambiguity about permitted use, not widespread text generation. [[nash-preservice-teachers-classroom-ai-policies-2026|Nash and Burriss (2026)]] show how such ambiguity is produced at the classroom level. Coding 27 [[teacher-education|preservice]] English language arts teachers' own classroom [[educational-policy-ai|AI policies]], they found that 26 of 27 permitted some [[generative-ai|generative AI]] use but overwhelmingly on teacher-specified terms, with 22 of 27 allowing AI for ideation and brainstorming while disallowing AI composition of sentences, paragraphs or papers. Limits were rarely operationalized — one participant allowed AI "to get your thinking started" and declared "this is where the line should be drawn" without saying where — and many policies simultaneously prohibited submitting AI text and held students responsible for the AI text they submitted, a contradiction that leaves students unable to comply. Twenty-two of the 27 policies were silent on reading altogether, ceding AI-supported comprehension work to no guidance at all. ### Writing as thinking Because writing is a cognitive process, AI-in-writing research connects to [[cognitive-offloading]] (does AI writing support bypass thinking?), [[metacognition]] (does AI feedback improve self-assessment?), [[self-regulated-learning]] (do students regulate their use of AI feedback?), and [[ai-literacy]] (can students evaluate AI-generated writing critically?). The [[critical-thinking-genai-scaffolding|critical-thinking scaffolding]] and [[ai-feedback-critical-thinking-writing-2026|AI feedback for critical thinking]] research show that the [[pedagogy|pedagogical]] value of AI in writing depends on whether it prompts reflection and judgment rather than answer-replacement. [[layer-sensitive-cognitive-offloading-writing-2026|Chen (2026)]] sharpens this with a **layer-sensitive** account of [[cognitive-offloading|cognitive offloading]] in GenAI-assisted academic writing: delegating *deeper* layers (reasoning, argumentative logic) carries a stronger negative association with independent no-AI writing quality and [[critical-thinking|higher-order thinking]] than delegating surface layers (grammar, vocabulary). Open AI collaboration yielded the best supported product but the worst independent outcomes, while bounded support with reflection preserved competence — evidence that GenAI writing support is not uniformly harmful but its effect depends on which cognitive layer students delegate. [[lu-ai-multimodal-writing-critical-thinking-2026|Lu et al. (2027)]] extend this thinking to *multimodal* composing by younger writers. Having 60 [[k-12|Grade 5]] students externalize their narratives as AI-generated images and short videos produced sustained [[self-report-measures|self-reported]] gains in interpretation, analysis, evaluation, and explanation — the facets multimodal resemiotisation exercises — but **no gain in inference**. Making meaning visually explicit lowered the demand to infer implicit meaning from text, exactly the offloading mechanism Chen describes; only structured peer discussion restored occasions for inference. The study cautions that [[multimodal|multimodal AI]] composing helps young writers reflect on clarity and coherence while potentially skimming off the inferential work that text-only writing preserves — a design consideration for writing instructors pairing AI [[visualization|visuals]] with peer [[peer-assessment|feedback]]. A dimension-specific pattern recurs across this literature, and it is a useful diagnostic. In the primary-level L2 writing program, emotional and behavioral [[student-engagement|engagement]] rose while cognitive and metacognitive engagement did not, and the authors name diminished self-monitoring during writing as an explicit risk of GenAI support ([[genai-writing-program-primary-l2-motivation-engagement|Lu et al., 2026]]). Enjoyment and on-task activity are therefore not evidence that deeper processing is happening — the same distinction the offloading research draws when it asks which layer of cognitive work a student has delegated. ### Designing AI writing support: coaching, not composing Because writing has no single correct answer, AI writing tools require a different design from answer-verifiable tutors. The knowledge base's design guidance (see the worked **AI writing coach** example in the FAQ on [[developing-ai-tutor|Designing an AI Tutor]]) centers on preserving authorship and [[evaluative-judgment|evaluative judgment]] rather than producing finished text: - **Track writing capabilities, not just essay scores.** A writing coach's [[student-modeling|learner model]] can track argument (thesis specificity, claim–evidence alignment, counterargument), organization, evidence integration, revision, and style — so feedback targets capabilities that persist across essays. - **Ground feedback in the assignment.** Retrieve the actual prompt, rubric, course readings, citation and genre conventions, and AI-use policy so feedback references the specific assignment rather than inventing generic expectations. - **Treat writing stages differently.** AI involvement at planning reduces perceived ownership less than at drafting, and AI-generated drafting produces the largest ownership decrease. So a coach can ask questions and critique outlines at planning while requiring the learner to produce prose first at drafting. - **Make feedback prioritized and reflective.** Each round can offer one strength to preserve, one high-impact issue, one question requiring the writer's judgment, and one concrete revision goal — and the coach should ask learners to evaluate whether they agree with a suggestion, developing [[feedback-literacy|evaluative judgment]] rather than obedience. - **Preserve authorial voice and guard against homogenization.** The coach should distinguish errors, clarity issues, rhetorical choices, and style preferences — and not automatically "correct" the latter, especially for [[multilingual-learning|multilingual]] writers and non-standard rhetorical styles. This coaching-not-composing stance is the writing-domain expression of the knowledge base's overarching "[[coach-not-crutch-ai-writing|coach over crutch]]" boundary: identify the cognitive activity that produces learning (planning, drafting, evaluating, revising) and design the AI to support it without taking it away from the learner. ### Connections Writing education connects to [[automated-essay-scoring]], [[ai-feedback-quality]], [[academic-integrity]], [[cognitive-offloading]], [[ai-literacy]], [[language-learning]], [[formative-assessment]], [[peer-assessment]], [[metacognition]], [[self-regulated-learning]], and [[higher-ed]]. It is a domain where AI's capabilities and risks are both highly visible, making it a rich site for studying how AI transforms pedagogy, [[assessment]], and the very nature of authorship and [[agency]]. ## Implications for writing instructors - **Frame AI as a complement, not a replacement, for the writing process.** The knowledge base's research consistently treats AI as a scaffold for drafting, revision, and feedback while protecting the cognitive work and human audience that make writing a learning activity — [[coach-not-crutch-ai-writing|coaching over crutch]]. - **Use AI feedback within a human-centered process.** [[pairr-ai-peer-review-2025|PAIRR]] finds AI feedback is most useful combined with peer review and reflection; design feedback loops that keep the instructor and peer audience central. - **Audit automated feedback for bias.** [[marked-pedagogies-linguistic-bias-writing-feedback|Marked Pedagogies]] shows LLM feedback shifts in stereotype-aligned ways when personalized with student attributes — monitor for positive/withholding bias, and be explicit that personalization can be a bias vector. - **Guard the cognitive work of writing.** Watch for [[cognitive-offloading|over-reliance]] that bypasses planning, revision, and self-assessment; use AI at chosen stages ([[ai-writing-support-stage-ownership-2026|stage-based ownership]]) to protect student agency. - **Design for empowerment rather than enforcement.** A PLS-SEM study of 327 Chinese EFL undergraduates ([[empowerment-ai-assisted-deep-revision-efl-writing-2026|Li & Zhang, 2026]]) tested the two levers writing instructors actually hold and found only one of them works. AI [[prompt-engineering|prompting]] literacy strongly predicted perceived competence, psychological safety, and [[motivation|intrinsic motivation]], and all three psychological needs partially mediated its link to deep revision engagement, with intrinsic motivation the strongest single driver of deep revision. External mandates had no direct effect at all. The practical translation is that requiring deep revision does not produce it — the enforcing requirement may be needed to make revision happen at all, but the depth comes from building students' prompting capability and the intrinsic motivation and psychological safety that follow from it. - **Address academic integrity constructively.** Move from policing AI use toward building [[ai-literacy]] and ethical-use framing that lets students use AI without unintentional misconduct. ## Connected Concepts - [[automated-essay-scoring]] - [[ai-feedback-quality]] - [[academic-integrity]] - [[cognitive-offloading]] - [[ai-literacy]] - [[language-learning]] - [[higher-ed]] - [[metacognition]] - [[llm]] - [[generative-ai]] - [[formative-assessment]] - [[peer-assessment]] - [[self-regulated-learning]] - [[student-experience]] - [[feedback-literacy]] - [[feedback]] - [[discipline-specific-aied]] - [[english-education]] - [[assessment]] - [[agency]] ## Connected Articles - [[nash-preservice-teachers-classroom-ai-policies-2026]] — Preservice teachers' classroom AI policies: tensions and entanglements - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment - [[llm-comparative-judgment-writing-screening-2026]] — Validity of Large Language Model Comparative Judgment for Universal Writing Screening - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[your-brain-on-chatgpt-cognitive-debt-essay-writing]] - [[benali-genai-academic-writing-2026]] - [[coach-not-crutch-ai-writing]] — AI writing tools can improve writing skill despite reducing effort (Lira et al. 2025) - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans (Mathew et al. 2026) - [[pairr-ai-peer-review-2025]] — Peer and AI Review + Reflection (PAIRR) - [[posthumanist-ai-literacy-2025]] — A Posthumanist Approach to AI Literacy - [[choi-anchor-aes-prompting-2025]] — Anchor-Based Automated Essay Scoring - [[aiawe-automated-writing-evaluation]] — AIAWE: Automated Writing Evaluation - [[agency-gap-ai-writing]] — The Agency Gap in AI-Supported Writing - [[ai-writing-support-stage-ownership-2026]] — From Planning to Revision: AI Writing Support at Different Stages - [[genai-teacher-feedback-comparison]] — Comparing Generative AI and Teacher Feedback - [[student-rationalization-ai-writing]] — "It's OK Because...": The Wild West of Student Rationalization - [[care-full-feedback-genai]] — Care-Full Feedback Approaches - [[self-referential-l2-writing-llm-assessment]] — Self-Referential L2 Writing Assessment - [[becerra-aicofe-feedback-2026]] — AI Peer Feedback Systems - [[repeated-ai-writing-feedback-semester]] — Student Evaluation of Repeated AI Feedback - [[elementary-writing-genai-systematic-review-2026]] — Rethinking Elementary Writing Instruction - [[veriforge-narrative-drafting-scaffolding-2026]] — VeriForge: Narrative Drafting Scaffolding - [[ai-feedback-critical-thinking-writing-2026]] — Using AI-Generated Feedback to Improve Critical Thinking - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Marked Pedagogies: linguistic biases in personalized automated writing feedback - [[zuo-instructor-power-genai-writing-2026]] — Power relations perceived by college instructors grappling with GenAI in writing (Zuo, Xu & Dunning 2026) - [[bassett-ai-detectors-education-2026]] — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026) - [[academic-erasure-complexity-ai-writing-2026]] — Academic erasure: the disappearance of complexity under AI-supported writing - [[making-ai-annoying-constrained-writing-2026]] — Making AI annoying on purpose: constraint in AI-supported writing (Konradt, Boote & Taub 2026) - [[lu-ai-multimodal-writing-critical-thinking-2026]] — Multimodal AI composing and critical thinking in primary writing (Lu et al. 2027) - [[student-genai-use-views-writing]] — Student use of and views on GenAI for writing (Kuznetsov, Sheely & Baker 2026) - [[genai-writing-program-primary-l2-motivation-engagement]] — A GenAI-supported writing program for primary L2 motivation, engagement and performance (Lu et al. 2026) - [[automated-scoring-marketing-posts-agreement-2026]] — Agreement and error in automated scoring of student marketing posts - [[empowerment-ai-assisted-deep-revision-efl-writing-2026]] — Prompting literacy and intrinsic motivation drive deep revision while external mandates have no direct effect (Li & Zhang 2026) - [[swim-student-writing-simulation-2026]] — student writing simulation conditioned on trait-level proficiency profiles --- ## [Language Learning](https://edtechdev.github.io/aied/concepts/language-learning/) > **Language Learning** — the study of how AI supports second language (L2) acquisition, writing development, and linguistic diversity in educational settings. [[ai-education|AI in education]] [[research-methods-aied|research]] in this knowledge base spans AI interlocutors for spoken dialogue, [[automated-essay-scoring|automated writing evaluation]] for L2 learners, reading support, and concerns about language bias in AI scoring systems. ## Questions to Consider - Language is inherently interactive, which makes it well-suited to [[conversational-ai|conversational AI]] — but AI's linguistic capabilities also raise risks of bias against non-native patterns. Where have you seen this tension between opportunity and risk play out? - One study found AI scoring systematically underestimates linguistically weaker students, while another proposed comparing students to their own prior work rather than native-speaker norms. How does the 'reference point' for evaluation change whether [[ai-feedback-quality|AI feedback]] helps or penalizes a learner? - AI interlocutors can extend communicative practice at scale, but the page warns they should pair with human interaction so fluency transfers to real conversation. What might you gain from practicing with an AI that a human partner can't give — and what would you lose? - Teacher support — not just the AI tool — was shown to drive engagement in AI-assisted language learning through students' achievement goals. How does the social and [[pedagogy|pedagogical]] context shape whether learners keep engaging with an AI practice tool? - A [[meta-analysis-systematic-review|meta-analysis]] found small-to-moderate, level-dependent gains from emerging tech, with productive skills (speaking, writing) gaining more than receptive ones. Why might speaking and writing benefit more than listening and reading from AI tools? - If AI privileges standard English and can penalize non-native or diverse language patterns, how should language instructors design evaluation and feedback so AI supports linguistic diversity rather than erasing it? ## Introduction Language learning has emerged as a significant AI in education domain because language is inherently interactive — making it well-suited to conversational AI — and because AI's linguistic capabilities raise both opportunities ([[personalized-learning|personalized language practice]] at scale) and risks (systematic bias against non-native language patterns). The articles in this knowledge base explore both sides of this equation. Where the target language is **English specifically** — especially [[english-education|English for Academic Purposes (EAP)]] and EFL/ESL/L2 English [[teacher-role|teaching]] — see the dedicated [[english-education]] concept page, which distinguishes English-specific research from general L2 acquisition and general writing. **AI as language tutor and interlocutor** is the most developed theme. **[[ai-interlocutor-l2-spoken-dialogue|What Changes When the Interlocutor Is an AI?]]** examines interactional fluency and linguistic uptake when L2 learners converse with AI versus humans. **[[tact-pedagogically-adaptive-esl-tutoring|TACT]]** provides pedagogically adaptive ESL tutoring. **[[llm-children-reading-story-generation]]** explores AI-generated stories for children's reading development. These connect to [[intelligent-tutoring]] and [[generative-ai]]. **[[llm-agents-5e-esl-grammar-2026|Yang, Weng and Yang (2026)]]** designed two [[llm]]-based agents — a conventional AI English teacher and one using the **5E framework** (engage, explore, explain, elaborate, evaluate) for inquiry-based grammar learning. Across **37 ESL students** in a randomized comparison, **high-performing students responded positively to the AI teacher** while low-performing students showed mixed attitudes, and the conditions differed in intrinsic motivation, cognitive change, and performance — indicating that LLM-agent design should be matched to learner proficiency. **AI in language assessment** is emerging as LLMs support [[automated-question-generation|item generation]] and evaluation. **[[gpt-item-generation-l2-listening-2026|Aryadoust and Wong (2026)]]** compared [[prompt-engineering|prompt engineering]] against fine-tuning for automatic item generation in L2 listening assessment: iterative prompt refinement improved item quality but plateaued, while **fine-tuning GPT-4.1 on the optimized prompt** (holding prompt design constant) yielded further gains — a template for when assessment developers should invest in model adaptation over prompt iteration. **Automated writing evaluation for L2 learners** evaluates AI's ability to assess non-native writing. **[[self-referential-l2-writing-llm-assessment|Bannò et al.]]** proposed a self-referential approach comparing student writing to their own prior work rather than native-speaker norms. **[[ai-scoring-language-bias-physics|Feser & Tschisgale]]** found AI scoring systematically underestimates linguistically weak students — a finding that connects to [[assessment-validity]] and [[bias-mitigation]] concerns. **[[genai-linguistic-diversity-academic-writing]]** explores how AI affects linguistic diversity in academic contexts. **[[accessibility]] for language learners** connects to [[inclusive-learning]]: **[[dyslexlens-dyslexic-learners-ai|DysLexLens]]** analyzed how dyslexic learners use AI for literacy support, and **[[ai-tools-arab-english-classrooms]]** explored AI tools in Arabic-English classroom contexts. These studies connect language learning to [[equity-in-ai-education]] and [[special-education]]. **Motivational mechanisms in AI-assisted language learning** examine why learners engage with AI for language practice. **[[wang-goal-setting-ai-engagement-2026|Wang & Wang (2026)]]** used goal-setting theory with 758 Chinese university English learners to show that **teacher support** enhances engagement in AI-assisted learning through students' mastery-approach and performance-approach goals (not avoidance goals) — evidence that the pedagogical and social context, not just the AI tool, determines whether learners stay engaged with AI-assisted language practice. This connects language learning to [[motivation]] and [[student-engagement]]. **GenAI-supported writing at the primary level.** [[genai-writing-program-primary-l2-motivation-engagement|Lu et al. (2026)]] ran a nine-week opinion-writing program with 301 Grade 5 and 6 learners in Eastern China, with eight intact classes randomly assigned to the program or to conventional instruction. The program raised learners' ideal L2 writing self (adjusted mean difference 0.20) and academic buoyancy (0.17), and lifted behavioral and emotional engagement, but it did not move growth mindset, cognitive or metacognitive engagement, or rubric-scored organization — among writing dimensions only language use improved. Two features of the design matter for language teachers: prompting was taught explicitly, through a categorized bank of prompts tied to specific writing goals, and GenAI feedback was used alongside comparison with teacher feedback and repeated revision. The authorship gains learners reported rested on that instructional structure rather than on the tool by itself, and the authors name reduced [[metacognition|self-monitoring]] and shortcut-oriented strategies as the standing risks. ## Implications for language instructors - **Emerging [[ai-technologies|technologies]] yield small-to-moderate, level-dependent gains.** A [[liu-emerging-tech-tefl-review-2026|meta-analysis of 33 TEFL studies]] (N = 3,181) finds an overall effect of Hedges' g = 0.38 that rises with educational level (primary 0.29, secondary 0.35, tertiary 0.44), with VR/AR yielding the largest effects and productive skills (speaking, writing) gaining more than receptive skills — supporting the use of emerging tech, especially at tertiary level, while keeping expectations realistic. - **Use AI to extend communicative practice, not replace it.** [[ai-interlocutor-l2-spoken-dialogue|AI interlocutors]] and [[tact-pedagogically-adaptive-esl-tutoring|adaptive ESL tutors]] expand interactional practice at scale — pair them with human interaction so fluency and uptake transfer to real conversation. - **Prioritize feedback quality over quantity in ASR-supported speaking.** [[asr-english-speaking-feedback-metacognition-2026|Chen et al. (2026)]] find that accurate error correction and structured reflection tasks improve [[feedback]] internalization and reflective behavior in college English speaking, while frequent ASR use and recognition accuracy boost motivation or reflection only partially — technical precision alone does not drive deeper cognitive engagement, and language proficiency moderates the gains (stronger learners internalize feedback more effectively). This argues for pedagogically sound feedback (e.g., articulatory explanations over simple error flags), scaffolded reflection, and proficiency-differentiated support. - **Support learners' psychological adaptation to AI-assisted study.** [[wu-psychological-adaptation-ai-japanese-learning-2026|Wu (2026)]] tracks learners of Japanese over a semester and finds they sort into maladaptive, moderate, and positive adaptation profiles driven by the balance of technostress and resilience, with most learners gradually shifting toward positive adaptation and reporting higher [[self-efficacy]] and lower burnout — a signal to design AI-mediated language practice that manages technological strain, not just tool access. - **Be alert to scoring and feedback bias against learners.** [[ai-scoring-language-bias-physics|AI scoring]] can penalize non-native patterns; [[genai-linguistic-diversity-academic-writing|linguistic-diversity research]] warns AI privileges standard English — use self-referential or human-moderated evaluation. - **Support the full spectrum of learners.** [[dyslexlens-dyslexic-learners-ai|Dyslexia and accessibility studies]] and [[culturally-relevant-pedagogy|culturally responsive]] design ([[ai-tools-arab-english-classrooms|Arab-English contexts]]) show AI must be adapted to diverse learner needs, not assumed universal. - **Educators value GenAI for preparatory work, not live classroom use.** A [[li-language-educators-genai-review-2026|PRISMA systematic review of 23 studies]] (Li et al. 2026) finds language educators most value GenAI for behind-the-scenes preparation — [[curriculum-design|lesson planning]], materials creation, and writing support/feedback — yet remain hesitant about direct, classroom-facing implementation, reflecting a theory–practice gap between approving AI in principle and using it live. Adoption is shaped by professional-identity, pedagogical, technical, [[governance|institutional]], and [[academic-integrity]] factors, with educators falling on a spectrum from non-adoption to comprehensive integration; attitudes tend to evolve from initial insecurity toward confident, selective use with exposure. - **Prepare language teachers' [[ai-literacy|AI literacy]].** [[governing-unseen-ai-literacy-language-teachers-2026|Systematic reviews]] find AI literacy among language teachers is a key gap — invest in teacher [[educational-development|professional development]] alongside tool adoption. As AI reshapes language education, AI literacy is also crucial for teachers to engage critically with the technology: the Teachers' AI Literacy Scale (TAILS) was developed for language [[teacher-education|teacher education]], operationalizing the six-dimension ED-AI framework (knowledge, evaluation, collaboration, contextualization, [[agency|autonomy]], [[ethics]]) and validated with preservice English language teachers. - **Four interaction profiles in a high-pressure bilingual task.** [[student-ai-interaction-consecutive-interpreting-2026|Kuang, Li and Weng (2026)]] used eye-tracking, pen-recording and voice-recording with 22 interpreting trainees to show that students divide attention between AI output and their own note-taking in four distinct ways — Intensive Engagers, Fast Scanners, Traditionalists and Frequent Switchers — and that 58.3% of stage-level observations changed profile between the comprehension and production stages of the same task. Only comprehension-stage patterns predicted product quality, and the AI-heaviest cluster scored lowest on fluency of delivery and target language quality, which makes the case for teaching learners to describe and reflect on their own strategy rather than prescribing one way of working with the tool. ## Connected Concepts - [[eportfolio]] - [[writing-education]] - [[ai-literacy]] - [[equity-in-ai-education]] - [[assessment-validity]] - [[bias-mitigation]] - [[inclusive-learning]] - [[special-education]] - [[intelligent-tutoring]] - [[generative-ai]] - [[student-experience]] - [[higher-ed]] - [[k-12]] - [[discipline-specific-aied]] - [[english-education]] - [[speech-and-voice-technologies]] ## Connected Articles - [[student-ai-interaction-consecutive-interpreting-2026]] — Student-AI Interaction in Computer-Assisted Consecutive Interpreting - [[wu-psychological-adaptation-ai-japanese-learning-2026]] — Profiles and Transitions of Psychological Adaptation in AI-Assisted Japanese Language Learning - [[llm-agents-5e-esl-grammar-2026]] — LLM agents with 5E framework for ESL grammar acquisition (Yang, Weng & Yang 2026) - [[gpt-item-generation-l2-listening-2026]] — Prompting vs. fine-tuning GPT for L2 listening item generation (Aryadoust & Wong 2026) - [[bert-discourse-english-teaching-2026]] — BERT discourse classification for English teaching - [[alharbi-ethical-genai-eap-2026]] - [[sutama-chatgpt-eportfolio-speaking-2026]] - [[ni-lam-multiliteracies-ai-portfolio-2026]] - [[llms-text-linguistics-teaching-2026]] — LLMs in text linguistics teaching - [[ai-vs-human-assessment-efl-tpck-2026]] — AI-generated vs human-developed assessment tasks in EFL - [[governing-unseen-ai-literacy-language-teachers-2026]] — Governing the unseen: AI literacy among language teachers - [[ai-guided-learning-audiovideo-2026]] - [[ai-interlocutor-l2-spoken-dialogue]] - [[robot-assisted-language-learning-meta-analysis-2026]] — Meta-analysis of AI-enhanced embodied robot-assisted language learning - [[self-referential-l2-writing-llm-assessment]] - [[ai-scoring-language-bias-physics]] - [[genai-linguistic-diversity-academic-writing]] - [[dyslexlens-dyslexic-learners-ai]] - [[tact-pedagogically-adaptive-esl-tutoring]] - [[ai-tools-arab-english-classrooms]] - [[structural-silence-underrepresented-language-ai-2026]] - [[bilingual-llm-lecture-companion-srl-2026]] - [[instructor-designed-ai-tutors-foreign-language-sdt-2026]] — Instructor-Designed AI Tutors in University Foreign Language Education: A Mixed-Methods Study of Learner Motivation and Reflective Learning Experience Based on Self-Determination Theory - [[lukesova-clue-before-correction-2026]] — Clue Before Correction: ChatGPT for Autonomous Language Learning - [[chatgpt-english-language-learning-malaysia]] — Students' ChatGPT experiences in English language learning - [[tts-dialogue-lessons-learner-characteristics-2026]] — Learner characteristics × TTS dialogue-format interactions - [[liu-emerging-tech-tefl-review-2026]] — Meta-analysis of emerging tech for EFL - [[wang-goal-setting-ai-engagement-2026]] — Goal-setting theory: teacher support, achievement goals, and engagement in AI-assisted English learning (758 Chinese students) - [[language-teachers-ai-literacy-edai-2026]] — Teachers' AI Literacy Scale (TAILS) psychometric study (ED-AI framework) - [[li-language-educators-genai-review-2026]] — Language educators' practices and development with GenAI - [[asr-english-speaking-feedback-metacognition-2026]] — ASR technology in college English speaking: feedback internalization and metacognitive strategies - [[genai-writing-program-primary-l2-motivation-engagement]] — A GenAI-supported writing program for primary L2 learners (Lu et al. 2026) --- ## [English Education (EAP / EFL / ESL)](https://edtechdev.github.io/aied/concepts/english-education/) > **English education** — the application of AI to the [[teacher-role|teaching]] and learning of English, especially **English for Academic Purposes (EAP)** and English language teaching more broadly (EFL/ESL/L2). This is a [[discipline-specific-aied|discipline-specific]] [[ai-education|AIEd]] strand distinct from both general [[language-learning]] (second/foreign-language acquisition of any language) and [[writing-education]] (writing as a general skill): it centers on English as a target language and academic register, with its own signature pedagogies — communicative competence, genre-based academic writing, corrective feedback, and reading/writing in an academic register — that shape how AI is designed, used, and evaluated. ## Questions to Consider - English is the language large language models handle best. That gives English learners powerful scaffolding for academic writing — but could it also quietly entrench a bias toward standard academic English that marginalizes [[multilingual-learning|multilingual]] and World Englishes writers? Which outcome do you think dominates? - English for Academic Purposes (EAP) centers on genre-based academic writing, corrective feedback, and academic register — distinct from general language learning or general writing instruction. Why might the specific register of academic English change what AI tools need to do versus generic writing support? - If the AI that revises your English is itself strongest in the very 'standard academic English' you're trying to master, when does that help and when does it flatten your own voice or dialect? How would you tell the difference? - Automated feedback and AI tutors are increasingly common in English writing and speaking instruction. What might AI feedback miss about communicative competence, register, and audience that a human instructor or peer would catch? ## Introduction English is one of the most AI-affected discipline strands because [[llm|LLMs]] are English-dominant: they are strongest at generating, revising, and evaluating English text, which is exactly what EAP and EFL/ESL instruction centers on. That English-advantage creates a distinctive double edge — powerful [[scaffolding]] for academic English on one hand, and an entrenched bias toward standard academic English that can marginalize multilingual learners on the other. ## Scope and focus This concept organizes AI [[research-methods-aied|research]] in **English education** — the subset of language learning where the target language is English (including EFL/ESL/L2 contexts) and the academic-English register (EAP). Core themes: - **Academic English (EAP):** AI support for the genre-based, discipline-specific English used in higher-education writing, reading, and feedback — distinct from general writing instruction. - **English language teaching (EFL/ESL/L2):** [[intelligent-tutoring|AI tutors]], interlocutors, and [[feedback]] tools for learners acquiring English. - **English-specific [[assessment]]:** automated evaluation and feedback on English writing and speaking, including EAP writing revision and L2 writing assessment. - **Reading and literature readability:** Bird (2026) fuses transformer text classification with computational-linguistics features to classify English literature by UK Key Stage, reaching an F1 of 0.996 — a scalable, data-driven complement to genre-based EAP reading support and reading-level alignment. - **Linguistic equity:** the tension between AI's English dominance and the needs of multilingual and World Englishes writers. ## How English education differs from Language Learning [[language-learning]] is the broader umbrella for acquiring any second/foreign language — spoken, written, and literate — via AI interlocutors, pronunciation tools, and conversational practice. English education is the **English-specific** case, and within it, EAP is a **register-specific** case: | Dimension | [[language-learning]] | **English education (this page)** | |-----------|----------------------|-----------------------------------| | Target language | Any L2 (French, Spanish, Japanese, …) | English specifically | | Focus | L2 acquisition generally: spoken dialogue, pronunciation, literacy | English as a target + the academic-English register | | Signature contexts | Conversation, pronunciation, general fluency | **EAP**: academic writing, reading, feedback, genre | | Representative AI | L2 interlocutors, pronunciation feedback, robot-assisted L2 | EAP writing tools, EFL peer-feedback, English academic writing assessment | The two overlap heavily (most English learning is also L2 acquisition), but English education foregrounds English as the target and the academic register — e.g., [[alharbi-ethical-genai-eap-2026|ethical GenAI integration in EAP]], [[feedback-literacy-scripts-eap-writing|GenAI EAP writing revision]], and [[genai-differentiated-eap-reading-materials-2026|EAP reading-material adaptation]] are EAP-specific in ways generic language-learning research is not. ## How English education differs from Writing Education [[writing-education]] concerns writing as a general cognitive and rhetorical skill — across all languages and disciplines, from composition to academic integrity. English education focuses specifically on **English** and, within EAP, on the **academic register**: | Dimension | [[writing-education]] | **English education (this page)** | |-----------|----------------------|-----------------------------------| | Scope | Writing in general (any language, any genre) | English as target language + academic English register | | Signature concern | Composition, revision, [[agency]], authorship | EAP genre, academic register, L2/EFL writing, [[feedback-literacy|feedback literacy]] in English | | Assessment angle | [[automated-essay-scoring|Automated essay scoring]], writing feedback broadly | English-specific assessment (EAP writing, EFL [[peer-assessment|peer feedback]], L2 writing evaluation) | | Equity angle | Bias in writing feedback | Bias + the English-dominance/multilingual tension (World Englishes) | Many writing-education articles are English-first (e.g., [[marked-pedagogies-linguistic-bias-writing-feedback|Marked Pedagogies]]), but they are framed as general writing research; English education re-centers the **English-as-target** and **academic-English** dimensions that generic writing and generic language-learning pages underemphasize. ## Articles in this cluster - **EAP-specific:** [[alharbi-ethical-genai-eap-2026|Ethical GenAI integration in EAP]], [[feedback-literacy-scripts-eap-writing|GenAI EAP writing revision]], [[genai-differentiated-eap-reading-materials-2026|EAP reading-material adaptation]]. - **EFL/ESL/L2:** [[tact-pedagogically-adaptive-esl-tutoring|TACT ESL tutoring]], [[sutama-chatgpt-eportfolio-speaking-2026|ChatGPT EFL e-portfolio speaking]], [[irwin-muller-efl-peer-feedback-literacy|EFL peer-feedback literacy]], [[ai-vs-human-assessment-efl-tpck-2026|AI vs human EFL assessment]], [[acceptance-ai-english-tools-2026|acceptance of AI English tools]], [[ai-tools-arab-english-classrooms|AI in Arab English classrooms]]. - **L2 English writing/assessment:** [[self-referential-l2-writing-llm-assessment|self-referential L2 writing assessment]], [[ai-interlocutor-l2-spoken-dialogue|L2 spoken-dialogue interlocutors]]. - **Linguistic equity / World Englishes:** [[genai-linguistic-diversity-academic-writing|GenAI and linguistic diversity in academic writing]], [[governing-unseen-ai-literacy-language-teachers-2026|AI literacy among language teachers]], [[structural-silence-underrepresented-language-ai-2026|underrepresented languages in AI infrastructure]]. ## Why it matters AI's English dominance is a defining feature of this strand. Because models are strongest in English and in standard academic English specifically, English education both benefits disproportionately (powerful EAP scaffolds) and carries distinctive risks (monolingual bias, discrimination against World Englishes and multilingual writers). Research here connects to [[equity-in-ai-education|equity]], [[bias-mitigation|bias mitigation]], [[automated-assessment|automated assessment]], [[ai-feedback-quality|AI feedback quality]], and [[academic-integrity|academic integrity]]. ## Implications for English / EAP / EFL-ESL instructors - **Exploit AI's strength for academic English, deliberately.** Because models are strongest in English and standard academic English, EAP instructors can deploy AI for genre-based writing, reading-material differentiation ([[genai-differentiated-eap-reading-materials-2026|EAP materials]]), and revision feedback — but should frame AI as a drafting/feedback partner, not an answer engine. - **Protect academic-English register and feedback literacy.** [[feedback-literacy-scripts-eap-writing|EAP writing revision]] shows feedback is only as productive as the learner's feedback literacy — teach students to interpret, judge, and act on AI feedback, and use second-rater mechanisms to check AI quality. - **Watch the English-dominance equity tension.** Models privilege standard academic English, marginalizing World Englishes and multilingual writers ([[genai-linguistic-diversity-academic-writing|World Englishes]], [[marked-pedagogies-linguistic-bias-writing-feedback|Marked Pedagogies]]) — audit feedback for monolingual bias and lowered expectations. - **Integrate AI ethically into EAP.** [[alharbi-ethical-genai-eap-2026|Ethical GenAI in EAP]] calls for transparent, responsible use in [[higher-ed]] English teaching that preserves academic integrity. - **Expect cautious, preparatory-first adoption.** A [[li-language-educators-genai-review-2026|systematic review of 23 studies]] (Li et al. 2026) finds language educators value [[generative-ai|GenAI]] most for behind-the-scenes preparation — lesson planning, materials creation, and writing support/feedback — while hesitating on direct classroom use, with primary concerns centering on [[academic-integrity|academic integrity]] (plagiarism and [[assessment-validity|assessment validity]]). Adoption is shaped by professional-identity, [[pedagogy|pedagogical]], technical, [[governance|institutional]], and integrity factors, and competency gaps map to episteme (understanding AI's capabilities/limits), techne ([[prompt-engineering]], AI-enhanced task/assessment design, detecting AI-generated text), and phronesis ([[ethics|ethical]] judgment, bias/privacy handling) — so EAP/EFL instructors should build these competencies deliberately and plan a "back-end then classroom" implementation. - **Differentiate by proficiency and need.** [[ai-vs-human-assessment-efl-tpck-2026|EFL assessment]] and adaptive tutoring research support tailoring AI support and evaluation to learners' level rather than one-size-fits-all. ## Connected Concepts - [[language-learning]] - [[writing-education]] - [[multilingual-learning]] - [[higher-ed]] - [[k-12]] - [[generative-ai]] - [[llm]] - [[automated-assessment]] - [[ai-feedback-quality]] - [[feedback-literacy]] - [[equity-in-ai-education]] - [[bias-mitigation]] - [[academic-integrity]] - [[discipline-specific-aied]] ## Connected Articles - [[alharbi-ethical-genai-eap-2026]] — Ethical Generative AI Integration in EAP within Higher Education - [[feedback-literacy-scripts-eap-writing]] — Feedback Literacy Scripts and a Second-Rater Mechanism in GenAI EAP Writing Revision - [[genai-differentiated-eap-reading-materials-2026]] — From Unified to Differentiated Materials: GenAI-Supported Adaptation of EAP Reading Materials - [[tact-pedagogically-adaptive-esl-tutoring]] — TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring - [[sutama-chatgpt-eportfolio-speaking-2026]] — Aligning ChatGPT with E-Portfolio Assessment as EFL Learning Model - [[irwin-muller-efl-peer-feedback-literacy]] — Positioning Generative AI in EFL Peer Feedback - [[ai-vs-human-assessment-efl-tpck-2026]] — AI-Generated versus Human-Developed Assessment Tasks in EFL Context - [[acceptance-ai-english-tools-2026]] — Acceptance of AI-Assisted English Language Learning Tools in Higher Education - [[ai-tools-arab-english-classrooms]] — AI tools in Arab University English classrooms - [[self-referential-l2-writing-llm-assessment]] — Toward Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs - [[ai-interlocutor-l2-spoken-dialogue]] — What Changes When the Interlocutor Is an AI? L2 Spoken Dialogue - [[genai-linguistic-diversity-academic-writing]] — Generative AI and Linguistic Diversity in Academic Writing and Publishing - [[governing-unseen-ai-literacy-language-teachers-2026]] — Governing the Unseen: AI Literacy among Language Teachers - [[structural-silence-underrepresented-language-ai-2026]] — Structural Silence: Underrepresented Languages in AI Infrastructure - [[liu-emerging-tech-tefl-review-2026]] — Emerging technologies for TEFL - [[wang-goal-setting-ai-engagement-2026]] — Goal-setting theory: teacher support, achievement goals, and engagement in AI-assisted English learning (758 Chinese students) - [[bird-multimodal-educational-literature-2026]] — Multimodal fusion for classifying educational literature - [[li-language-educators-genai-review-2026]] — Language educators' practices and development with GenAI --- ## [Environmental Education](https://edtechdev.github.io/aied/concepts/environmental-education/) ## Questions to Consider - "AI and environmental education" can mean two very different things: using AI to teach *about* climate and sustainability, and reducing the environmental cost of *using AI* in classrooms. Which do you hear in your institution — and which has a budget line attached? - If a university teaches climate science with an energy-intensive [[llm|large language model]] on every student laptop, has it advanced environmental education or undercut it? What would have to be true for the answer to be "both"? - [[ai-assisted-inquiry-ssi-climate|One experiment]] found AI-assisted inquiry improved climate decision-making over inquiry alone, with gains concentrated on the weakest steps. Would you trust an AI-supported climate unit in your classroom, and what would you need to see first? - Nearly no [[ai-education|AIED]] papers report their compute or carbon footprint. Is unreported [[sustainability]] impact a research problem, a procurement problem, a teaching problem, or none of these? ## Introduction **Environmental education** develops learners' understanding of ecological systems, climate, and human–environment interdependence, together with the dispositions and competences to act on them — often framed institutionally as Education for Sustainable Development (ESD) or "green education." In an AI-in-education context the term carries a tension the corpus runs into without always naming: **AI as a tool for environmental education** (teaching climate and sustainability content with generative tools) and **the environmental footprint of AI itself** (the carbon, water, and energy cost of the models institutions deploy). These are different questions with different evidence, actors, and remedies; conflating them — as several papers and most strategy documents do — makes it impossible to say who is accountable for what.([[daniel-ai-sustainability-scoping-review-2026]])([[aied-carbon-footprint-reporting]]) ## Two things called "environmental education with AI" - **AI for environmental education** — using AI to teach environmental content, build sustainability consciousness, or support green skills. This half has frameworks, a [[teacher-role|teacher]]-survey study, and one classroom experiment. - **The footprint of AI in education** — emissions and resource use of the models and infrastructure used for teaching, whatever the subject. This half has one review of reporting practice, one engineering [[benchmark]], and one small interface study — and nothing on the footprint of AI used for environmental education specifically. A third, weaker strand treats "sustainable learning" as a [[pedagogy|pedagogical]] property rather than an environmental one: learning that persists and transfers rather than being short-circuited by [[cognitive-offloading]]. It shares sustainability's vocabulary but is not an environmental claim.([[zhu-e3-hot-embodied-intelligence-sustainable-learning]]) ## Teaching environmental and climate topics with AI The strongest evidence is a three-group quasi-experiment using climate change as a socio-scientific issue. Students working with an AI partner inside a structured [[inquiry-based-learning|inquiry]] task outperformed inquiry-only peers (d = 0.69) and traditional instruction (d = 1.88) on a decision-making rubric, with the largest gains on steps students entered weakest — monitoring and adaptive management, and generating alternatives. The AI condition was *additional to* inquiry rather than a replacement, and data collection and analysis did not separate the groups at all, suggesting the AI supported reasoning about trade-offs rather than evidence gathering.([[ai-assisted-inquiry-ssi-climate]]) At [[curriculum-design|curriculum]] level, the AI-SEE framework integrates AI across an [[engineering-education|engineering curriculum]] on four pillars (intelligence-driven, green-empowered, responsibility-leading, practice-integrated) rather than as a bolt-on sustainability module; a 144-student pilot reported gains in sustainability consciousness and behavioral [[student-engagement|engagement]] across personal, academic, professional, and social levels. The authors caution that this is single-institution, self-reported, single-time-point interview evidence from one Chinese transportation program.([[liu-ai-sustainable-engineering-education-2026]]) Two smaller studies cover the design side. AI-assisted [[learning-design|instructional design]] under Sustainable Development Pedagogy constraints improved pre-service teachers' lesson plans with an implausibly large effect (d = 2.80, 28 teams, no control group). Separately, a structural-equation study of 122 in-service teachers found *practical* AI use on science and green-energy tasks plus involvement in developing ESD-aligned materials predicted AI-integration capability, while abstract AI knowledge and attitudes did not — though a dichotomous [[self-report-measures|questionnaire]] and a weak knowledge construct limit how far it travels.([[talebzadeh-ai-green-education-2026]])([[riandi-teacher-ai-green-energy-education-2026]]) ## The environmental footprint of AI in education This half is where the corpus is thinnest. A review of all 396 AIED 2025 proceedings papers found "LLM adoption without disclosure": most projects use [[llm|LLMs]], 85 papers reported any computational cost, and only 57 mentioned environmental impact — using incompatible metrics, so the field cannot aggregate its own evidence. The authors argue that failing to report environmental cost is itself an [[ethics|ethical]] concern, and propose an [[open-source]] method (CodeCarbon plus a two-parameter FLOPs estimate for proprietary models).([[aied-carbon-footprint-reporting]]) On the learner side, an eco-feedback interface exposing the latency–carbon trade-off during live LLM use was studied with 89 computer science undergraduates in a computing ethics course: students chose the lower-carbon mode in roughly 45% of low-latency interactions but under 5% at high latency, and sustainability awareness significantly raised that choice. The sample is small and unusually technically informed, but it is the only direct evidence in the corpus that footprint information moves learners at all — and it suggests the binding constraint is patience, not values.([[llm-environmental-impact-student-usage-2026]]) Mitigation evidence comes from an on-premise [[cs-education|CS]] knowledge-base assistant on a single consumer GPU (12 GB VRAM) with openly licensed content: 1.8 mWh per query at best — about 0.54 Wh for a class of 30 students submitting ten queries each — with quantization-aware fine-tuning containing both the accuracy loss and the [[hallucination-risk|hallucination]] rise that compression alone caused. It is an engineering benchmark, not a learning study, but it shows the footprint question has design levers (grounding, quantization, deployment location), not only usage-discipline levers.([[shen-sustainable-ai-knowledge-base-cs-education-2026]]) ## Green skills, sustainability consciousness, and teacher capacity Together these studies describe a green-skills agenda delivered unevenly: sustainability consciousness as a curricular outcome, ESD-aligned material development as the mechanism for teacher capability, and climate decision-making as a measurable reasoning skill. The value-critical strand supplies the caveat: a conceptual analysis argues AI's contribution to sustainable education is **conditional and governance-mediated**, supporting sustainability only when adoption is subordinated to explicit educational values and human-centered purposes — which places [[governance]] and [[critical-pedagogy|critical pedagogy]] inside environmental education rather than beside it.([[alsuhaymi-sustainable-education-ai-digitalization-2026]]) The [[meta-analysis-systematic-review|scoping review]] organizing this literature supplies the field's own verdict on scale: applications cluster around energy management, climate monitoring, and green-campus programs, but are limited in scale, often lack ethical or environmental guidelines, and much of the research remains conceptual or small-pilot. Coverage is skewed toward North America and Europe, with almost nothing from the [[global-south|Global South]] beyond a few South African studies.([[daniel-ai-sustainability-scoping-review-2026]]) ## What the evidence does not yet establish - **No footprint evidence for environmental education specifically.** Carbon and energy figures come from general LLM studies and a CS-context benchmark; nothing measures the environmental cost of an AI-supported climate or ESD curriculum. - **No learning-gain evidence for footprint interventions.** Eco-feedback changed choices in a lab-like study; nothing shows it changes habits, assessment outcomes, or procurement. - **No causal evidence for the curriculum frameworks.** AI-SEE is a single-site self-report case and E3-HOT a design blueprint with no implemented study; the SDP lesson-design result has no control group. - **No large-scale or cross-context evidence.** Every empirical result here is single-site, and the teacher studies rest on small purposive samples with weak instruments. - **The two halves are rarely studied together**, even though both bear on the same classroom. ## Implications for AI in education 1. **Name which question you are answering.** Strategies that use "sustainability" for both AI-for-climate-teaching and AI's own footprint hide the accountability gap [[daniel-ai-sustainability-scoping-review-2026|Daniel et al. (2026)]] identify; keep the agendas separate, with separate owners. 2. **Treat footprint disclosure as field infrastructure.** Reporting compute and carbon beside accuracy — with a sustainability statement even when measurement is imperfect — is the only route out of the incompatible-metrics problem. 3. **Design for the impatient learner.** If the lower-carbon option costs perceived latency, most students will not take it: keep latency low, expose controls, and describe impact in concrete outcome terms, not abstract carbon units. 4. **Prefer lighter deployment where pedagogy allows.** Grounding a model in a licensed local corpus and quantizing with fine-tuning gave usable accuracy at a fraction of the energy — reuse that pattern before scaling cloud inference. 5. **Build teacher capability through material development, not awareness campaigns,** and fill the [[global-south|Global South]] gap rather than importing evidence from North American and European contexts. ## Connected Concepts - [[sustainability]] - [[global-south]] - [[critical-pedagogy]] - [[inquiry-based-learning]] - [[science-education]] - [[engineering-education]] - [[teacher-education]] - [[curriculum-design]] - [[ethics]] - [[governance]] ## Connected Articles - [[daniel-ai-sustainability-scoping-review-2026]] — Separates AI for sustainability from sustainable AI; Global South gap - [[ai-assisted-inquiry-ssi-climate]] — AI-assisted climate inquiry, d = 0.69 over inquiry alone - [[aied-carbon-footprint-reporting]] — AIED 2025 disclosure review plus an open-source footprint method - [[llm-environmental-impact-student-usage-2026]] — Eco-feedback on latency–carbon trade-offs with 89 CS students - [[shen-sustainable-ai-knowledge-base-cs-education-2026]] — On-premise quantized assistant at 1.8 mWh per query - [[liu-ai-sustainable-engineering-education-2026]] — AI-SEE: sustainability consciousness in engineering education - [[riandi-teacher-ai-green-energy-education-2026]] — Practical AI use and ESD material development predict integration - [[talebzadeh-ai-green-education-2026]] — AI-assisted design under Sustainable Development Pedagogy constraints - [[alsuhaymi-sustainable-education-ai-digitalization-2026]] — AI's contribution to sustainable education as governance-mediated - [[zhu-e3-hot-embodied-intelligence-sustainable-learning]] — "Sustainable learning" as durable cognitive agency, not an environmental claim --- ## [Business Education](https://edtechdev.github.io/aied/concepts/business-education/) > **AI in business education** — the application of artificial intelligence to the teaching and learning of business, economics, and management, and the preparation of students for a GenAI-integrated workplace. Business schools face a double imperative: integrating AI *into* the curriculum as a subject and [[pedagogy|pedagogical]] tool, while also preparing students to *use* generative AI responsibly and effectively in professional practice. ## Questions to Consider - Business schools face a double imperative: integrating AI into the curriculum as subject and tool, while preparing students to use AI responsibly in professional practice. Before reading, which of those two do you think most business programs currently get right, and which gets shortchanged? - A student-informed framework found students were highly engaged and recognized they needed GenAI skills for careers — but also wanted to use AI without unintentionally committing academic misconduct. Why might a student want to use AI and simultaneously worry about breaking rules around it? - [[research-methods-aied|Research]] on constructive alignment found that whether GenAI benefits or risks dominated was shaped by the degree of curriculum integration — integrate, don't append. What does 'integration' look like in a business program as opposed to just adding AI tools on top? - A decade of research flags persistent gaps in curriculum coherence, educator readiness, and [[assessment-validity|assessment validity]]. If you were redesigning a business course for an AI-integrated workplace, which of these three gaps would you tackle first and why? - Business is one of the fastest fields for generative AI adoption, so graduates face acute pressure to be AI-ready. How might the AI skills that employers in finance, marketing, or management actually value differ from what students are currently taught in the classroom? ## Introduction AI in business education is a growing [[discipline-specific-aied|discipline-specific]] strand of [[ai-education|AI in education]]. Business schools must prepare students for a future workforce where generative AI is pervasive, requiring both AI literacy and domain-specific application (finance, marketing, management, economics, entrepreneurship). The research highlights a tension between integrating AI as a pedagogical and professional tool and managing [[academic-integrity|academic integrity]], [[ethics|ethical]] use, and [[assessment]] redesign. ## AI in business education across the knowledge base's research - **Student-informed frameworks.** [[drummond-genai-business-schools-framework-2026|Drummond & Dale (2026)]] present a student-informed conceptual framework for integrating GenAI throughout business degree programs, based on a 149-student UK case study. Students were highly engaged, recognized the need for GenAI skills for their careers, and wanted to use AI without unintentionally committing academic misconduct. - **A decade of research.** [[espino-ai-business-education-review-2026|Espino & Espino (2026)]] bibliometrically mapped 213 articles (2015–2024) on AI in business education, identifying four clusters: AI-driven business education transformation, innovative digital pedagogies, AI-enhanced [[personalized-learning|personalization]], and business education aligned with the digital economy — with persistent gaps in curriculum coherence, educator readiness, and assessment validity. - **Constructive alignment.** [[zhou-constructive-alignment-genai-business-2026|Zhou et al. (2026)]] analyzed 17 cases of GenAI adoption at a UK business school, finding the balance between pedagogical benefits and risks was shaped by the **degree of curriculum integration** — constructive integration produced better outcomes. - **Student insights on curricula.** [[rook-plumb-genai-curricula-student-insights-2026|Rook & Plumb (2026)]] drew on 166 undergraduate business students in a capstone unit, finding strong support for integrating GenAI into curricula with three priority areas: understanding/optimizing GenAI functionality, exploring applications across contexts, and navigating ethical/legal dimensions. - **Organizational adoption.** [[alrahmi-org-drivers-ai-adoption-he-2026|Al-Rahmi (2026)]] examined organizational and technological drivers of AI adoption in [[higher-ed|higher education]] (using TOE + Diffusion of Innovations frameworks, n=300 staff), relevant to how business schools and their institutions adopt AI-driven decision support and smart learning platforms. - **Conversational agents in business [[simulation]] games.** [[conversational-agents-business-simulation-gaming-2026|Wenzel, Geiger, and Liening (2026)]] develop and evaluate an AI-enhanced [[conversational-ai|conversational agent]] (Lara) for adaptive instructional support in business simulation games used for experiential entrepreneurial learning. Grounded in an [[equity-in-ai-education|equity]]-by-design stance and [[universal-design-for-learning|universal design for learning]], the CAIS-GBL framework targets cognitive, [[motivation|motivational]], [[affective-computing|affective]], and [[sociocultural-learning|socio-cultural]] [[student-engagement|engagement]], with evaluations among student teachers and BSG participants showing positive perceptions of cognitive/[[community-of-inquiry|social presence]] and [[self-regulated-learning]] support — a model for scaling [[formative-assessment|formative]] feedback in [[game-based-learning|simulation-based]] business education. - **Agentic GAI for complex applied tasks.** [[ilieva-agentic-genai-higher-education-2026|Ilieva et al. (2026)]] develop the AGAI-HE framework around an **e-commerce** course, where the learning task requires comparing business models, weighing market and operational constraints, and defending strategic recommendations — multi-step applied work that prompt-response chatbots support poorly. Their 130-student perception study found both chatbot and agent support rated above traditional [[online-teaching-and-learning|e-learning]] on learning enhancement, personalization, decision-making support, and workflow organization, yet no significant agent-versus-chatbot difference, and the strongest endorsement went to combining all three modes. For business programs the reading is that adoption is an [[curriculum-design|instructional-design]] question about how AI support is sequenced and governed, not simply a tooling upgrade ([[agentic-ai]]). - **Assessment and grading innovation is under-researched.** [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] surveyed 58 assessment-related articles over 20 years across four leading management-education journals (AMLE, JME, Management Learning, IJME) — roughly three per year, many in 2010 and 2014 special issues — finding assessment and grading under-researched relative to their central role in shaping learning. Self- and peer-assessment dominate the discourse (nearly half the corpus, chiefly for [[summative-assessment|summative]] evaluation of [[group-work|group work]]); [[authentic-assessment|authentic assessment]] appears mainly via technology-mediated [[simulation|simulations]] and is often conflated with [[experiential-learning|experiential learning]]; while reassessment, standards-based grading, and ungrading are virtually absent. The authors urge management educators to engage more actively and reciprocally with these innovations, recommending incremental experimentation (ungraded assignments, reassessment for a single task, standards-based rubrics) backed by program-level coordination and documented Scholarship of Teaching and Learning evidence. ## Economics and management education AI in business education spans **economics** and **management** as core disciplines. The knowledge base's coverage focuses on the pedagogical and curriculum dimensions (how AI is taught and used in business programs) and the professional-readiness dimension (preparing students for AI-integrated workplaces). Key themes include [[ai-literacy|AI literacy]], [[generative-ai|generative AI]] application across business functions, [[ethics|ethical]] and [[academic-integrity|integrity]] considerations, and [[curriculum-design|curriculum]] and [[assessment]] redesign to reflect AI-integrated professional practice. ## Why AI in business education matters Business is one of the fields where generative AI adoption is fastest, so business schools face acute pressure to prepare students for an AI-integrated workplace. The research emphasizes that effective integration is **curriculum-driven and student-informed** — not just adding AI tools but redesigning programs so that GenAI literacy, ethical judgment, and authentic application are woven through degree structures. This connects business education to the knowledge base's broader themes of [[ai-literacy|AI literacy]], [[teacher-role|educator]] preparation, [[assessment]] redesign, and [[curriculum-design|curriculum]] reform in the AI era. ## Implications for business instructors - **Integrate AI curriculum-driven and student-informed.** [[drummond-genai-business-schools-framework-2026|Student-informed frameworks]] and [[rook-plumb-genai-curricula-student-insights-2026|student insights]] show students want GenAI skills for careers and to use AI without unintentional misconduct — design programs that weave AI literacy, ethical judgment, and authentic application through degree structures. - **Align curriculum constructively.** [[zhou-constructive-alignment-genai-business-2026|Constructive alignment research]] finds the degree of curriculum integration shapes whether GenAI benefits or risks dominate — integrate, don't append. - **Address the persistent gaps.** [[espino-ai-business-education-review-2026|A decade of research]] flags gaps in curriculum coherence, educator readiness, and assessment validity — prioritize these in program design. - **Prepare students for an AI-integrated workplace.** Emphasize AI literacy, ethical use, and domain application (finance, marketing, management, economics) as core competencies, not electives. ## Connected Concepts - [[ai-education]] - [[discipline-specific-aied]] - [[generative-ai]] - [[ai-literacy]] - [[curriculum-design]] - [[assessment]] - [[ethics]] - [[academic-integrity]] - [[technology-acceptance-model]] - [[teacher-role]] - [[higher-ed]] - [[stem-education]] ## Connected Articles - [[ilieva-agentic-genai-higher-education-2026]] — AGAI-HE: agentic GAI support in an e-commerce course, perceived benefits and risks (Ilieva et al. 2026) - [[drummond-genai-business-schools-framework-2026]] — Student-informed framework for GenAI in business schools (Drummond & Dale 2026) - [[espino-ai-business-education-review-2026]] — A decade of AI in business education (Espino & Espino 2026) - [[zhou-constructive-alignment-genai-business-2026]] — Constructive alignment of GenAI in business higher education (Zhou et al. 2026) - [[rook-plumb-genai-curricula-student-insights-2026]] — Student insights on integrating GenAI into curricula (Rook & Plumb 2026) - [[alrahmi-org-drivers-ai-adoption-he-2026]] — Organizational drivers of AI adoption in higher education (Al-Rahmi 2026) - [[conversational-agents-business-simulation-gaming-2026]] — CAIS-GBL framework for AI conversational agents in business simulation games (Wenzel et al. 2026) - [[mesny-innovative-assessment-grading-management-2026]] - [[automated-scoring-marketing-posts-agreement-2026]] — Agreement and error in automated scoring of student marketing posts - [[shi-genai-experiential-learning-management-education-2026]] — three generative AI mechanisms for reconfiguring management education pedagogy --- ## [Humanities and Social Science Education](https://edtechdev.github.io/aied/concepts/humanities-education/) > **Humanities and Social Science (SSH) Education** — the [[teacher-role|teaching]] of disciplines concerned with human culture, values, meaning, and social life, including history, philosophy, literature, languages, sociology, and the arts. AI in SSH education raises distinctive questions because these fields center on interpretation, critical judgment, authorship, and meaning-making — processes that [[generative-ai|generative AI]] both supports and disrupts. AI here functions less as a tutor of factual content and more as an epistemic mediator that reshapes how students interpret texts, construct arguments, and understand their own [[agency]] as writers and thinkers. ## Questions to Consider - The humanities prize interpretation, authorship, and critical judgment — exactly the capabilities generative AI most challenges. How does that make AI integration in these disciplines different from in STEM? - The page describes AI as an 'epistemic mediator' that can externalize your interpretive thinking and hybridize your authorship. When AI helps you interpret a text, who is actually doing the interpreting? - It identifies three trajectories for learners: AI-dependent, AI-enhanced, and AI-critical interpretation. Where do you think most humanities students currently land — and what would move them toward the AI-critical end? - In history education, [[research-methods-aied|research]] shows AI can shape reasoning and source interpretation, including through filters that 'protect' students. How would you teach students to interrogate an AI's mediation of historical claims? - If original authorship and interpretive autonomy are the core values of the humanities, what does it mean when AI can produce a plausible interpretation instantly? What human capability becomes more, not less, valuable? - The page warns against treating AI as a content tutor in fields centered on meaning-making. What would it look like to use AI deliberately to deepen, rather than flatten, students' critical analysis? ## Introduction SSH education is a distinct subject area in the knowledge base, complementary to [[stem-education]] and [[language-learning]]. Because the humanities prize interpretive depth, authorship, and contextual judgment, they pose different AI-integration challenges than STEM — and connect to [[higher-ed]] and [[ai-literacy]] in [[discipline-specific-aied|domain-specific]] ways. ### How AI appears in humanities and social science education - **Restructuring interpretive cognition.** [[voicu-ai-interpretive-cognition-ssh-2026|Voicu]] argues that generative AI acts as an *epistemic mediator* that reconfigures meaning-making, authorship, and learner [[agency]] in SSH education. It identifies three transformations — externalization of interpretive cognition, hybridization of authorship, and emergence of distributed epistemic agency — and three developmental trajectories (AI-dependent, AI-enhanced, AI-critical interpretation). This grounds a developmental-critical [[pedagogy|pedagogical]] model for preserving interpretive depth. - **History education.** [[paternalistic-filter-llm-history-education|Research on LLMs in history education]] examines how AI shapes historical reasoning and source interpretation. - **Humanities and social-science students' experiences.** [[genai-impact-chinese-students-hss|GenAI's impact on Chinese humanities and social science students]] and [[acceptance-ai-english-tools-2026|AI acceptance among language/English learners]] document student-facing effects across SSH disciplines. - **Critical and philosophical dimensions.** [[voicu-ai-interpretive-cognition-ssh-2026|Critical AI literacy]] and the [[philosophy-of-ai-in-education|philosophy of AI in education]] are especially salient in SSH, where questions of meaning, values, and epistemic authority are central. - **[[governance|Institutional]] digital transformation.** Qin (2026) documents how Lingnan University repositioned itself as a "Research-Intensive Liberal Arts Institution in the Digital Era," mandating GenAI literacy for all undergraduates (including a required first-year Common Core course on generative AI covering latent spaces, GANs, diffusion models, [[prompt-engineering|prompting]], fine-tuning, bias, and misinformation). The position paper argues the AI-for-education shift is an *intellectual transformation* rather than technocentric augmentation, positioning [[ai-literacy|digital fluency]] as a core liberal arts competency while a [[human-in-the-loop-ai|human-in-the-loop]] model foregrounds [[ethics|ethical]] reasoning, critical judgment, and social responsibility — a concrete blueprint for [[higher-ed|higher education]] balancing [[generative-ai|GenAI]] innovation with humanistic foundations. ### Why it matters SSH education foregrounds the very capabilities generative AI most challenges — original authorship, interpretive judgment, critical analysis, and context-sensitive meaning-making. The knowledge base treats this domain as a critical counterweight to instrumental, skills-based framings of AI: it asks whether AI-supported learning preserves [[critical-thinking]], epistemic responsibility, and interpretive autonomy, connecting to [[critical-pedagogy]] and [[ai-literacy]]. ## Implications for humanities and social-science instructors - **Treat AI as an epistemic mediator, not a content tutor.** [[voicu-ai-interpretive-cognition-ssh-2026|Voicu]] shows AI reconfigures meaning-making, authorship, and agency in SSH — design pedagogy around the three trajectories (AI-dependent, AI-enhanced, AI-critical) and aim for the AI-critical end. - **Protect interpretive depth and authorship.** Because the humanities prize judgment and original authorship, guard against AI flattening analysis; make critical [[ai-ed-evaluation|evaluation of AI]] output a learning goal ([[critical-thinking]], [[ai-literacy]]). - **Use AI deliberately in source work.** [[paternalistic-filter-llm-history-education|History research]] shows AI can shape reasoning and source interpretation — teach students to interrogate AI-mediated historical claims and the filters it applies. - **Consider student-facing impacts.** [[genai-impact-chinese-students-hss|Student experience studies]] document how GenAI affects SSH learners; adapt support and integrity framing to real usage rather than assumptions. ## Connected Concepts - [[higher-ed]] - [[critical-thinking]] - [[ai-literacy]] - [[critical-pedagogy]] - [[philosophy-of-ai-in-education]] - [[agency]] - [[language-learning]] - [[stem-education]] - [[ai-education]] - [[discipline-specific-aied]] ## Connected Articles - [[voicu-ai-interpretive-cognition-ssh-2026]] — A developmental-critical model of interpretive cognition for SSH education - [[paternalistic-filter-llm-history-education]] — LLM use and historical reasoning in history education - [[genai-impact-chinese-students-hss]] — GenAI's impact on humanities and social science students - [[acceptance-ai-english-tools-2026]] — AI acceptance among language and humanities learners - [[ai-digital-transformation-liberal-arts-lingnan-2026]] — Digital transformation of a liberal arts university toward a research-intensive model in the GenAI era (Qin 2026) --- ## [Arts, Design and Media Education](https://edtechdev.github.io/aied/concepts/arts-design-and-media-education/) > **Arts, design and media education** — the studio- and performance-based disciplines in which learning happens by making: architecture and spatial design, interior design, music performance, [[writing-education|composition]] and analysis, visual art and image generation, digital media and [[storytelling-in-education|digital storytelling]], and stage and performance technology. Across the articles in this knowledge base the recurring questions are whether [[generative-ai]] erodes craft and skill development or moves it up a level, how critique and the [[assessment]] of creative work change when a polished artifact is cheap to produce, how much of each discipline rests on [[embodied-learning|embodied]] and material practice, and who is included when creative tools become generative. Coverage is concentrated in [[higher-ed|higher education]] and in single-studio or single-course studies, so most findings below are case evidence rather than discipline-wide claims. ## Questions to Consider - In an architectural studio study, teams using a GenAI-plus-XR pipeline reported *lower* confidence in their own design ability afterward, and blinded raters scored their presentations no better than a control group's. What would you need to know before concluding the tools harmed learning — and what would make you conclude it anyway? - Studio teaching has always run on the critique and on visible process: sketches, models, successive iterations. If a student can produce a finished-looking image in a minute, what should the critique actually examine? - In a project-based digital storytelling capstone, students' ideas were rated more original than their finished works were coherent. Where does the hard learning really sit — generating ideas, or integrating them? - A music education review distinguishes autonomous generators such as Suno and Udio from interactive composition assistants such as the Continuator, and argues that only the second is educationally productive. Do you agree, and what is your criterion for the distinction? - Designers in Malaysia are described as shifting from primary form-generators to critical curators of machine output. If that is the destination, what should a first-year design course teach first? - Blind and low-vision string players learn bowing through touch and proprioception because visual demonstration is unavailable to them. What does that imply for [[feedback|feedback systems]] in the arts that default to a screen? ## Introduction What these disciplines share is the studio or the rehearsal room as the site of learning, with a made artifact and the process behind it as the object of assessment. AI enters in several shapes: an image generator in the architecture and interior design studio, a composition and analysis tool in music, a spatial partner inside immersive environments, an automated scorer of written analysis, and — in stage lighting — an instructor-facing aid that turns spoken teaching intent into executable demonstrations. The sections below trace what the studies report, then draw out the cross-cutting questions of craft, assessment and access. This page is the discipline home for studio and performing-arts subjects, and it is narrower than its neighbors. [[discipline-specific-aied]] holds the general argument that [[ai-education|AI in education]] research must be domain-aware across all subjects; this page is about what is specific to making. [[creativity]] covers the cognitive construct wherever it appears — divergent thinking, the homogenization risk of single-model assistance, protecting the learner's generative act — including in [[math-education|mathematics]] and creative coding; this page covers the disciplines in which creative production *is* the [[curriculum-design|curriculum]]. [[design-thinking]] covers ideation and user-centered problem framing as a general pedagogy; this page covers the studios where that process is taught, supervised and judged. [[humanities-education|Humanities and social science education]] treats interpretation, authorship and argument in text-centered fields and lists the arts among its constituent areas; here the arts and media are lifted out of that interpretive frame and centered on material, spatial, sonic and performative practice. [[design-education]] carves the professional half out of this same territory: this page spans the studio and performing arts whose shared feature is making, while design education follows the narrower formation pipeline for product, service, interaction, interior and architectural designers — portfolio assessment, accreditation expectations, employability — and treats the studio process behind an artifact rather than the artifact itself as the object of assessment. ## Studio Pedagogy and the Critique Under Generative AI Generative tools reach these disciplines through an existing ritual: work made publicly, shown incomplete, and criticized. What the sources show is that AI changes the *middle* of that process most reliably and the *ends* least. In the GenARch study, [[genai-xr-architectural-design-education-2026|Xiao et al. (2026)]] put a GenAI plus multi-user XR pipeline into a real undergraduate architectural design studio and found a complementary division of labor: [[generative-ai]] externalized ideas, giving teams something concrete to point at, compare and argue over, while XR supported evaluation, letting students judge proportion, adjacency, lighting and site context at full scale. The tooling was awkward at both ends of the process. At the start, teams had no ideas yet to encode in a prompt; later, richer concepts made [[prompt-engineering|prompting]] easier but raised expectations of constraint-following, and when precision mattered students went back to SketchUp and Revit. They reported a dimensional-fidelity failure — "I asked it to make a 10 feet wall of 3D model, it wasn't 10 feet" — asset latency of seconds to minutes, and shared attention that drifted when team members looked at different parts of a model. Giving every participant a headset did not by itself produce [[collaborative-learning|coordinated collaboration]]. The studio's accountabilities did shift. In a longitudinal study of generative models in design studios, [[genai-architectural-design-studios|Gül et al. (2026)]] distinguish using GenAI as visual stimulus in early ideation from a combinatorial use that expands the solution space during development, and locate the [[pedagogy|pedagogical]] problem in competence rather than access: without the discernment to differentiate between models and use them deliberately, abundant output produces [[design-thinking|design fixation]] and aesthetic lock-in. Their conclusion moves the [[teacher-role|instructor's role]] from transmitting craft toward coaching when, how and why to delegate creative exploration. ## Architecture, Spatial and Interior Design A Derby focus-group study gives the studio a more favorable result. [[genai-architecture-education|Kapsalis (2026)]] ran two groups of eight Level 3–5 architecture students through a 90-minute session with a locally executed image generation-and-editing workflow fine-tuned for residential interiors. Students produced roughly 80 images per session, about ten each, with acceptance rates above 60%. Self-reported idea generation (4.5) and experimentation (4.2) scored highest on the creative subscale, while refinement of final decisions (3.7) was less uniformly supported — some students found the tool surfaced unconsidered combinations, others that it distracted with details that did not fit the concept. The author frames this as a constructionist "microworld" that expands the creative search space without replacing design thinking, and notes survey evidence that nearly 70% of architecture students already use AI tools independently while over 95% report no formal [[ai-literacy|AI education]]. Interior design raises the professional-formation question most sharply. [[ai-interior-design-malaysia-2026|Syed Abdul Rahman (2026)]] describes Malaysian designers shifting from primary form-generators toward critical mediators and curators of machine output, with [[visualization]] platforms compressing timelines and widening the range of options producible within a budget. Citing empirical work on client preference (Lan et al., 2025), the article notes that AI output is rated highly for uniqueness when aesthetic criteria predominate, but that human designers retain a clear advantage when functional, ergonomic and contextual requirements are emphasized — and that generated designs often lack cultural specificity, climatic responsiveness and reliable constructability in a market negotiating tropical conditions, multi-generational living and Islamic spatial principles. Its curriculum recommendation is sequencing rather than substitution: teach CAD/BIM, construction knowledge and human factors before generative exploration, and assess [[critical-thinking]] and [[ethics|ethical]] reflection alongside visual quality. ## Music and Performing Arts Music is treated here as the art most exposed to economic change that began before generative AI. [[musical-education-ai-digital-transformation-2026|Briot (2026)]] argues that dematerialisation collapsed per-unit recorded-music value by roughly two orders of magnitude well before generative models were commercially relevant, that per-stream payouts remain on the order of \$0.003–\$0.005, and that generative AI prolongs an already-broken value model rather than creating it. His distinction that matters for teaching contrasts autonomous generators such as Suno and Udio — complete, stylistically coherent pieces from a text prompt, with opaque and uncontrollable structure — against interactive composition assistants such as FlowComposer and the Continuator, where the musician imposes constraints and retains intentionality and the Continuator learns a player's style in real time as an improvisational partner. Only that second kind, he argues, is educationally productive, because the first positions students as consumers of machine output. The recommended curriculum makes production and DAW fluency core competencies and defends ensemble performance as the embodied, social dimension most resistant to substitution. Assessment in music is being pushed toward automation, with limits worth noting. [[gpt4o-mini-music-analysis-scoring|Lin, Jin and Min (2026)]] benchmarked GPT-4o-mini against teacher mean scores on 300 university-level music analysis responses across Harmony, Form, Reasoning and Terminology, finding that few-shot chain-of-thought prompting agreed most strongly with teacher means and that [[rag|retrieval augmentation]] systematically over-scored; agreement was also weaker on Terminology than on Reasoning, so the authors argue for strategy-specific calibration, [[assessment-validity|dimension-level validation]] and continued [[human-in-the-loop-ai|human oversight]]. Vocal pedagogy supplies the clearest statement of the measurement-versus-value problem. [[ai-vocal-pedagogy-2026|Li (2026)]] argues that AI-assisted vocal teaching should not be judged by how precisely it measures pitch, stability, vibrato and timing, because the same measured deviation may indicate technical inaccuracy, expressive inflection or a recording artifact, and because measurable outputs do not represent a learner's [[embodied-learning|embodied]] coordination. The framework separates the evidence AI makes visible, the human learning processes that interpret it (bodily awareness, [[metacognition|metacognitive]] monitoring, [[self-regulated-learning|self-regulated practice]], motivation), and the outcomes that follow, judged by effectiveness, [[equity-in-ai-education|equity]] and [[sustainability]]. Instrumental teaching for disabled musicians takes the embodiment point further: [[embodied-string-learning-blindness-low-vision-musicians|Shi et al. (2026)]] worked with four advanced blind and low-vision string musicians and three instructors, finding that bowed-string technique is conventionally taught by visual demonstration — unavailable to these learners — and that tactile and kinesthetic cues substitute for it. [[inclusive-learning|Inclusive instruction]], on their account, should be built with disabled learners rather than retrofitted for them. Stage and performance technology is where AI-assisted instruction has been designed and evaluated most concretely. [[luminote-llm-vr-stage-lighting-education-2026|Liang et al. (2026)]] built LumiNote, an [[llm]]-assisted VR system that turns an instructor's spoken intent, anchored by a laser pointer, into reviewable spatial annotations, executable lighting demonstrations and jargon explanations. Across 55 instructor prompts and 531 generated action pairs, 380 (71.6%) were applied; visual effects formed the largest prompt category (27 of 55) with high adoption (212 of 245 actions, 86.5%), while fixture-specific requests proved weakest — 26 of 28 rejections traced to directional-reference misreads such as "left light". Instructors used suggestions as a refinement process, not a finished lesson plan: of 147 rejected or modified suggestions with follow-up, 127 (86.4%) triggered a new prompt and only one a direct manual adjustment. Their workload fell (NASA-TLX 3.50 to 2.22) and sessions shortened (19 minutes 34 seconds to 12 minutes 33 seconds) as effort shifted from low-level configuration toward expression. The sharpest finding is a representation mismatch: the cues instructors rated most useful for externalizing expert reasoning were not the ones 24 learners could follow, who preferred a laser pointer (M = 5.50) and console tagging with values (M = 5.33) over spatial arrows and an avatar (both M = 4.58), though student outcome measures showed no significant differences. ## Visual Arts, Image Generation and Digital Media Text-to-image tools are the case where the studio's iterative workflow is most directly at stake, and the evidence is not a simple story of gain or loss. [[t2i-competence-paradox-2026|Liu, Meng and Zhang (2026)]] surveyed 417 art and design students alongside instructor focus groups and interviews, finding that performance expectancy, social influence, novelty value and creative competence all positively predicted intention to use text-to-image tools, while effort expectancy and facilitating conditions predicted negatively — which they read as shortcut-oriented use in coursework where ease and availability enable quick output rather than sustained engagement. Their competence paradox is the notable result: creative competence supports intention but also predicts more selective, restrained actual use, as students weigh authorship, originality and skill preservation against efficiency. They argue for pedagogy addressing [[ai-literacy|AI literacy]] around creative process, prompt crafting and output evaluation, with studio process and effort as the assessed object. Digital media production shows what a structured process can do. [[project-based-digital-storytelling-art-design-2026|Tian et al. (2026)]] studied a 15-week project-based digital storytelling capstone, "Creative Shanzhou", in which 426 final-year undergraduates, guided by 48 mentors at roughly a 1:9 ratio, translated local cultural heritage into [[multimodal]] narratives. Expert ratings of the 92 resulting works were strongest on novelty (M = 4.21, SD = 0.72), followed by effectiveness (3.96) and wholeness (3.68), with 43% of projects scoring below 3.5 on wholeness — students generated original ideas more successfully than they integrated them. In a 31-student Animation subgroup, figural creative-thinking scores rose from 77.23 to 94.68 (mean difference 17.45, t(30) = 4.55, p < .001, Cohen's d = 0.82). Across the analyzed projects AI served as information organizer, visual reference, ideation aid and editing tool, while students retained topic selection, narrative interpretation, cultural meaning-making and final decisions; the highest-scoring example attributed its coherence to sustained human engagement with source material rather than AI polish. The project ended in a three-day public exhibition, and 31 works were taken up by the local Cultural and Tourism Bureau. Immersive environments provide the more favorable counter-case in vocational design. [[ai-ive-pbl-vocational-design-creativity-2026|Jin et al. (2027)]] implemented AI-IVE-PBL — [[project-based-learning|project-based learning]] inside an AI-enabled immersive [[virtual-and-augmented-reality|virtual environment]] with an [[agentic-ai|agent]] teaching assistant — in a first-year vocational interior design course, as a five-phase loop of discovery, envisioning, modeling, communication and refinement. In a 12-week two-group quasi-experiment (63 valid responses, 31 versus 32), the immersive-plus-agent condition scored higher on design ability (η²p = .138) and creative ability (η²p = .111) under ANCOVA with pretest as covariate, and lifted cognitive (d = 0.90) and behavioral (d = 0.75) [[student-engagement|engagement]], [[motivation]] (d = 0.74) and satisfaction (d = 0.69) while lowering reported cognitive load (d = −0.52). Innovative thinking and affective engagement moved in the expected direction without reaching significance. All outcomes were [[self-report-measures|self-report]], with no performance artifacts or expert ratings, which is why the authors contrast their result with the architectural studio study above. ## Craft, Skill and the Assessment of Creative Work The craft question appears in several forms, and the sources disagree about how worried to be. [[genai-architecture-education|Kapsalis (2026)]] reports a levelling mechanism at entry level: students who described themselves as weak at drawing found the tool opened access to visual expression without removing the need for judgment. Against that, [[ai-interior-design-malaysia-2026|Syed Abdul Rahman (2026)]] warns that uncritical adoption risks graduates without foundational spatial reasoning, material knowledge or independent critical evaluation, and names the prevention of deskilling among early-career practitioners as a core [[regulation|regulatory]] concern. The disagreement is partly about sequencing: both favor technical foundations before generative exploration. What is consistent is a repositioning of what gets assessed. The Derby study measured procedural confidence in-session at 3.7 and 3.5 out of 5 but confidence in transferring those skills beyond the studio at only 2.6 — a gap its author attributes to single-session exposure and calls a curriculum-level problem. The Malaysia analysis recommends studio projects requiring comparative [[ai-ed-evaluation|evaluation of AI]] and non-AI design pathways, plus criteria that reward critical thinking alongside visual quality. Music's automated-scoring results point the same way: agreement with teacher means was dimension-specific, and the authors insist on human oversight rather than full delegation. [[ai-vocal-pedagogy-2026|Li (2026)]] adds that treating measured output as educational value in itself narrows vocal training into output correction and score optimization. The converging proposal is that creative assessment keep examining process, iteration and justified decision-making precisely because the finished artifact no longer evidences them. ## Access and Inclusion in Arts Learning [[accessibility]] in the arts is treated less as compliance than as a design and material-practice question. [[embodied-string-learning-blindness-low-vision-musicians|Shi et al. (2026)]] is the clearest case: because bowed-string technique is normally taught by visual demonstration, blind and low-vision musicians depend on tactile and kinesthetic channels, and the study's disability-led co-design produced strategies rooted in those musicians' own practice rather than adaptations bolted onto a visual default. Two other findings complicate any simple story of generative tools as an equaliser. In the Derby study about a third of participants declared a disability and nearly a fifth reported a specific learning difficulty; inclusivity items clustered at 3.8–4.1 and correlated strongly with feeling the session was approachable (ρ = .74) and with keeping up regardless of prior AI use (ρ = .81), and students less confident at drawing reported not feeling disadvantaged — a modest equalizing effect in a single 90-minute session, not demonstrated learning. The LumiNote representation gap shows the other side: cues that help experts externalize their reasoning are not the cues novices can follow, so a design that serves the instructor well can leave learners adrift, and the authors position the LLM as a mediation layer that must translate between the two. The GenARch study records a related equity decision: after data collection the [[ai-technologies|technologies]] were released to the control students so they were not withheld. Whose aesthetic and cultural references generative models reproduce, and whether wider access to design services expands demand or intensifies mid-market competition, are raised in the Malaysia analysis and remain open. ## Connected Concepts - [[creativity]] - [[design-thinking]] - [[generative-ai]] - [[project-based-learning]] - [[embodied-learning]] - [[authentic-assessment]] - [[ai-literacy]] - [[equity-in-ai-education]] - [[virtual-and-augmented-reality]] - [[discipline-specific-aied]] - [[design-education]] ## Connected Articles - [[genai-xr-architectural-design-education-2026]] — GenAI plus multi-user XR in a real architectural studio: complementary roles, declining design self-efficacy, no portfolio advantage (Xiao et al. 2026) - [[genai-architecture-education]] — A locally executed, discipline-specific image workflow in an architecture studio: wider creative search space and reduced crit anxiety (Kapsalis 2026) - [[genai-architectural-design-studios]] — Longitudinal studio study and the GAI-A platform: GenAI as stimulus and solution-space expansion, with design fixation as the risk (Gül et al. 2026) - [[ai-interior-design-malaysia-2026]] — Interior designers as curators of machine output, and what curriculum and professional regulation should do about it (Syed Abdul Rahman 2026) - [[ai-ive-pbl-vocational-design-creativity-2026]] — Project-based learning inside an AI-enabled immersive environment in vocational interior design (Jin et al. 2027) - [[project-based-digital-storytelling-art-design-2026]] — A 15-week digital storytelling capstone with 426 students: novelty strong, integration weak, AI as tool not author (Tian et al. 2026) - [[t2i-competence-paradox-2026]] — Text-to-image acceptance among 417 art and design students, and the competence paradox of selective use (Liu, Meng & Zhang 2026) - [[musical-education-ai-digital-transformation-2026]] — Streaming economics, autonomous generators versus composition assistants, and rethinking the music curriculum (Briot 2026) - [[gpt4o-mini-music-analysis-scoring]] — Benchmarking GPT-4o-mini against teacher means on 300 music analysis responses (Lin, Jin & Min 2026) - [[ai-vocal-pedagogy-2026]] — Why measurement precision is not educational value in AI-assisted vocal teaching (Li 2026) - [[embodied-string-learning-blindness-low-vision-musicians]] — Disability-led design of tactile and kinesthetic string instruction for blind and low-vision musicians (Shi et al. 2026) - [[luminote-llm-vr-stage-lighting-education-2026]] — LLM-assisted VR instruction for stage lighting, and the gap between expert and novice representations (Liang et al. 2026) --- ## [Design Education](https://edtechdev.github.io/aied/concepts/design-education/) > **Design Education** — the professional formation of designers in studio-based disciplines: product, service, interaction, interior and architectural design, taught through iterative project work in which the assessed object is the process behind an artifact as much as the artifact itself. Its distinguishing features are the studio and the critique as the site of learning, visible process evidence (sketches, intermediate models, iterations) as the primary indicator of learning, and the fact that [[generative-ai]] now makes the polished output students are graded on cheap to produce. That makes design education a test case for [[creativity]], [[design-thinking]], [[assessment-validity]] and [[professional-training]]. ## Questions to Consider - If a studio can no longer read learning off the finished artifact, what should replace visible process as evidence — mandatory process artifacts, oral examination, something that survives a cohort of 140? - Unsupported novices in one experiment never revised a single idea until an adversarial agent surfaced stakeholder pushback. Should design teaching manufacture friction deliberately — and where does productive difficulty end? - Designers are described as shifting from form-generators to curators of machine output. If so, what comes first in a curriculum: manual craft, or critical evaluation of model output? - If an LLM can approximate instructor, peer and grant-reviewer readings of the same poster, what is left for the human critic? ## Introduction Design education is narrower than [[arts-design-and-media-education]], which covers studio and performance disciplines together — architecture, music, visual art, media and digital storytelling. This page covers the professional design disciplines and their formation pipeline: portfolio assessment, accreditation expectations, employability, and a practice culture that has absorbed generative tools faster than its curricula have. It is also distinct from [[design-thinking]], a human-centered problem-solving method taught across business, education and law as a transferable skill. In design education the process is not a method applied to a discipline; it is the discipline, and the studio is where it is taught, supervised and assessed. [[engineering-education]] is the closest professional neighbour — also design-heavy, also accountable to practice — but engineering assessment rests on calculable artifacts and public-safety responsibility, whereas design's rests on critique and on process evidence that generative tools can visibly simulate. ## How AI appears in design education - **AI is an ideation resource whose characteristic risk is fixation.** [[genai-architectural-design-studios|Gül et al. (2026)]] built the GAI-A platform across a longitudinal studio study and found two productive uses: GenAI as visual stimulus in early ideation, and as a combinatorial expander of the solution space during development. The problem, they argue, is competence rather than access — without the discernment to read models deliberately, abundant output produces [[design-thinking|design fixation]] and aesthetic lock-in — which moves the [[teacher-role]] from transmitting craft to coaching when and why to delegate creative exploration. - **A discipline-specific local tool widened the search space and levelled participation.** [[genai-architecture-education|Kapsalis (2026)]] ran two Derby focus groups of eight architecture students through a 90-minute session with a locally executed ComfyUI workflow: creative scores were highest for idea generation (4.5) and experimentation (4.2) and lower for refining decisions (3.7), while inclusivity items ran 3.8–4.1 and correlated with keeping up regardless of prior AI use (ρ = .81). Confidence gained in the session did not transfer — confidence in using these skills professionally was 2.6, a gap the author calls a curriculum-level problem. - **Unsupported novices do not revise; an adversarial agent changes that.** [[ai-agents-constructive-conflict-design-education-2026|Han & Martelaro (2026)]] ran a between-subjects experiment with 45 novice interaction designers on a civic reporting site. Their Self Reflection group never revised or deleted a single idea; agent-engaged students made roughly 3.6 times more edits, reached stakeholder concerns the other groups missed ([[accessibility]], [[privacy]], reluctance toward automation) and produced stronger proposals. The agent was rated as contributing (M = 5.33), but one participant felt criticized "in every aspect", and the authors insist simulated pushback is no substitute for engaging real publics. - **A structured process keeps authorship with students at scale.** [[project-based-digital-storytelling-art-design-2026|Tian et al. (2026)]] studied a 15-week storytelling capstone in which 426 undergraduates and 48 mentors (about 1:9) turned local heritage into [[multimodal]] narratives. Expert ratings of the 92 works were strongest on novelty (M = 4.21) and weakest on wholeness (M = 3.68) — ideas came more easily than integration. Creative-thinking scores in a 31-student subgroup rose from 77.23 to 94.68 (d = 0.82). AI organized information and edited; students kept topic choice, interpretation and final decisions, and 31 works were adopted by a regional Cultural and Tourism Bureau. - **Adoption is frequent, uneven across the process, and suspicious.** [[genai-usage-design-students-survey|Broadbent et al. (2026)]] surveyed 280 Politecnico di Milano design students and read 100 Masters AI-use journals: 71% used GenAI daily or weekly, 65% distrusted it, 85% modified its output. Use clustered in research and writing — editing writing 77.5%, finding information 74.5%, brainstorming 68.0% — with whole-assignment use rare (18.6%). Ownership largely survived: 31% reported lost ownership or creative agency, and only prototype generation was significantly associated with it. Planning and interpreting primary research stayed with students, which the authors read as evidence that students may not count research as part of the creative process. - **The net effect is divergence rather than uplift.** [[ai-mediated-cognitive-divergence-2026|Crolla, Xia & Jiang (2026)]] drew on 24 interviews and a 32-instructor survey in a built-environment faculty to describe *cognitive divergence*: AI amplifies prior differences, letting stronger students extend reasoning while weaker ones delegate formative work and produce coherent outputs without understanding — a "false sense of capability" that collapses under questioning. Loss of process visibility, in a field that relied on visible process as its primary evidence of learning, and erosion of the frictional stages that act as [[desirable-difficulties]] compound it. Faculty positivity tracked pedagogical relevance (β = 0.534, p = .004), not seniority, and the faculty's mitigations do not scale past studio-sized cohorts — which argues for institutionalized process-evidence requirements. Limits: single institution, n = 32, faculty perception. - **Disciplines differ in which behaviours pay off.** [[same-ai-different-pathways|Shen (2026)]] compared 221 design students with 222 business students and found [[ai-literacy]] predicted both prompting proficiency and verification behavior (p < .001). Verification directly predicted task quality in the design cohort, calibration accuracy in business — open-ended ideation against audit-like analytic standards — and verification functioned as a metacognitive safeguard against [[cognitive-offloading]]. - **Studio judgement can be approximated when a rubric mediates.** [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] raised LLM–human agreement on 80 design posters from 54.75% to 81.25% by iteratively refining rubric descriptors; instructor, peer-reviewer and grant-reviewer prompts produced epistemically different feedback on the same work. The authors treat the rubric as a revisable interface between pedagogical intent and machine inference, retaining [[human-in-the-loop-ai|human oversight]] against hallucinated rationale. - **Students negotiate the tools rather than simply accept them.** [[t2i-competence-paradox-2026|Liu, Meng & Zhang (2026)]] surveyed 417 art and design students: performance expectancy, social influence, novelty value and creative competence predicted intention to use text-to-image tools, while effort expectancy predicted negatively, read as shortcut-oriented coursework use. Their competence paradox is that creative competence supports intention yet predicts more selective use as students weigh authorship and skill preservation — an [[assessment-validity]] problem in a studio where process is the assessed object. [[rana-genai-design-thinking-2025|Rana et al. (2025)]] reached a compatible conclusion from 112 reflections in a 12-week [[design-thinking]] course: benefits dominated (86% positive sentiment), ethical concerns drove 62% negative sentiment, and scaffolded integration moved students from skepticism to 72% positive orientation. ## Professional formation and regulation Assessment problems here are inseparable from professional ones. [[ai-interior-design-malaysia-2026|Syed Abdul Rahman (2026)]] describes Malaysian interior designers shifting from primary form-generators toward critical mediators and curators of machine output, with [[visualization]] platforms compressing timelines while generated schemes still lack cultural specificity, climatic responsiveness and reliable constructability. Its curriculum recommendation is sequencing rather than substitution — CAD/BIM, construction knowledge and human factors before generative exploration, and criteria rewarding [[critical-thinking]] and ethical reflection. The professional side is unresolved: interior design regulation is less formalized than architecture's, leaving open questions of disclosure, accountability for algorithmic error, and deskilling among early-career practitioners. The analysis is document-based and single-country. ## Connected Concepts - [[design-thinking]] - [[creativity]] - [[generative-ai]] - [[ai-literacy]] - [[higher-ed]] - [[arts-design-and-media-education]] - [[engineering-education]] - [[professional-training]] - [[assessment-validity]] - [[assessment]] - [[cognitive-offloading]] - [[critical-thinking]] - [[scaffolding]] - [[desirable-difficulties]] - [[human-in-the-loop-ai]] - [[curriculum-design]] - [[teacher-role]] - [[equity-in-ai-education]] - [[discipline-specific-aied]] - [[regulation]] ## Connected Articles - [[ai-agents-constructive-conflict-design-education-2026]] - [[ai-mediated-cognitive-divergence-2026]] - [[ai-interior-design-malaysia-2026]] - [[genai-architecture-education]] - [[genai-architectural-design-studios]] - [[genai-usage-design-students-survey]] - [[project-based-digital-storytelling-art-design-2026]] - [[t2i-competence-paradox-2026]] - [[same-ai-different-pathways]] - [[yasar-llms-iterative-pedagogical-design-2026]] - [[rana-genai-design-thinking-2025]] --- ## [Medical and Health Professions Education](https://edtechdev.github.io/aied/concepts/medical-education/) > **Medical and Health Professions Education (HPE)** — the teaching and training of medical, nursing, pharmacy, and allied health professionals. AI is reshaping this domain through clinical [[simulation]], [[reinforcement-learning|reinforcement learning]] trainers, [[adaptive-learning|adaptive learning]], and the application of foundational learning principles (experiential, situated, and distributed cognition) in health-professions contexts. Because HPE is high-stakes, competency-based, and clinically embedded, it raises distinct questions about AI's role in skill acquisition, patient safety, and the educator's judgment. Nursing, the most heavily studied program in HPE, now has a page of its own — [[nursing-education]] — where the competence boundary, identity-formation, and workforce evidence that this page carries only in outline is developed in full. ## Questions to Consider - In medicine, AI benefits like scalable practice and adaptive feedback must be balanced against risks like erosion of hands-on clinical skill — where errors carry direct patient consequences. Where would you draw the line between what AI should do and what a trainee must practice for themselves? - The page argues AI should be used to operationalize age-old learning principles — experiential, situated, distributed cognition — rather than replace the educator's guiding role. What makes the educator's judgment indispensable even when AI can simulate or personalize the practice? - Reinforcement-learning trainers and [[agentic-ai|agentic AI]] are now used for clinical and procedural skills in residency. If an AI agent trains a resident on a procedure, how would you verify they've actually learned it safely before they do it on a patient? - Because health-professions education is high-stakes and competency-based, assessment questions carry particular weight. How might AI-assisted assessment both improve and threaten the evaluation of clinical competence? - Over-reliance on AI is a specific concern in medicine. How do you think training with AI could produce a clinician who is more confident but less able to reason independently — and what could guard against that? ## Introduction AI in medical and health-professions education is a growing strand of the knowledge base's subject-area coverage. Unlike general [[higher-ed|higher education]], HPE is oriented toward the development of clinical competencies, procedural skills, and professional judgment, which shapes how AI tools are designed and evaluated. ### How AI appears in health-professions education - **Operationalizing foundational learning principles.** [[fowlin-operationalizing-learning-principles-ai|Fowlin et al.]] (developed at the Medical University of South Carolina) argue that AI should be used to operationalize age-old learning principles — Dewey's [[experiential-learning|experiential learning]], [[situated-learning|situated cognition]], and [[distributed-cognition|distributed cognition]] — rather than replace the educator's guiding role. AI enhances [[personalized-learning|personalized]] and [[adaptive-learning|adaptive]] learning while the teacher remains central to [[student-engagement|student engagement]] and outcomes. - **Clinical skills and reinforcement learning.** [[residencyrl-clinical-rl-training-2026|ResidencyRL]] uses reinforcement learning to train clinical reasoning and procedural skills in residency, demonstrating AI as a skills-training partner in real clinical workflows. - **Simulation and agentic AI.** [[hdr-brachytherapy-agentic-ai-simulation-2026|Agentic AI simulation in brachytherapy]] shows how AI agents support hands-on procedural training in medical specialties. - **Gamification of medical learning.** [[medgame-llm-medical-education-gamification|MedGame]] applies [[game-based-learning|gamification]] and LLMs to engage medical students. - **Nursing and interdisciplinary education.** [[alrazeeni-transforming-nursing-education-ai-2026|AI transformation of nursing education]] documents how AI reshapes nursing curricula and instruction. This page keeps nursing as one program among medicine, pharmacy, dentistry, and allied health, organized around the shared problem of scalable clinical competence; the nursing-specific record — a competence boundary where AI's documented benefits stop, an explicitly professional-identity framing, and students' workforce anxiety — belongs to [[nursing-education]] and is cross-referenced here rather than restated. - **AI-powered simulation in nursing.** [[jiang-ai-powered-simulation-nursing-education-2026|Jiang et al. (2026)]]'s [[mixed-methods-research|mixed-methods]] [[meta-analysis-systematic-review|systematic review]] of 19 studies (N = 1,253) finds AI-driven [[simulation|simulations]] (GenAI/LLMs, virtual patients/mannequins, AI-enhanced VR/MR, and [[conversational-ai|chatbots]]) significantly improve cognitive knowledge and [[affective-computing|affective]] outcomes ([[self-efficacy]], communication confidence) in the strongest designs, but are inconsistent for complex psychomotor skills — one [[rct]] found AI-assisted simulation *inferior* to standardized patients. Their concept of an **"authenticity gap"** — a learner-perceived shortfall in emotional resonance, nonverbal cues, and tactile examination — grounds why AI is best for highly structured objectives (foundational communication, history-taking) and should sit in a **stepped simulation continuum** alongside, not instead of, human-standardized patients and clinical placement. This is a distinctive, evidence-anchored refinement of the domain's [[simulation]] strand and parallels the acceptance-trust finding that learners favor human feedback perceived as "benevolent" over AI seen as merely "competent." - **LLMs and the reconstitution of [[learner-identity|professional identity]] in nursing (2026).** [[sun-llm-nursing-education-professional-identity-2026|Sun et al. (2026)]]'s critical integrative review of 489 studies across 47 countries (Whittemore & Knafl framework) shifts the question from what [[llm|LLMs]] can do in nursing education to what their integration does to the *developmental processes* through which a nurse is formed. Framed through Benner's skill acquisition, cognitive load theory, automation bias, and Wenger's identity formation, they find the same technology both enhances and erodes competence depending on whether the displaced cognitive work is extraneous to, or constitutive of, the target competence — unstructured reliance produced measurable deficits in ethical reasoning and clinical judgment. Their **Professional Identity Tension Model** distinguishes LLM integration that supports professional formation from integration that silently substitutes for it, and **Structural Empathy Suppression** reframes apparent AI "outperformance" on relational metrics as a symptom of systemic overwork rather than genuine AI empathy. An evidence gap map shows the highest-policy-consequence domains (professional identity, relational/ethical competency, long-term outcomes) rest on the thinnest rigorous evidence — a caution against confident deployment claims in nursing education. - **Multi-agent AI standardized patients in clinical interview training.** In a 2026 [[rct|randomized controlled trial]] with 95 analyzed medical students ([[ai-standardized-patient-scaffolding-medical-2026|Yang et al.]]), multi-agent AI standardized patient training raised final OSCE-aligned examination scores over a structured non-LLM progressive-disclosure control (71.8% vs. 55.6%; β = 16.4 percentage points; P = 3.30e-4) while leaving binary diagnostic accuracy statistically identical (84% vs. 86%; P = 1.000). The dissociation indicates that early clinical-training gains can surface first in consultation process quality — communication (mean 3.53 vs. 2.64 on a 1–5 OSCE scale; P = 4.50e-4), empathic expression (+31 percentage points on the "expressing empathy" checklist item), and selected history-taking behaviors — before any improvement in diagnostic endpoints, and that AI patients still need human raters and faculty for professionalism and readiness judgments. - **Task allocation and the SCAN framework.** [[ai-teammate-task-distribution-medical-training-2026|Tsim et al.]] reframe AI integration from learner "misuse" to *misclassification* — a failure of real-time [[metacognition|metacognitive]] evaluation. Their SCAN framework (Substitute, Complement, Aid, Non-Negotiable), grounded in Vygotsky's [[sociocultural-learning|zone of proximal development]], allocates [[generative-ai|generative AI]] tasks by the individual learner's developmental state and identifies passive engagement within AI-scaffolded tasks as a hidden pathway to mis-skilling that requires re-identification from AI to expert assistance with human [[human-in-the-loop-ai|epistemic auditors]]. - **Human-in-the-loop instructional asset generation.** [[gen-mentor-dental-radiography-2026|Gen-Mentor]] (Dong et al. 2026) integrates a vision-language-model backbone into a dental-radiography workflow: Faster R-CNN localizes four target findings (Filling, Implant, Impacted Tooth, Cavity), a conditional diffusion model generates class-specific synthetic ROI candidates, a VLM produces evidence-linked captions, and an [[llm]] reformats them into case descriptions, comparisons, and quiz prompts — all before structured expert review. Evaluated with 45 dental students (mean SUS 72.7), it demonstrates how [[human-in-the-loop-ai|human-in-the-loop]] review of AI-generated instructional assets can expand case diversity and immediate-feedback support while retaining expert oversight over what students see. - **[[automated-assessment|AI grading]] of pharmacy exams.** [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat, Das, Bhaumik & Thambi (2026)]] evaluated ChatGPT-5 against human faculty grading of a 21-item pharmacy exam (16 students) across multiple-choice, select-all-that-apply, fill-in-the-blank, listing, short-answer, and essay items. The model matched faculty closely on objective items (CCC 0.935–1.000) but was unreliable for listing, short-answer (CCC ≈0), and essay (0.341–0.854) responses, and a structured rubric did not consistently improve agreement — evidence that in high-stakes, competency-based [[assessment]] in health professions, AI suits well-specified items while [[human-in-the-loop-ai|human review]] remains necessary for subjective, open-ended clinical reasoning. - **GenAI in scenario-based healthcare education.** [[genai-scenario-based-healthcare-education-2026|Neto and colleagues (2026)]] systematically reviewed 23 studies of GenAI across scenario-, case-, problem-, and simulation-based learning in healthcare education (PRISMA 2020). Their central finding is that **[[prompt-engineering|prompt design]] functions as instructional specification** — encoding the cognitive targets and quality criteria implicit in expert authoring — yet only 34.8% of studies aligned generated content with instructional frameworks and only 34.8% reported prompting in enough detail to reproduce. GPT-4 dominated implementations (44.4%), hybrid [[human-ai-collaboration|human-AI collaboration]] outperformed fully automated approaches, and evidence was strongest for higher-order cognitive skills but inconsistent elsewhere. This grounds [[simulation|scenario/simulation-based]] and [[problem-based-learning|problem-based]] [[pedagogy]] in medical education with validation- and integration-quality [[benchmark|benchmarks]]. - **AI scoring of open-ended exam questions.** [[olvet-genai-scoring-open-ended-medical-2026|Olvet et al. (2026)]] tested whether [[generative-ai|GPT-4]] could reliably score open-ended questions on pre-clerkship [[assessment|assessments]] at two US medical schools. With faculty iteratively refining scoring rubrics across three rounds of error-pattern analysis, AI–faculty inter-rater reliability reached substantial-to-almost-perfect agreement on three of four questions (weighted kappa up to 0.94) but only moderate on the holistic-rubric item (κw = 0.54); discrepancies traced to both raters (GPT-4 over-scoring multiple-answer or rubric-absent-vocabulary responses; faculty being overly generous) and occasional feedback inaccuracies keep [[human-in-the-loop-ai|humans in the loop]]. The authors argue the case for automated OEQ scoring is strengthened because ~82% of US medical schools grade pre-clerkship work pass/fail, where exact AI score agreement is not always required. - **[[ai-ed-evaluation|Evaluating AI]] teaching agents, not just deploying them.** [[zhang-platform-scores-miss-ai-teaching-agents-2026|Zhang et al. (2026)]] deployed eight LLM teaching agents across four role-play paradigms (patient, student, expert, family member) in an endocrinology [[curriculum-design|curriculum]] and scored 167 student dialogues with an 8-dimension teaching-quality rubric. Platform scores diverged sharply from rubric quality; agents differed most on knowledge dimensions and least on role enactment; adaptive difficulty calibration was a shared weakness; and strengthening an agent's empathy raised role-play quality without improving knowledge coverage. The finding that role-play paradigms can be separated (beyond the default patient-doctor script) points to designing and evaluating agents against explicit pedagogical dimensions. ### Why it matters HPE is a high-stakes, competency-based domain where AI's benefits (scalable practice, adaptive feedback, simulation) must be balanced against risks ([[cognitive-offloading|Over-Reliance]], erosion of hands-on clinical skill, ethical and safety concerns). The knowledge base's general concepts — [[teacher-role]], [[assessment]], [[feedback]], [[equity-in-ai-education]], and [[ethics]] — apply with particular intensity in health professions, where errors carry direct patient consequences. ## Implications for health-professions educators - **Use AI to operationalize learning principles, not replace the educator.** [[fowlin-operationalizing-learning-principles-ai|Fowlin et al.]] argue AI should operationalize experiential, situated, and distributed-cognition learning while the teacher remains central to engagement and outcomes. - **Leverage AI for clinical skills training.** [[residencyrl-clinical-rl-training-2026|ResidencyRL]] and [[hdr-brachytherapy-agentic-ai-simulation-2026|agentic simulation]] show AI as a skills-training partner in real clinical workflows — embed it where it adds safe, scalable practice. - **Balance high-stakes benefits against over-reliance.** HPE is competency-based and high-stakes; guard against AI substituting for hands-on clinical skill and judgment, and apply [[feedback]], [[assessment]], and [[ethics]] considerations with particular care. - **Adapt gamified and interdisciplinary AI thoughtfully.** [[medgame-llm-medical-education-gamification|Gamified LLM learning]] and [[alrazeeni-transforming-nursing-education-ai-2026|nursing-education transformation]] show promise but need evaluation for safety and skill outcomes; for nursing specifically, that safety-and-skill evidence — including the RCT in which AI-assisted simulation underperformed standardized patients — is gathered on [[nursing-education]]. ## Connected Concepts - [[problem-based-learning]] - [[higher-ed]] - [[simulation]] - [[adaptive-learning]] - [[personalized-learning]] - [[game-based-learning]] - [[experiential-learning]] - [[situated-learning]] - [[distributed-cognition]] - [[teacher-role]] - [[assessment]] - [[feedback]] - [[ai-education]] - [[discipline-specific-aied]] - [[nursing-education]] — the nursing strand of health-professions education - [[virtual-and-augmented-reality]] — immersive and AR clinical training ## Connected Articles - [[ai-standardized-patient-scaffolding-medical-2026]] — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training - [[akbaba-nursing-ai-experiences-tam-2026]] — Nursing students' and faculty AI experiences (TAM; psychosocial support) - [[zhang-platform-scores-miss-ai-teaching-agents-2026]] — 8-dimension rubric evaluation of AI teaching agents in medical education - [[sun-llm-nursing-education-professional-identity-2026]] — LLMs, nursing education structural gaps, and the reconstitution of professional identity (Sun et al. 2026) - [[ai-teammate-task-distribution-medical-training-2026]] — SCAN framework: rethinking AI task distribution in medical training (Tsim et al. 2026) - [[genai-simulate-patient-history-pbl-2026]] - [[fowlin-operationalizing-learning-principles-ai]] — Operationalizing experiential, situated, and distributed cognition with AI in health-professions education - [[residencyrl-clinical-rl-training-2026]] — Reinforcement-learning training for clinical skills in residency - [[medgame-llm-medical-education-gamification]] — Gamified LLM-based learning for medical education - [[hdr-brachytherapy-agentic-ai-simulation-2026]] — Agentic AI simulation for brachytherapy training - [[alrazeeni-transforming-nursing-education-ai-2026]] — Transforming nursing education with AI - [[jiang-ai-powered-simulation-nursing-education-2026]] — AI-powered simulation in nursing: mixed methods systematic review - [[wang-safety-gap-productive-struggle-2026]] — The Safety Gap: Restoring Productive Struggle - [[gen-mentor-dental-radiography-2026]] — Gen-Mentor: human-in-the-loop dental radiography instruction (Dong et al. 2026) - [[genai-scenario-based-healthcare-education-2026]] — Systematic review of GenAI in scenario-based healthcare education (Neto et al. 2026) - [[falahat-chatgpt-grading-pharmacy-exams-2026]] - [[olvet-genai-scoring-open-ended-medical-2026]] - [[sophie-clinical-communication-ai-assessment-2026]] — Scalable AI-based clinical communication training and automated assessment --- ## [Nursing Education](https://edtechdev.github.io/aied/concepts/nursing-education/) > **Nursing Education** — the preparation of nurses for licensed practice, and the subfield of [[medical-education|health-professions education]] where AI is currently most studied. AI enters nursing curricula through [[simulation]] with virtual patients and mannequins, [[llm|LLM]]-based study and clinical-reasoning support, [[automated-assessment|automated assessment]], and adaptive platforms for at-risk learners. Its object is distinctive: nursing competence fuses psychomotor skill, relational practice, and [[learner-identity|professional identity]] formation, so a technology that raises measured performance can simultaneously erode the developmental work that produces a nurse. ## Questions to Consider - AI-supported simulation reliably improves knowledge and [[self-efficacy]] but shows inconsistent, sometimes negative effects on complex psychomotor skill. Where is AI the right [[teacher-role|teacher]], and where must a human body in the room remain non-negotiable? - Nursing students describe AI as emotional comfort during clinical stress. Is that a support to design for, or a signal that the relational [[sociocultural-learning|apprenticeship]] is under-resourced? - AI anxiety among health-sciences students tracks job-search anxiety. Is teaching AI literacy the remedy, or does it raise the salience of a threat that [[educational-policy-ai|institutional policy]] should address instead? ## Introduction Nursing education is the clinical strand of the knowledge base with the densest empirical record: system-wide reviews of AI applications, a [[mixed-methods-research|mixed-methods]] synthesis of AI-powered simulation, a [[qualitative-research|qualitative]] study of student and faculty experience, and a critical integrative review of LLM research sit alongside one another. That makes nursing a useful test case for claims the wider [[ai-education|AI in education]] field makes about clinical training. It must be kept distinct from [[medical-education|Medical and Health Professions Education]], which treats nursing as one program among medicine, pharmacy, dentistry, and allied health and organizes its account around the shared problem of scalable clinical competence. The nursing literature adds three things that page cannot carry: a competence boundary at which AI's documented benefits stop; an explicitly *professional-identity* framing, in which the question is who becomes a nurse rather than what a trainee can do; and a workforce layer — students' anxiety about AI displacing nursing roles, and the faculty view of nurses as shapers rather than recipients of healthcare's digital transformation. The field also refines three adjacent concepts. From [[simulation]] it inherits the fidelity debate but renames the deficit an "authenticity gap." From [[affective-computing]] it takes the measurement of affect, yet its strongest claim is that affective outcomes improve while emotional depth in the interaction does not. From [[self-efficacy]] it takes the central mediating variable, while cautioning that confidence built in a low-risk setting may not transfer to the clinical one. Beneath all of it sits the ordinary [[higher-ed|higher education]] problem of workload, access, and assessment integrity, sharpened by licensure and patient consequence. ### How AI appears in nursing education - **Four application areas, recurring risks.** [[alrazeeni-transforming-nursing-education-ai-2026|Alrazeeni et al. (2026)]] reviewed 28 empirical studies (2010–April 2025) and group AI applications into [[personalized-learning|personalized learning]], [[simulation|simulation-based training]], [[automated-assessment]], and institutional [[curriculum-design|curriculum]] management with predictive [[learning-analytics|analytics]]. Technological inequity, faculty preparedness gaps, and privacy and bias concerns recur. Their recommendations are concrete — embed AI simulation in emergency-care training, deploy [[adaptive-learning|adaptive platforms]] for at-risk learners, use automated tools for real-time [[formative-assessment|formative]] [[feedback]], adopt diagnostic accuracy as the impact measure — within 6–12 month multi-site pilots tracking [[learning-gains|learning outcomes]] and trust. - **Simulation: strong on knowledge and confidence, conditional on skill.** [[jiang-ai-powered-simulation-nursing-education-2026|Jiang et al. (2026)]]'s PRISMA-guided mixed-methods review of 19 studies (N = 1,253, mostly prelicensure students) covers [[generative-ai|generative AI]]/LLMs, AI-driven virtual patients and mannequins, AI-enhanced VR/[[virtual-and-augmented-reality|mixed reality]], and [[conversational-ai|chatbots]]. The strongest designs — three [[rct|RCTs]] plus controlled quasi-experiments — show significant gains in cognitive knowledge and affective outcomes including [[self-efficacy]] and communication confidence, but effects on complex psychomotor skills are inconsistent, and one RCT found AI-assisted simulation *inferior* to standardized-patient simulation. Qualitative meta-aggregation explains both the appeal and the limit: learners value safe, repeatable, nonjudgmental practice that bridges the theory–practice gap, but report an **authenticity gap** — robotic dialogue, missing nonverbal cues, no tactile examination — plus technical instability that raises extraneous [[cognitive-offloading|cognitive load]] and state anxiety. The authors recommend a stepped continuum: AI for pre-learning, history-taking, and foundational reasoning; human simulators and standardized patients for complex psychomotor and emotionally loaded work; human-facilitated debriefing alongside [[ai-feedback-quality|AI feedback]]. - **Acceptance is real, and relational.** [[akbaba-nursing-ai-experiences-tam-2026|Akbaba and Calik Kus (2026)]] interviewed 28 participants (16 students, 12 faculty) and analyzed transcripts deductively through the [[technology-acceptance-model|Technology Acceptance Model]]. TAM held: ease of use, usefulness, intention, and use shaped adoption. The role split is practical — students used AI for presentations, visual content, clinical case analysis, and care planning; faculty for course materials, [[meta-analysis-systematic-review|literature review]], [[writing-education|academic writing]], and administration — implying tailored rather than uniform training. Two findings extend TAM's cognitive frame. Participants described AI as psychosocial support, a "companion" offering reassurance and a confidential space to reflect during clinical stress. And where AI *scored higher than nurses* on empathy in a high-volume context, [[sun-llm-nursing-education-professional-identity-2026|Sun et al. (2026)]] read it as **Structural Empathy Suppression**: overwork makes authentic empathic expression unsustainable, algorithmic consistency fills the gap, and identity-constituting significance transfers to the algorithm — reframing the policy question from "how effective is this tool?" to "what conditions made it appear necessary?" - **Faculty read the risk as professional, not procedural.** [[dabkowski-nursing-academics-genai-2026|Dabkowski et al. (2026)]] interviewed 22 nursing academics across Australia and New Zealand and found them navigating policy that was absent, late, or written without them, sharply divergent collegial attitudes, and a line drawn at *replacement* rather than use. The objection was developmental before it was procedural: generating answers for a deteriorating-patient scenario means the reasoning is never practiced, and participants warned of a future cohort not fit for practice — "Copilot won't teach you to be a nurse." Assessment practice was already moving toward invigilated exams, oral vivas and process evidence, and misconduct was read as a rehearsal for unsafe clinical shortcuts, which reframes [[academic-integrity]] as a question about who is safe to practice rather than a compliance process. The academics asked for nursing-specific guidance tied to the profession's [[ethics|ethical]] codes, GenAI literacy embedded across the [[curriculum-design|curriculum]], and [[teacher-ai-competency|staff development]] with unit co-design — not blanket prohibition. - **Anxiety as a workforce signal.** [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al. (2026)]] surveyed 821 health-sciences students (nursing among them) and found a moderate positive correlation between [[anxiety-and-stress|AI anxiety]] and job-search anxiety (r = 0.233, p < 0.001), with AI anxiety a significant predictor after controlling for socio-demographics (β = 0.234, p < 0.001). AI anxiety, they argue, is not a technology attitude but a psychological factor shaping how students see their professional futures — pointing to [[ai-literacy]] and career counseling, not tooling, as the intervention. - **Groupthink and [[collaborative-learning|interprofessional]] teams.** [[genai-counter-learner-groupthink-2025|Wiss et al. (2025)]] placed a generative [[agentic-ai|AI agent]] (CALIE) into twelve newly formed interprofessional teams across seven health-professions programs, including nursing, during a 180-minute virtual [[problem-based-learning|problem-based learning]] session, [[prompt-engineering|prompting]] it to inject controversial viewpoints. Of 165 learners, 158 completed the survey; the agent was rated most useful as a tool (M = 3.49), then for feedback helpfulness (M = 3.21), and least as part of the team (M = 2.88), all pairwise differences significant (F(2, 156) = 26.01, p < .001). Facilitator stance mattered, and learners who rejected the agent's responses still used them as a socially permissible way to speak up — disagreement with the AI did team work. ### Competence, identity, and the evidence gap - **Displacement is the design variable.** [[sun-llm-nursing-education-professional-identity-2026|Sun et al. (2026)]]'s critical integrative review of 489 studies across 47 countries reframes LLM research in nursing around one criterion: whether the cognitive and participatory work an LLM displaces is *extraneous to* or *constitutive of* the competence being developed. Displacing documentation and routine retrieval frees [[cognitive-psychology|working memory]]; displacing a full reasoning chain, an ethical justification, or an individualized care plan removes the object of learning. The evidence is not hypothetical — students using ChatGPT as a sole resource scored significantly below textbook controls on ethical standards and clinical reasoning, and an AI-integrated curriculum produced higher scores alongside less individualized, weaker-logic care plans, a performance–learning dissociation. Their **Professional Identity Tension Model** formalizes this across task, competency, and identity layers: outcomes measured with a tool available may index fluent performance rather than durable capability, so programs should use delayed, no-tool post-tests and transfer tasks. - **The evidence is inverted.** The same review's gap map finds no randomized or quasi-experimental study of professional identity or [[career-development-and-readiness|career development]], only three of relational and ethical competency, and no follow-up beyond 12 months — while controlled evidence clusters in cognitive and technical outcomes. - **Surveillance is documented as perception, not as improved conduct.** [[harerimana-remote-proctoring-nursing-scoping-2026|Harerimana et al. (2026)]] mapped [[remote-proctoring]] in nursing assessment across six studies (1,567 students) and found deterrence reported almost universally by proctored students (98–100% agreement in one graduate program), while the single comparative performance study found an in-person proctored cohort scoring significantly higher on a readiness exam than a remotely proctored one. Anxiety ran both ways — some students calmer at home and relieved of travel, others unable to concentrate while watched and fearful of [[legal-issues-and-risks|wrongful accusation]] — and four of the six studies reported connectivity failure, one load shedding, making [[equity-in-ai-education|equitable]] access the operative constraint rather than exam security. The authors' position is that remote proctoring should be one instrument among several, governed by data-protection standards and paired with integrity-by-design and [[authentic-assessment|authentic assessment]] rather than treated as the default safeguard. - **Boundary conditions.** The strands converge on the same restraint. AI is best evidenced for structured, repeatable objectives — foundational communication, history-taking, health education, knowledge acquisition — and least evidenced for complex psychomotor skill, emotionally loaded interaction, and long-term professional formation. Acceptance is consistently moderate to high, but acceptance of AI *feedback* is conditioned by trust: learners prefer human sources perceived as benevolent over AI perceived as merely competent. The reviews are also geographically narrow, dominated by uncontrolled designs and [[self-report-measures|self-report]], and — in the simulation review — show a near-uniform positive pattern that raises publication-bias concerns. ## Connected Concepts - [[medical-education]] - [[simulation]] - [[affective-computing]] - [[self-efficacy]] - [[higher-ed]] - [[learner-identity]] - [[professional-training]] - [[equity-in-ai-education]] - [[cognitive-offloading]] - [[technology-acceptance-model]] - [[anxiety-and-stress]] - [[ai-literacy]] ## Connected Articles - [[alrazeeni-transforming-nursing-education-ai-2026]] — Systematic review of AI in nursing education (2010–2025) - [[jiang-ai-powered-simulation-nursing-education-2026]] — AI-powered simulation in nursing education: mixed methods systematic review - [[sun-llm-nursing-education-professional-identity-2026]] — LLMs, structural gaps, and the reconstitution of professional identity - [[akbaba-nursing-ai-experiences-tam-2026]] — Nursing students' and faculty AI experiences through the TAM - [[dag-ai-perceptions-career-anxiety-health-2026]] — AI perceptions and career anxiety among health sciences students - [[genai-counter-learner-groupthink-2025]] — Generative AI to counter groupthink in interprofessional PBL - [[dabkowski-nursing-academics-genai-2026]] — Nursing academics on GenAI: policy ambiguity, professional values, and the drift to invigilated assessment - [[harerimana-remote-proctoring-nursing-scoping-2026]] — Remote proctoring in nursing assessment: surveillance, equity, and a deterrence case built on perception --- ## [Legal Education](https://edtechdev.github.io/aied/concepts/legal-education/) > **Legal Education** — the professional preparation of lawyers, and the discipline in which [[generative-ai|generative AI]] raises the strongest version of the question every program faces: whether assisted performance on law-school tasks is evidence of the analysis the license depends on. Its structure is unusual. In the United States, [[governance]] of the [[curriculum-design|curriculum]] runs through ABA accreditation and a licensure [[summative-assessment|examination]] rather than a ministry syllabus; teaching leans on the case method and [[socratic-method|Socratic]] dialogue, which depend on students arriving having done the preparatory reading themselves; clinics and legal writing courses carry the [[experiential-learning|experiential]] weight; and the professional conduct rules students will be bound by already apply to the tools they are being taught to use. AI enters through legal research platforms, drafting support, hypothetical generation, and bar preparation, and its signature failure is not a wrong grade but a fabricated citation. ## Questions to Consider - If a student briefs cases with an AI assistant but cannot reproduce the reasoning unaided, what has the [[socratic-method|Socratic]] classroom been assessing all along — the preparation or the tool that produced it? - Law schools are being asked to govern AI while the profession they feed is still deciding what lawful AI-assisted practice looks like. Should legal education follow the bar, or lead it? - Legal education has more policy documents than empirical studies on this question. Is a discipline that teaches evidence law unusually well placed to demand better evidence about its own AI adoption? ## Introduction Legal education prepares students for a licensed profession with its own rules of conduct, its own accreditation body, and its own gatekeeping examination, which makes it the discipline in the knowledge base where [[educational-policy-ai|education policy]], professional [[regulation]] and assessment validity intersect most tightly. It sits alongside [[medical-education|Medical and Health Professions Education]] and [[nursing-education|Nursing Education]] as a [[professional-training]] discipline rather than a school subject, but its distinguishing feature is procedural: much of what law schools must decide is not *how* to teach with AI but *what the professional ethics rules require* of a graduate who will practice with it. The evidence base is currently thin and policy-heavy. One substantial article anchors the page: [[gutowski-hurley-genai-policy-legal-education-2025|Gutowski and Hurley (2025)]] surveyed and compared institutional generative AI policies across ABA-approved US law schools and proposed a governance framework. That imbalance is itself a finding, and the page flags it rather than papering over it. ## How AI appears in legal education - **Legal research.** Commercial platforms such as Lexis+ AI and Westlaw Precision with CoCounsel embed generative features in the tools students are already trained on, which makes "AI use" hard to separate from ordinary database searching. Gutowski and Hurley note how much of the familiar workflow this compresses: identifying authorities and secondary sources in minutes instead of hours. - **Drafting and writing support.** Initial case briefs, outlines, memoranda, first-pass syntheses and sentence-level feedback on clarity and grammar. The authors place AI-generated work in the same supervisory relationship as work by a paralegal or junior associate: a lawyer remains responsible for its accuracy and legal sufficiency. - **Study and bar preparation.** Generating practice questions, fact patterns, and "hypotheticals" for timed practice, and tutoring on recurring weaknesses. Gutowski and Hurley pair this with the observation that generative AI now passes both the Bar Exam and the Multistate Professional Responsibility Exam, which they read as saying more about the minimal-competency bar than about the model. - **Journals, moot court and advising.** Screening submissions, preparing advocacy, and academic support programs using custom models trained on past exams, model answers and course materials. - **Assessment and integrity.** The recurring problem cases: undisclosed drafting, citation to non-existent authority, and exams that no longer measure unaided analysis. ## What the policy evidence shows Gutowski and Hurley assess policies on five dimensions with 0–5 rubrics: **prohibitiveness**, **permissiveness**, **educational integration**, **transparency and accountability**, and **depth**. Their canvass found that most law schools adopted generally prohibitive stances, often explained as buying time to study the technology, while almost all policies reserved discretion to individual instructors. The permissiveness and prohibitiveness ratings correlated inversely, as expected, and the authors report no single accepted approach: policies range from comprehensive governance to no stated policy at all. They also document how thin the sector's evidence is. The ABA's 2024 AI and legal education survey drew responses from only 29 schools, roughly 15% of ABA-approved law schools, so it can hardly be treated as representative; and a LexisNexis survey of 800 law students found only 9% reporting current use of generative AI for their studies, with 25% planning to adopt it. The authors read low reported use partly as a response to unclear rules, since clarity reduces the fear of breaking them. Two further points shape their recommendations: - **Faculty governance is a law-school-specific complication.** Law faculties, individually and collectively, hold unusually direct authority over curriculum and course standards, so policy has to be built with them rather than announced to them. - **Clinical education raises a [[pedagogy|pedagogical]] objection, not a technical one.** Citing Karr and Schultz, they report the position that AI tools designed to mimic human responses do not develop the original judgment, client interaction and [[ethics|ethical]] decision-making that clinics exist to produce, and therefore should not be used in clinical courses at all. The recommendations are procedural: clear and comprehensive guidelines whatever the institutional stance, full stakeholder involvement in drafting, proactive training for students and faculty, flexible governance reviewed periodically, and self-regulation through information sharing rather than waiting for the ABA to dictate policy. ## Why the discipline is distinctive Three things make legal education more than one more subject area. **Assessment validity is inseparable from professional competence.** [[cognitive-offloading|Over-reliance]] during law school is feared to leave graduates underprepared for work that demands independent legal analysis and advocacy, and grading that rewards fluent output cannot distinguish a student's comprehension from a model's. The [[assessment-validity|validity]] question is therefore not academic. **Disclosure norms are unsettled and consequential.** Gutowski and Hurley report disagreement about whether any use of AI must be disclosed, proposals for student disclosure forms, and no consensus on citation practice, with the Bluebook still silent on citing generative AI. They also predict that disclosure requirements will eventually look as pointless as noting that a student used a search engine, which is a prediction about the shelf life of the rules now being written. **Professional ethics travel with the graduate.** The ABA's duty of technological competence and its guidance on confidentiality, supervision and candour toward the tribunal apply to practicing lawyers using these tools, and law schools inherit the job of teaching them. That is why the discipline connects directly to the knowledge base's work on [[ai-use-disclosure|AI use disclosure]] and on the [[legal-issues-and-risks|legal issues and risks]] that arise when institutions govern AI badly. ## Open Questions - Does AI-assisted legal writing coursework still develop the analysis that licensure assumes, or does it train a fluency the bar exam cannot detect either? - What happens to Socratic teaching if the preparatory reading is routinely delegated, and is the classroom's silence about it a measurement problem or a design one? - Legal education produces the people who will litigate and regulate [[ai-education|AI in education]]. Does that make it the natural place for the knowledge base's legal-risk work to be tested, or does it distort attention toward the law schools with the loudest policies? ## Connected Concepts - [[academic-integrity]] — the conduct framework applied to assisted work - [[assessment-validity]] — whether assisted performance measures the intended competence - [[authentic-assessment]] — designs that require unaided analysis - [[ai-use-disclosure]] — attribution and reporting norms still without consensus - [[governance]] — faculty governance and policy creation as law schools practice it - [[educational-policy-ai]] — institutional AI policy as an object of study - [[critical-thinking]] — the reasoning the case method is meant to build - [[socratic-method]] — dialogue dependent on student preparation - [[experiential-learning]] — clinics and practice-based training - [[professional-training]] — the wider family of licensed professions - [[career-development-and-readiness]] — practice readiness as the stated purpose - [[ai-literacy]] — technological competence as a professional duty - [[hallucination-risk]] — fabricated citations as the field's signature failure - [[legal-issues-and-risks]] — what institutions face when governance fails ## Connected Articles - [[gutowski-hurley-genai-policy-legal-education-2025]] — Five-factor comparison of law school generative AI policies and a governance framework - [[llm-turing-test-italian-legal-exams-2026]] — Blind Turing test on Italian bar, judges and notary examinations, benchmarked against expert marking --- ## [Information Technology Education](https://edtechdev.github.io/aied/concepts/information-technology/) > **Information Technology Education** — the branch of computing education that prepares practitioners to select, deploy, secure, administer, and govern the socio-technical systems organizations actually run, rather than to study computation as a discipline in its own right. Its closest neighbor, and the source of most boundary confusion, is [[cs-education]]: computer science education centers on algorithms, programming, and formal foundations, whereas IT education centers on applied configuration, cybersecurity, data and information management, and the organizational [[governance]] of those systems. [[generative-ai|Generative AI]] reaches the field twice over — as an object students must learn to evaluate, secure, and regulate, and as an instrument that tutors them, classifies their questions, and quietly rewrites the [[professional-training|professional]] pathways they are being prepared for. ## Questions to Consider - If AI compresses the build-fail-debug cycles that historically produced IT expertise, which foundational competencies should a curriculum deliberately protect, and which can be delegated to the tool? - Students know their institution's rules and still cannot say whether their own use complies. Is that a communication failure, or is a rule-based instrument the wrong lever for a practice that happens privately on personal accounts? - IT education sits between computer science, business, and vocational education. Who should own AI governance competence — a course, a curriculum thread, or a program-level outcome? - If a transformer classifier separates higher-order learner questions at roughly 79% precision, should IT programs automate formative feedback on questioning at all? - Most student mental models of GenAI are declarative and shallow. Does that predict [[ai-misuse-learning-harm|misuse]], or only an inability to explain decisions students are in fact making correctly? ## Introduction Information technology education is the applied wing of computing. It trains people to make systems work inside organizations — networks, databases, security operations, health information management, e-government services — and its graduates are usually assessed by professional bodies and employers rather than only by the academy. The distinctiveness relative to [[cs-education]] is the unit of analysis. CS education takes the program and the algorithm as its objects; IT education takes the deployed system and the practitioner's judgment about it. Where CS debates whether AI code generation erodes programming skill, IT debates whether AI troubleshooting erodes diagnostic skill, whether AI-drafted policy documents bind anyone, and whether graduates can govern the data systems they administer. The neighboring disciplines overlap in different ways. [[stem-education]] is the parent category and shares its instruments, but IT education is more often a professional master's or an applied undergraduate degree whose graduates enter regulated workplaces. [[business-education]] is adjacent through information systems and e-government programs, which frequently share courses and students with IT curricula. [[vocational-education]] is the other occupational neighbor: both fields train for practice, but vocational programs target technician-level competence in defined trades, while IT education assumes abstract systems reasoning and produces the people who write the governance documents as well as follow them. [[higher-ed]] names the level rather than the field, and the field carries an accreditation and compliance surface — health informatics accreditation, HIPAA and FERPA obligations for the data students handle — that shapes what counts as legitimate curriculum. What is distinctive about AI in this field is that the same technology is both the curriculum content and the pedagogy. The bundled evidence converges on one uncomfortable theme: IT education's current AI agenda is dominated by integrity and tool adoption, while the governance, security, and data-ethics competencies the field's own workplaces demand appear mostly outside the documents students actually receive. ### How AI appears in Information Technology Education - **Gamified cybersecurity training.** Li and colleagues built several short, mobile-friendly games spanning password security through text and phone scam recognition, combining quiz-based, narrative-based, and [[simulation]]-based designs with interactive formats such as TikTok Mini-Games, motivated by the low engagement and limited effectiveness of conventional video training ([[ai-gamification-security-education-2026]]). Their two-tier evaluation with 59 college students (9 technical experts, 50 general users) reports potential to improve engagement and attention to cybersecurity rather than demonstrated learning gains. - **GenAI as scaffold or shortcut in self-regulated learning.** A mixed-methods study of 267 postgraduate IT students in Australia distinguishes scaffolded [[cognitive-offloading]] (learners clarify goals, generate ideas, obtain [[feedback]] they then critique and adapt, so [[agency]] stays with them) from substitutional offloading (outputs accepted with minimal verification, control shifting to the tool) ([[atif-dickson-deane-scaffold-shortcut-genai-srl-2026]]). Confidence shaped orientation: confident students exercised autonomy in goal setting and monitoring, while less confident peers read GenAI as a shortcut or as [[academic-integrity|misconduct]]. The cohort was differentiated, not homogeneously AI-literate; the authors recommend requiring students to justify or adapt AI outputs as an explicit [[learning-design]] move. - **Compressed expertise pathways in professional practice.** Fourteen semi-structured interviews with IT professionals found GenAI acting as both a mentor-like tutor and a ladder-shortening tool in troubleshooting, scripting, and system verification ([[genai-expertise-pathways-sysadmin]]). Accelerated unfamiliar-domain performance reduces exposure to the build-fail-debug cycles that historically built expertise; AI-assisted speed also resets team and self-expectations, producing a two-speed culture and productivity guilt. The study carries classroom concerns about [[metacognition|metacognitive]] cost and skill decay into [[professional-training]] and [[lifelong-learning|workplace learning]]. ### Assessment and the learner side of IT AI use - **Learner questions as diagnostic signals.** Lee, Atif, and Kang classified 434 authentic student queries from 12 IT courses into three constructivist instructional roles — knowledge transmitter, facilitator, and co-learner — reaching consensus labeling with [[human-in-the-loop-ai|human-in-the-loop]] review (Fleiss' kappa 0.60 rising to 0.83) and augmenting the corpus to 582 balanced questions ([[lee-learner-question-types-ai-education-2026]]). DeBERTa led at 86.36% accuracy and 96.67% precision on factual questions, but facilitator precision fell to 78.79%, and fine-tuned BERT reached 92.00% recall on co-learner items at only 74.19% precision. Errors came from conceptual similarity between roles, ambiguous learner intent, and domain phrasing misread as cognitive depth; the authors warn that augmentation may have introduced lexical shortcuts, and the 11-student, IT-only corpus cannot yet generalize to other fields. - **Wide but shallow mental models.** From 64 usable concept maps drawn from 86 undergraduates in a required technology-ethics course, five mental-model categories emerged: technical-process based, educational-tool based, transitional, consequence-aware, and integrated ([[student-mental-models-genai]]). Every map showed declarative knowledge, 25 showed procedural, 17 conditional, and only 9 integrated all three. Technical and social-regulatory clusters sat far apart, which the authors read as [[ai-literacy]] and [[ethics|ethical]] awareness developing separately; integrity-centered guidelines, they argue, address only one dimension of how students conceptualize the tool. - **Rules known, compliance uncertain.** A survey of 151 undergraduates in Business Information Systems and E-Government programs found most students actively using GenAI but over half unsure whether their usage complied with institutional regulation, with only weak to moderate associations between [[regulation|regulatory]] awareness and actual behavior, and reliance mostly on privately accessed tools rather than institutional ones ([[student-regulatory-awareness-genai]]). Knowing the rules did not strongly predict what students did. ### Institutional policy and the governance gap - **Guidance, not policy.** An environmental scan of all 48 accredited health informatics and health information management master's programs found 40 (83%) with at least one publicly available AI document, but the modal artifact was advisory guidance (21, 53%) rather than formal policy (7, 18%) ([[institutional-ai-policy-health-informatics-2026]]). Academic integrity dominated the vocabulary (n = 139), ahead of citation (n = 59) and [[assessment]] (n = 50), while HIPAA (n = 5), FERPA (n = 11), equitable access (n = 2), and disclosure requirements (n = 1) were nearly absent, and electronic health records were not mentioned at all. Latent Dirichlet Allocation produced four themes around integrity, student GenAI use, university research tools, and ChatGPT engagement. Because only public documents were analyzed, the authors treat the absence of privacy and equity language as a finding about published guidance, not about institutional practice. - **The equity blind spot.** Inclusion (n = 9), accessibility (n = 9), accommodations (n = 4), and equitable access (n = 2) appear at rates that cannot support any claim that [[equity-in-ai-education]] has been addressed, even though several programs require AI use in coursework. Together with the thin treatment of professional [[governance]] competence, this is the clearest design task the bundle leaves open: connecting integrity rules to the access, privacy, and data-governance provisions those graduates will be responsible for enforcing. ## Connected Concepts - [[cs-education]] - [[stem-education]] - [[business-education]] - [[vocational-education]] - [[higher-ed]] - [[professional-training]] - [[ai-literacy]] - [[academic-integrity]] - [[governance]] - [[cognitive-offloading]] - [[equity-in-ai-education]] ## Connected Articles - [[ai-gamification-security-education-2026]] - [[atif-dickson-deane-scaffold-shortcut-genai-srl-2026]] - [[genai-expertise-pathways-sysadmin]] - [[institutional-ai-policy-health-informatics-2026]] - [[lee-learner-question-types-ai-education-2026]] - [[student-mental-models-genai]] - [[student-regulatory-awareness-genai]] --- ## [Learning Sciences](https://edtechdev.github.io/aied/concepts/learning-sciences/) > **Learning sciences** — the interdisciplinary research field that studies how people learn and how to design environments in which learning happens, drawing on [[cognitive-psychology|cognitive psychology]], [[learning-theories|learning theory]], computer science and linguistics, and judging its designs with empirical evidence rather than theory alone. In this knowledge base it is the research field around [[ai-education|AI in education]] rather than one of the school subjects: it supplies the mechanisms that AI systems operationalize (knowledge components, [[mastery-learning|mastery thresholds]], [[transfer-of-learning|transfer]]), the design objects they are embedded in ([[learning-design|planned course sequences]], [[intelligent-tutoring|tutors]], [[feedback]] regimes) and the standards by which they are judged ([[learning-gains]], [[assessment-validity]], [[equity-in-ai-education|equity]]). Its organizing question is not whether a tool performs well but whether a learner changed. ## Questions to Consider - A learner passes every practice item, so the mastery threshold ends the set — then misapplies the rule where the action should be withheld. Whose error is that: the learner's, the model's, or the stopping rule's? - Sequence mining can describe 554 courses as patterns without observing a classroom. What does the field gain, and lose, by studying designed intentions instead of enacted activity? - Demographic sensitivity in an LLM's feedback looks like adaptation when it tracks a learner's stated education level and like bias when it shifts sentiment. Should a field that cannot separate the two keep using open-ended models to assess? - Does the learning sciences' appetite for causal design — randomized assignment, counterfactual audits, executable models of the learner — narrow what counts as evidence in AI in education? ## Introduction The learning sciences study learning and the design of learning environments, and they are defined by their methods as much as their topics: experiments, classroom trials, [[quantitative-research|quantitative]] modeling of student data, [[qualitative-research|qualitative]] analysis of designs and contexts, and design-based research that builds an intervention and revises it in use. That breadth separates this page from the neighbours that supply the frameworks, the practice and the instruments; the section below sets out each boundary and what the field has established on the other side of it. This page covers the substantive knowledge such methods have produced — what learners do with a generative model, which arrangements change outcomes, and where the field's own instruments fail. [[discipline-specific-aied]] takes the opposite cut, holding that subject matter changes what support should do; the learning sciences take the cross-cutting view, and the mechanisms they test are gathered under [[cognitive-psychology]]. ### How AI appears in the learning sciences - **Mechanism first, then the model.** [[deceptive-overgeneralization-adaptive-learning-2026|An, McLaren and Stamper (2026)]] ran eleven experiments (N = 192) with [[intelligent-tutoring]] systems for Riichi Mahjong and showed that learners who compiled an overgeneralised production — the action without its application constraint — misapplied it on the first "do-not-act" item at 61.5%–100%, against 12% expected error under [[knowledge-tracing|Bayesian knowledge tracing]]. With a 95% mastery threshold the system stopped practice before learners met a case requiring the action to be withheld, so the defect went undetected. Short do-not-act practice with [[feedback]] naming the missing constraint cut misapplication to 0.0%–23.1% (Cohen's h 1.70–2.44), and a secondary analysis of thirteen K-12 *Decimal Point* datasets found the same structure in whole-number bias (84%–88% of comparison errors). - **The optimum depends on the content.** [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit, Koedinger and Carvalho (2025)]] treat the example–problem ratio as a content–treatment interaction: in a 2×2 experiment with 95 participants on geometry-area material, practice-only training produced larger [[learning-gains|learning gains]] for verbatim facts while example-integrated training produced larger gains for generalizable skills (β = 0.41, p = .038, d = 0.38). A simulated learner (Apprentice Learner) reproduced the crossover only when given an ACT-R-style [[cognitive-psychology|memory-and-forgetting mechanism]]. More practice is not uniformly better: memory-oriented content warrants retrieval, and induction-oriented skills warrant integrated examples. - **Design as an analysable object.** [[learning-paths-patterns-learning-design-2026|Divjak, Svetec and Horvat (2026)]] turned [[learning-analytics]] on [[learning-design]] itself, coding 29,064 activities across 554 courses planned in a free course-design tool. Acquisition was the most common learning type and the most common entry point; the strongest Markov transition was Assessment → Discussion (0.332) and the highest-confidence rule was Acquisition → Assessment → Practice → Practice (0.743, lift 1.45). Learning type tracked intended outcome level, Acquisition falling from about 50% of activities at Bloom level 1 to around 20% at level 6. The authors stress these are pre-implementation designs: resemblance to flipped, [[inquiry-based-learning|inquiry-based]] or [[project-based-learning|project-based]] sequences is not evidence of intent. - **Auditing the models that assess.** [[demographic-signals-llm-student-assessment-2026|Rooein, Benedetto and Hovy (2026)]] audited six [[llm|LLMs]] across [[automated-essay-scoring|essay scoring]], [[formative-assessment|formative feedback]] and question answering, holding the task input fixed while varying only demographic context (192,480 calls). Scoring was stable under explicit personas, but Llama-70B inflated its own scores by 1.57 points under implicit conversational history (p < 0.001), and higher education drew less readable and more positive responses — a sentiment gap of roughly four standard deviations. Readability effects shrank while length effects grew, and some coefficients changed sign between conditions. The authors offer it as an audit instrument, not a deployment verdict, and read the entanglement of demographic and topical signal as a threat to [[assessment-validity|validity]] and [[equity-in-ai-education|equity]]. - **Measuring competence, and its limits.** [[competent-generative-ai-use-measures-review-2026|Verí (2026)]] organizes instruments for competent [[generative-ai]] use into four domains — knowledge and use, epistemic oversight, reliance calibration, and control of tool-using agents — refusing to collapse them into one proficiency continuum. Three same-sample correlations between self-rated and demonstrated [[ai-literacy|AI literacy]] pooled to r = .055 (95% CI [-.047, .156], reported N = 2,765), which the author reads as enough to reject treating [[self-report-measures|self-ratings]] as interchangeable with performance scores, though not enough to set a cutoff. No validated instrument covered the full set of decisions that tool-using agents create; the proposed layered battery is a design hypothesis. - **Self-report about the learner's own offloading.** [[pause-ai-cognitive-offloading-self-reflection-2026|Alam (2026)]] translates the [[cognitive-offloading]] literature into PAUSE, a browser-only self-check with four domains, every LLM-era item anchored to a source, no composite, no storage and no model in production; its bands are descriptive rather than normed, and a reading must not justify assessment, admissions or hiring decisions. Its stated limits matter: self-report of offloading is vulnerable to the faculty it concerns, a respondent who deliberately uses AI as a [[scaffolding|scaffold]] reads as offloading on several items, and whether AI-associated offloading is distinct from general technology dependence remains open. - **Where human expertise sits.** [[wang-tutor-copilot-human-ai-live-tutoring-rct-2024|Wang et al. (2024)]] report the clearest division of labor: in a two-month [[rct|randomized controlled trial]] with about 900 novice K-12 tutors and roughly 1,800 students, real-time suggestions drawn from experienced tutors' reasoning raised topic mastery by 4 percentage points (62% to 66%, p < 0.01), and by 9 points for lower-rated tutors, at about \$20 per tutor per year, shifting tutoring toward guiding questions. Gains were proximal — year-end tests did not move. [[reichert-human-centered-llm-chatbot-design-teachers-2026|Reichert et al. (2026)]] find teachers arriving at the same position by design: six secondary teachers prototyping chatbots specified a bounded expert, holding authority boundaries (responsibility for learning and safety is not delegable) and expertise boundaries (the model lacks their knowledge of individual students), and delegating content presentation, practice and corrective [[feedback]] while reserving objective-setting and [[summative-assessment|summative assessment]]. - **Capability at the level of the field.** [[sutedjo-faculty-genai-tpack-21-2026|Sutedjo, Chowdhury and Liu (2026)]] surveyed 127 faculty with a [[tpack|TPACK]] instrument adapted for generative AI: strong content and pedagogical content knowledge (M = 4.70–5.15) beside markedly lower technology-integrated knowledge, with holistic TPACK lowest at 2.55, content knowledge uncorrelated with any technology-integrated domain, and the three integrated domains correlating so highly (r = .81–.91) that they may function as one factor. [[perrotta-zero-shot-governance-2026|Perrotta (2026)]] reads the governance layer through a discontinued UK civil-service prototype whose codebase was a system prompt plus a retrieval pipeline over commercial models, arguing that the generality of foundation models both enables rapid repurposing into [[educational-policy-ai|policy]] tools and makes aberrant output a permanently only-mitigable risk — oversight that peers over the loop rather than sitting inside it. ## How the learning sciences relate to their neighbours [[design-based-research|Design-based research]] is the method this field developed rather than borrowed: an intervention is built and revised inside a working classroom, with its theoretical rationale revised alongside it, so one study yields both an artifact and a design principle. That is what separates it from a laboratory experiment, which isolates a cause by holding the context still, and it is why the field's findings arrive as design knowledge rather than as effect sizes. [[research-methods-aied]] takes the other cut: that page surveys the whole repertoire — experiments, surveys, qualitative work, benchmarks, reviews, consensus methods — as a choice among instruments, weighed for the validity of the claim each can support. This page reads the same corpus from the substantive side, asking what the repertoire has established about learning and judging a method by whether its design claim survives contact with learners. [[learning-theories]] collects the candidate frameworks — behaviorism, cognitivism, constructivism, sociocultural accounts, motivation and self-regulation — as lenses for reading AI. The learning sciences share that vocabulary but not that stance: here a theory is a claim about mechanism that a design must either instantiate or refute, and the field's standing rests on empirical and design work rather than on the coherence of a framework. The theory page is the one to open for what a framework asserts; this page is the one for the evidence a framework has accumulated. The field also builds theory rather than only testing borrowed frameworks: [[theory-development-aied]] covers the conceptual work that explains how learners, teachers and AI systems interact, and it is where the field's own constructs are argued before they are measured. What its designs are usually asked to produce is [[transfer-of-learning|transfer]] — knowledge and skill that survive past the tutor, subject or task they were learned in — which is why a gain measured inside a tool counts as a weaker claim than one measured without it. And because a designed environment is a compound intervention, attributing an outcome to one component is the field's standing measurement problem: [[educational-measurement]] supplies the psychometric apparatus that makes the attribution arguable at all, which is why measurement questions arrive early here rather than after the fact. [[cognitive-psychology]] is the mechanism-level discipline the field draws on most heavily, supplying bounded working memory, encoding and retrieval, decomposable knowledge components and the diagnostic language of learner modeling. The learning sciences use those mechanisms without reducing to them: their unit of analysis is a designed environment carrying social, motivational and contextual variables that a laboratory account of memory does not, and their tests are run on whole interventions rather than on isolated cognitive effects. [[pedagogy]] and [[learning-design]] cover practice — which teaching strategy to use, and how to sequence objectives, activities and assessment into a course. Both are what the learning sciences study from the outside, as objects of description and evaluation; the field does not tell a teacher which tactic to reach for next, it reports what the tactics have been shown to do. Learning design is the closer relative, since both produce something that can be implemented and tested, but the designer's output is a teachable course and the field's output is knowledge about designs in general. The findings only matter once they reach teaching, and that journey runs through three pages. [[educational-development]] is the institutional practice that carries them — faculty development, standards, policy and identity work decide whether a validated design ever reaches a classroom, which is why the field's evidence routinely leads what institutions have implemented. [[teacher-education]] is where the knowledge has to land before a teacher enters the room, and [[teacher-role]] is where it lands afterwards, in the moment-to-moment judgment about when to intervene, which instrument to use, and when to leave a learner alone. None of the three produces learning-science findings; all three decide whether those findings change practice. ## Connected Concepts - [[learning-theories]] - [[cognitive-psychology]] - [[pedagogy]] - [[learning-design]] - [[research-methods-aied]] - [[theory-development-aied]] - [[design-based-research]] - [[teacher-education]] - [[educational-development]] - [[discipline-specific-aied]] - [[intelligent-tutoring]] - [[learning-analytics]] - [[assessment-validity]] - [[educational-measurement]] - [[cognitive-offloading]] - [[learning-gains]] - [[transfer-of-learning]] - [[teacher-role]] - [[equity-in-ai-education]] - [[educational-policy-ai]] ## Connected Articles - [[competent-generative-ai-use-measures-review-2026]] — Review and exploratory meta-analysis of measures for competent generative-AI use (Verí 2026) - [[deceptive-overgeneralization-adaptive-learning-2026]] — Correctness masking an incomplete rule: mastery stopping rules in adaptive learning (An, McLaren & Stamper 2026) - [[demographic-signals-llm-student-assessment-2026]] — Counterfactual audit of demographic signals in LLM student assessment (Rooein, Benedetto & Hovy 2026) - [[learning-paths-patterns-learning-design-2026]] — Markov chains and pattern mining over 29,064 activities in 554 courses (Divjak, Svetec & Horvat 2026) - [[pause-ai-cognitive-offloading-self-reflection-2026]] — A privacy-preserving, non-diagnostic self-check for AI-associated offloading (Alam 2026) - [[perrotta-zero-shot-governance-2026]] — Zero-shot governance: general-purpose AI in policy, read through the Redbox codebase (Perrotta 2026) - [[rachatasumrit-example-problem-ratio-2026]] — Why the best example–problem ratio depends on content (Rachatasumrit, Koedinger & Carvalho 2025) - [[reichert-human-centered-llm-chatbot-design-teachers-2026]] — Teachers design bounded-expert chatbots, with selective delegation of instruction (Reichert et al. 2026) - [[sutedjo-faculty-genai-tpack-21-2026]] — Faculty GenAI TPACK: strong content knowledge, weak technology-integrated knowledge (Sutedjo, Chowdhury & Liu 2026) - [[wang-tutor-copilot-human-ai-live-tutoring-rct-2024]] — Tutor CoPilot: a randomized trial of human–AI live tutoring at scale (Wang et al. 2024) --- ## [K-12](https://edtechdev.github.io/aied/concepts/k-12/) > **K-12** — the use of artificial intelligence in primary and secondary education, spanning [[ai-literacy|AI literacy]] curricula, [[intelligent-tutoring|AI tutoring]], [[teacher-role|teacher support]], and safety considerations unique to younger learners. A 2026 PRISMA-guided systematic review of 197 studies (2016–2024) organizes the field around four opportunity domains — [[personalized-learning|personalization]], [[student-engagement|motivation]], [[assessment]], and innovative [[pedagogy|teaching practices]] — while flagging persistent gaps in [[teacher-education|teacher training]], ministerial [[ethics|ethical]] guidelines, and discipline balance. ## Questions to Consider - K-12 AI use demands stronger safety protections than [[higher-ed|higher education]] because younger learners are least equipped to detect unsafe or manipulative AI. How does that change what 'safe' must mean for tools used by children? - [[meta-analysis-systematic-review|Systematic reviews]] find the K-12 evidence base concentrates overwhelmingly on [[stem-education|STEM]], leaving the arts and [[humanities-education|humanities]] largely unstudied. What might a discipline-balanced K-12 AI evidence base reveal that STEM-only studies miss? - [[research-methods-aied|Research]] finds teachers systematically overestimate their AI competency — a roughly 40% gap between [[self-report-measures|self-report]] and performance — yet short training yields substantially higher integration. What might that gap mean for how schools prepare teachers? - Most AI tools center dominant perspectives, yet 78% of teachers in one study found AI helpful for culturally relevant pedagogy. How can a tool that defaults to dominant views be steered to broaden representation rather than narrow it? - Evidence suggests outsourcing AI work can harm learning, especially for older students — yet younger students may be less affected. Why might the learning penalty of AI use differ by age, and what does that imply for design? - The K-12 review identifies eight under-explored research gaps, from the missing concrete teaching-unit examples to the underuse of [[educational-robotics|robotics]], wearables, and mobile communication. Which of these gaps most limits what schools can actually do today? - Institutional [[educational-policy-ai|AI policies]] often lack implementation guidance. What would it take to turn a K-12 AI policy document into concrete classroom practices and teacher training — and why might that translation so often fail? ## Introduction ### The evidence landscape A 2026 [[meta-analysis-systematic-review|PRISMA-guided systematic review]] (Marzano, [[generative-ai-k12-teaching-learning-systematic-review-2026|197 studies, 2016–2024]]) now anchors the K-12 evidence base, complementing earlier large syntheses such as the [[young-people-learning-generative-ai-rapid-review-2026|rapid review of GenAI and PreK-12 learners]] and the [[stanford-evidence-base-ai-k12-2026|aggregated K-12 AI evidence base]]. Its four-question structure organizes the field into four opportunity domains — personalizing learning, motivating students, improving assessment, and enabling innovative, immersive teaching — with [[conversational-ai|ChatGPT]] as the flagship case studied across a large share of the included articles. The review's central finding is a sharp asymmetry: the *promise* of [[generative-ai|GAI]] in schools is broad and well documented, but the *evidence of daily classroom practice* is thin, concentrated in [[stem-education|STEM]], and skewed by a shortage of concrete teaching-unit examples and practical teacher training. This gap between aspiration and implementation is the defining feature of K-12 GenAI research. A teacher-side synthesis adds the layer that tool-focused reviews leave out. A PRISMA review of 29 studies of [[teacher-intervention-k12-ai-based-instruction-2026|teacher intervention in K-12 AI-based instruction]] finds that AI output — alerts, [[visualization|dashboards]], automated scores, chatbot feedback — becomes teaching only through a cycle of monitoring, judgment, intervention and orchestration, and that the reported benefits for student performance, participation and confidence are conditional on the interpretability of the information, intervention timing, the level targeted, teachers' implementation feasibility and students' [[agency|autonomy]]. Its corpus is concentrated in the same places as the wider K-12 evidence base: [[math-education|mathematics]] and science/STEM subjects, middle and high school grades, and United States settings. ### Distinctive K-12 considerations - **Safety and guardrailing:** K-12 AI use demands stronger [[pedagogical-safety]] protections. [[eduzone-llm-safety-k12|EduZone]], [[eduguard-safe-rag-llm-tutor|EduGuard]], and [[hazra-safetutors-pedagogical-safety-2026|tutor harm research]] specifically address child-safe [[student-ai-interaction|AI interaction]]. - **AI literacy curricula:** [[aaai2026-prompting-literacy-k12|K-12 AI literacy modules]] develop age-appropriate AI understanding, while the [[gaide-vibe-coding-k12-teachers|vibe coding framework]] empowers teachers to create their own AI tools. A quasi-experimental study of an 18-hour, 5E-model critical media literacy program for fourth-grade Turkish students ([[demir-akar-ai-media-literacy-children-2026|Demir & Akar 2026]]) embedded [[generative-ai|generative AI]] (ChatGPT, Grammarly, Canva AI, Padlet) phase-by-phase as a [[pedagogical-agent|pedagogical agent]] aligned to the Turkish Language and Social Studies curricula, producing large gains in media reading (+3.50), writing (+1.67), and total media literacy (+5.17, all p < .01) with between-group effect sizes of Cohen's *d* = 1.12–1.31, and [[qualitative-research|qualitative]] growth across six domains of critical media literacy (digital self-protection and data privacy, purposeful and responsible media use, safe communication and boundary awareness, [[critical-thinking|critical evaluation]] and misinformation awareness, online risk awareness, and media ethics/digital citizenship) — evidence that age-appropriate, discipline-embedded AI literacy curricula can yield measurable critical-media gains in the elementary grades. - **Tutoring at grade level:** [[ecnuclaw-k12-personalized-companion|K-12 personalized companions]] and [[correct-answer-trap-ai-tutor|correct answer trap research]] examine tutoring effectiveness for younger students. - **Developmental appropriateness:** [[child-safety-genai|Child safety research]] and [[special-education]] considerations address the needs of diverse K-12 populations. - **Age-tailored conversational AI:** [[vahedian-children-attitudes-ai-chatbot-2026|Vahedian Movahed & Martin (2025)]] deployed AMA, a topic-bounded chatbot with age-varied prompt tailoring, with 63 students (grades 1 and 6–8). Children showed broad openness to and high trust in the chatbot as an information source and actively tested its credibility with known-answer questions, while gaps in digital-safety awareness (some were willing to confide secrets) point to age-sensitive [[scaffolding]] and explicit [[privacy]] instruction. ### Teacher-knowledge effects - **Teacher-knowledge effects.** A cross-level study ([[pedagogy-first-technology-second-teacher-knowledge-2026|46 teachers, 2,832 secondary students]]) found [[pedagogy|pedagogical]] AI knowledge — not technical AI knowledge — drove students' perceptions of AI for social good and their intention to learn AI, reinforcing a "pedagogy first, technology second" approach to K-12 AI teacher preparation. - **Continuous training as the binding constraint.** The Marzano review identifies continuous, targeted [[teacher-education|teacher training]] on ICT as the single most persistent integration challenge, and it must cover technical, pedagogical, and ethical dimensions while adapting to rapidly evolving tools. Initial and ongoing training should be continuous, collaborative, [[game-based-learning|game-based]], [[explainable-ai|transparency]]-focused, and grounded in [[educational-development|professional development]] for advanced tools like ChatGPT. This extends to classroom collaboration-support AI: the Community Builder ([[breideband-community-builder-cobi-2026|CoBi]]) feasibility study in middle-school classrooms found that high-integrity teacher use depended on substantial [[professional-training|professional learning]], with under-prepared teachers drifting toward treating the tool as a performance monitor or into general AI-discussion rather than its intended collaborative focus. - **The competency gap.** Research shows teachers systematically overestimate their [[teacher-ai-competency|AI competency]] — a roughly 40% gap between self-report and performance — yet brief training interventions (e.g., 4-hour [[prompt-engineering|prompting]] workshops) can yield substantially higher classroom AI integration. This makes teacher preparation and [[teacher-education]] central to successful K-12 AI integration. ### Four opportunity domains The systematic review consolidates what GAI can do in schools into four recurring domains: - **Personalized learning.** GAI tailors [[personalized-learning|learning experiences]] to individual needs, pace, and readiness — the flagship benefit documented across [[ecnuclaw-k12-personalized-companion|companion systems]] and [[adaptive-learning|adaptive platforms]]. - **Readability and curriculum alignment.** Bird (2026) fuses transformer text classification with computational-linguistics features to classify English literature by UK Key Stage (F1 0.996), packaged into a no-code web app so teachers can align reading materials to learners' levels — a concrete, discipline-relevant (English/literature) complement to the STEM-heavy K-12 evidence base. - **Motivation and engagement.** Tools like ChatGPT make learning more engaging for younger students, supporting [[student-engagement|engagement]] and [[motivation]] — though the review also warns this can tip into excessive reliance and technological dependency. - **Assessment.** GAI improves [[assessment]] methods, from [[formative-assessment|formative]] [[feedback]] to [[automated-assessment|automated scoring]], though [[marked-pedagogies-linguistic-bias-writing-feedback|bias in automated feedback]] remains a live concern. [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] extend this to item calibration in K-5: across 5,170 math and reading items, GPT-4o's zero-shot difficulty ratings correlated moderately-to-strongly with Rasch-calibrated difficulties (r = 0.83 math, r = 0.81 reading) but were uneven across grades and no better than a grade-mean dummy regressor for grades K and 1 — a range-restriction limitation that matters for early-grade assessment. A feature-based approach ([[llm]]-extracted features into tree-based models) reached correlations up to r = 0.87, with gains most pronounced for early-grade items. - **Innovative teaching practices.** Immersive, [[game-based-learning|game-based]], and [[collaborative-learning|collaborative]] approaches make learning more engaging, with [[code-to-learn-genai-artifact-construction-2026|constructionist]] and [[ai-pbl-computational-thinking-2026|project-based]] models showing particular promise. A concrete classroom-wide example is the Community Builder ([[breideband-community-builder-cobi-2026|CoBi]]), an AI collaboration-support system deployed across six middle-school classrooms that used speech recognition to visualize small-group "uplifting" discourse — evidence that real-time speech AI can work feasibly in noisy K-12 environments to support the relational dimension of learning. ### Cultural relevance imperative [[curriculum-design|LLM-supported curriculum design]] shows promise for diversifying materials — in one study 78% of teachers found AI suggestions helpful for [[culturally-relevant-pedagogy|culturally relevant pedagogy]]. However, most AI tools center dominant perspectives, requiring deliberate [[equity-in-ai-education|equity-centered design]] to avoid widening opportunity gaps. ### Policy-to-practice translation [[governance|Institutional]] [[generative-ai|GenAI]] policies largely lack implementation guidance. The systematic review underscores that ministerial guidelines addressing [[ethics]] and [[privacy]] remain underdeveloped, echoing the [[ai-literacy-assessment-misalignment|policy-to-practice]] translation problem: successful models transform policy documents into actionable teacher training modules, bridging the "what" (policy) and "how" (prompting instruction). ### Scale and equity K-12 AI deployment operates at societal scale — millions of students, compulsory education, and significant [[equity-in-ai-education|equity]] implications. Equity research examines whether AI widens or narrows opportunity gaps, including [[digital-divide|access]] and [[inclusive-learning|inclusive support]] for students with disabilities — one of the review's eight under-studied areas. ### Eight research gaps An innovative contribution of the Marzano review is its explicit enumeration of eight under-explored research gaps that shape a K-12 GenAI research agenda: (1) a lack of concrete examples of AI in teaching for constructing teaching units; (2) a need for practical, daily teacher training based on constant AI application; (3) skills in [[reinforcement-learning|machine learning]] and a balance between disciplines, with [[stem-education|STEM]] over-represented; (4) a lack of studies and experiments in Europe; (5) a need to define the [[teacher-role|teacher role]] and create specific pedagogical frameworks such as [[tpack]] or AI4K12; (6) inclusivity and support for students with [[special-education|disabilities]]; (7) a connection with pedagogical theories and innovative methodologies; and (8) applications of emerging [[ai-technologies|technologies]] such as [[educational-robotics|wearables, robot control]], and mobile communication. ### Connections K-12 connects to [[pedagogical-safety]] (child protection), [[ai-literacy]] (student competency), [[teacher-role]] (K-12 teacher transformation), [[equity-in-ai-education]] (access gaps), and [[special-education]] (diverse learner needs). Its systematic-review evidence base links to [[meta-analysis-systematic-review]] as a method, and its discipline-imbalance finding connects to [[stem-education]], [[humanities-education]], and [[early-childhood-elementary-ai-education]]. ## Implications for K-12 instructors - **Prioritize child-safe guardrailing.** K-12 AI use demands stronger [[pedagogical-safety|safety]] protections — use tools specifically built and vetted for children ([[eduzone-llm-safety-k12|EduZone]], [[eduguard-safe-rag-llm-tutor|EduGuard]]) and be alert to the harms of unguarded tutors ([[hazra-safetutors-pedagogical-safety-2026|tutor harm research]]). - **Build AI literacy age-appropriately.** Use curricula and frameworks designed for younger learners ([[aaai2026-prompting-literacy-k12|K-12 AI literacy modules]]), and don't assume students arrive with responsible-use knowledge. - **Close the teacher-competency gap.** Teachers systematically overestimate their AI competency (~40% gap between self-report and performance), yet short training (e.g., 4-hour prompting workshops) yields substantially higher classroom AI integration — invest in your own preparation ([[teacher-ai-competency]], [[teacher-education]]). - **Keep training continuous and practice-grounded.** The systematic review's core recommendation is that teacher training cannot be one-off — it must be continuous, collaborative, and based on constant daily AI application, not abstract theory. - **Diversify materials deliberately.** AI is strong at [[culturally-relevant-pedagogy|culturally relevant]] material generation (78% of teachers find suggestions helpful), but tools default to dominant perspectives — actively use AI to broaden, not narrow, representation, and guard against [[equity-in-ai-education|opportunity gaps]]. - **Watch the learning-penalty risk.** Evidence shows outsourcing AI work can harm learning, especially for older students ([[generative-ai-reduced-study-time-math|cognitive surrender]], [[stromberg-generative-ai-learning-penalty-secondary-2026|homework-outsourcing penalty]]) — design assignments that keep cognitive work in the loop and assess for durable understanding, not just completion. - **Design teacher tools for judgment, not volume.** More AI information is not better: real-time alerts expanded teachers' awareness but in some studies overloaded attention or pulled focus from their own observation, and a single teacher cannot act on more intervention targets than they can physically reach ([[teacher-intervention-k12-ai-based-instruction-2026|Lee 2026]]). Tools that prioritize and explain what is worth acting on, and let teachers accept, revise, defer or reject recommendations, fit the work better than dashboards that surface everything — and the structural conditions (review time, class size, support personnel) are part of the intervention, not background logistics. - **Extend beyond STEM.** The evidence base skews heavily toward [[stem-education|STEM]]; educators and researchers in the arts and [[humanities-education|humanities]] have an outsized opportunity to generate the discipline-balanced evidence the field is missing. - **Translate policy into practice.** Institutional GenAI policies often lack implementation guidance; turn policy into concrete, actionable classroom practices and training ([[ai-literacy-assessment-misalignment|policy-to-practice]]). ### AI Adoption and Skepticism in K-12 - As AI enters schools, the staff it affects are increasingly urged to consult [[conversational-ai|conversational AI]] about adoption. An algorithmic audit of ten frontier LLMs probed a rural K-12 AI-skeptic persona, finding eight of ten models acknowledged concerns before redirecting to AI-[[student-engagement|engagement]] framings — a model-dependent design outcome rather than an inevitable property of LLMs. [[trust|Trust]], skepticism, and [[human-in-the-loop-ai|human oversight]] are therefore central to how rural and other K-12 staff encounter AI. ## Connected Concepts - [[early-childhood-elementary-ai-education]] — Early childhood and elementary AI education - [[pedagogical-safety]] — Child-safe guardrails for AI interactions with younger learners - [[ai-literacy]] — Age-appropriate student understanding and responsible use of AI - [[teacher-role]] — Teacher transformation and AI integration in K-12 classrooms - [[equity-in-ai-education]] — Whether AI widens or narrows K-12 opportunity gaps - [[special-education]] — Diverse learner needs in K-12 AI use - [[scaffolding]] — Support structures for AI-assisted K-12 learning - [[personalized-learning]] — Tailored instruction enabled by AI at the school level - [[stem-education]] — AI literacy and tools across K-12 science, tech, engineering, math - [[generative-ai]] — The core technology behind K-12 AI tools and curricula - [[teacher-education]] — Training and upskilling teachers for classroom AI - [[sociocultural-learning]] — Social and cultural dimensions of K-12 AI learning - [[metacognition]] — Thinking about thinking; key to durable AI-era learning - [[educational-development]] — Ongoing professional development for teachers - [[learning-design]] — Designing lessons and materials with AI support - [[active-learning]] — Keeping cognitive work in the loop when using AI - [[computational-thinking]] — Foundational K-12 skill and AI curriculum target - [[intelligent-tutoring]] — AI tutoring systems for younger students - [[meta-analysis-systematic-review]] — The systematic-review method behind the K-12 evidence base - [[humanities-education]] — Discipline-balance counterpart to STEM-heavy K-12 AI evidence - [[parents-and-families]] - [[cognitive-surrender]] ## Connected Articles - [[generative-ai-k12-teaching-learning-systematic-review-2026]] — Systematic review of generative AI in K-12 teaching and learning (Marzano 2026) - [[pedagogy-first-technology-second-teacher-knowledge-2026]] — Cross-level effects of teacher AI knowledge (TAIK, TPAIK) on student learning in secondary AI education (Shen et al. 2026) - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[school-ai-education-readiness-gaps-agency-2026]] — School AI education narrows psychological but not cognitive readiness gaps - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[ai-pbl-computational-thinking-2026]] — AI-assisted project-based learning and computational thinking - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — Virtual tutoring with CAL: an experiment in take-up and learning - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[eduzone-llm-safety-k12]] — EduZone: child-safe LLM interaction platform for K-12 - [[eduguard-safe-rag-llm-tutor]] — EduGuard: safe RAG-based LLM tutor guardrailing - [[hazra-safetutors-pedagogical-safety-2026]] — Research on harms of unguarded AI tutors for children - [[aaai2026-prompting-literacy-k12]] — K-12 AI literacy modules on prompting - [[stanford-evidence-base-ai-k12-2026]] — The K-12 AI evidence base aggregated at scale - [[ecnuclaw-k12-personalized-companion]] — K-12 personalized AI companions - [[gaide-vibe-coding-k12-teachers]] — Vibe-coding framework empowering teachers to build AI tools - [[elementary-writing-genai-systematic-review-2026]] — Systematic review: GenAI and elementary writing - [[young-people-learning-generative-ai-rapid-review-2026]] — Rapid review: GenAI and PreK-12 learners (271 papers) - [[generative-ai-reduced-study-time-math]] — Cognitive surrender strongest for high schoolers, absent for Grade 5 - [[stromberg-generative-ai-learning-penalty-secondary-2026]] — The generative AI learning penalty: homework outsourcing harms learning - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Marked Pedagogies: bias in automated feedback for middle-school writing - [[code-to-learn-genai-artifact-construction-2026]] — CtL-GenAI: constructionism framework for artifact construction - [[all-girls-genai-makerspace-gender-equity-2026]] — All-girls GenAI makerspace workshops and gender equity in computing - [[frontier-ai-redirect-skeptical-rural-staff-2026]] — Algorithmic audit: how frontier LLMs redirect skeptical rural K-12 staff - [[demir-akar-ai-media-literacy-children-2026]] — AI-based critical media literacy program for children - [[bird-multimodal-educational-literature-2026]] — Multimodal fusion for classifying educational literature - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[breideband-community-builder-cobi-2026]] - [[vahedian-children-attitudes-ai-chatbot-2026]] - [[teacher-intervention-k12-ai-based-instruction-2026]] — Teacher intervention in K-12 AI-based instruction: a systematic review --- ## [Early Childhood Education](https://edtechdev.github.io/aied/concepts/early-childhood-elementary-ai-education/) > **Early childhood education** — the use of artificial intelligence in the education of young children, spanning preschool and the elementary (primary) years. This covers AI-literacy and [[computational-thinking|computational thinking]] curricula for young learners, AI-enabled toys and play, [[personalized-learning|personalized learning]] in elementary subjects, and the developmental, safety, and equity considerations unique to children rather than adolescents or adults. ## Questions to Consider - Young children now meet AI through toys, [[conversational-ai|chatbots]], and classroom robots. What makes a 5-year-old's relationship with AI fundamentally different from an adult's — and what should change about how we think about the risks? - Some early-childhood AI literacy is taught through 'unplugged' play — no computers at all. How can abstract AI concepts be learned through [[embodied-learning|embodied]], tangible activities rather than screens? - A concern called 'technological isomorphism' describes young students mimicking AI output without understanding. How would you distinguish a child who has genuinely learned from one who is simply parroting back a polished answer? - Because young learners are more vulnerable and less able to self-regulate, adult [[scaffolding]] becomes central. What does that mean for a parent or teacher deciding when and how a child should use AI? - If access to AI-rich early learning — or to protective adult guidance — is uneven, how does equity show up differently in early childhood than it does for older students? - The evidence base here is largely exploratory and design-oriented rather than causal. What would you want to know before trusting a claim about an AI toy or app improving children's learning? ## Introduction Young children interact with AI increasingly early — through AI-enabled toys, chatbots, [[adaptive-learning|adaptive learning]] platforms, and classroom robots — yet they differ developmentally from the K-12 and higher-education learners who dominate most AI-in-education [[research-methods-aied|research]]. This concept gathers the knowledge base's coverage of that younger band. It sits within the broader [[k-12|K-12]] umbrella but is distinct because the concerns are developmentally specific: [[game-based-learning|play]] as a primary learning mode, adult (parent/teacher) scaffolding, age-appropriate [[ai-literacy|AI literacy]], and heightened attention to [[well-being]] and [[pedagogical-safety|safety]]. ### How the research clusters - **AI literacy through unplugged play.** [[ai-play-framework-early-childhood-2026|The AI-Play framework]] teaches early-childhood AI concepts through *unplugged* (no-computer) activities, showing that abstract AI ideas can be introduced to young children through embodied, playful, and tangible activities (see [[game-based-learning]]) — a developmentally appropriate route into [[ai-literacy]] and [[computational-thinking]]. - **AI-enabled toys and child development.** [[ai-toys-child-development-2026|Research on AI in toys]] examines how commercial AI-enabled playthings shape child development and play. This raises open questions about [[pedagogical-agent|agents]] in play, [[trust-calibration|trust calibration]], [[agency]], and [[well-being]] for the youngest learners — an area where design guidance is thinner than for school-age curricula, and where parents/guardians become central stakeholders. - **Children's attitudes toward age-tailored chatbots.** [[vahedian-children-attitudes-ai-chatbot-2026|Vahedian Movahed & Martin (2025)]] deployed "Ask Me Anything" (AMA), a topic-bounded chatbot (astronomy, sneakers and shoes, dinosaurs) tailored by age via [[prompt-engineering|prompt engineering]], for 63 children (ages 6–14), revealing three patterns in child–[[student-ai-interaction|AI interaction]] — wonder and curiosity, testing trust and building confidence, and building relationships through anthropomorphization — alongside a broad openness to and high trust in AI as an information source. Their call for age-sensitive design and explicit [[teacher-role|teaching]] of [[privacy|digital-safety]] and data-use concepts reinforces the developmental and [[pedagogical-safety|safety]] concerns central to this concept. - **Elementary subject learning.** [[ai-powered-personalized-learning-elementary-fractions-2026|AI-powered personalized learning]] in elementary fractions and [[awareness-technological-isomorphism|AI in elementary math]] (including concerns about "technological isomorphism" — students mimicking AI output without understanding) show both the promise and the pitfalls of AI in early subject instruction. [[elementary-writing-genai-systematic-review-2026|A systematic review of elementary writing and GenAI]] maps how [[generative-ai|generative AI]] is reshaping [[writing-education|writing instruction]] in the early grades. - **Robots and young children.** [[tsingidou-ct-robotics-kindergarten-2026|Robotics in kindergarten]] supports computational thinking through [[educational-robotics|educational robots]], and [[icub-humanoid-storytelling-llm-hri-2025|LLM-powered humanoid storytelling]] explores whether parents will accept robots as narrative play partners for their children — foregrounding [[trust]] and adult attitudes. - **A "Creative Project Approach" framework for AI agents and robotics.** [[creative-project-approach-ai-early-childhood-2025|Yang, Li and Lee (2025)]] propose a five-step pedagogical framework that integrates [[generative-ai|generative AI]] agents and robots into the Project Approach to foster young children's [[creativity|creative learning]]. They argue physical agents are more appropriate than screens for young children, and pair two complementary paradigms with their [[learning-theories|learning theories]] — **coding robots** (Papert's [[constructivist|constructionism]]; [[computational-thinking|computational thinking]] through tangible programming with tools like Matatalab and KIBO) and **generative social robots** ([[sociocultural-learning|Vygotskian]] scaffolding within the child's Zone of Proximal Development, acting as conversational peers or tutors). The five steps — identify learning needs, facilitate teacher-guided child–robot interaction, situate AI in contexts, calibrate the automation/creativity balance, and evaluate outcomes — offer designers and ECE practitioners a concrete route for embedding [[agentic-ai|AI agents]] and [[educational-robotics|robots]] in play-based, project-based classrooms while preserving teacher facilitation and child [[agency]]. - **Play-centered AI literacy curricula.** Lee (2026) reports the [[design-based-research|design-based research]] of Play With AI (PL-AI), a developmentally appropriate [[curriculum-design|curriculum]] in which two pre-K and two kindergarten teachers co-designed seven sequenced activities blending unplugged play (composing and testing "how-to" algorithms), tangible coding (Bee-Bot, Ozobot), and guided dialogue with a social AI robot. [[formative-assessment|Formative]] evidence showed substantial growth in teacher confidence and [[pedagogy|pedagogical]] agency, with four design principles emerging — **embodied play, tangible coding, guided dialogue, and teacher co-design** — aligned with the AI4K12 Initiative and NAEYC frameworks. This is a model of how [[educational-robotics|robots and tangible tools]] can be embedded in play-based, teacher-led [[ai-literacy|AI literacy]] for the youngest learners. - **AI-assisted lesson planning for children's STEAM arts teachers.** An experimental study of children's art teachers ([[luo-tahir-chatgpt-steam-lesson-planning-2026|Luo and Tahir 2025]]) found ChatGPT-assisted lesson plans were rated significantly higher in quality than teacher-generated ones by six expert professors (median 20.5 vs. 17.6, p = .002, large effect), while mapping how teachers delegated — most often (60%) to fill content gaps in a self-outlined lesson rather than to generate whole plans. Teachers valued AI chiefly as an inspiration and gap-finding tool for interdisciplinary [[stem-education|STEAM]] integration and faster topic selection, but flagged outputs that were impractical for real classrooms, overlooked child-safety constraints (e.g., carving knives for young children), and carried Western-centric cultural bias — concrete, developmentally grounded caution for [[generative-ai|generative AI]] use with young learners that a Role–Instructions–End Goal prompt framework was built to temper. - **AI-supported critical media literacy in the elementary years.** Demir and Akar (2026) evaluate an 18-hour, 5E-model critical media literacy program for fourth-grade students in a Turkish public primary school, embedding [[generative-ai|generative AI]] (ChatGPT, Grammarly, Canva AI, Padlet) phase-by-phase as a pedagogical agent rather than an isolated add-on, with activities aligned to the Turkish Language and Social Studies curricula. The AI-supported group showed large gains in media reading (+3.50), writing (+1.67), and total media literacy (+5.17, all p < .01) with between-group effect sizes of Cohen's *d* = 1.12–1.31, while the control group advanced only modestly. [[qualitative-research|Qualitative]] analysis of interviews, student artifacts (posters, drawings, slogans), and classroom observation surfaced six domains of critical media literacy growth — digital self-protection and [[privacy|data privacy]], purposeful and responsible media use, safe communication and boundary awareness, critical evaluation and misinformation awareness, online risk awareness, and media [[ethics]]/digital citizenship — offering a rare quasi-experimental, curriculum-aligned model of how AI can support critical [[ai-literacy|AI literacy]] and [[critical-thinking|critical thinking]] in the elementary grades. - **AI-rated observation of teacher–child interaction quality.** [[ai-rated-classroom-observation-scores-2026|Fong et al. (2026)]] benchmark an [[llm]] against trained human raters on the full CLASS Pre-K framework in Hong Kong kindergartens — 87 video-recorded observations from 38 classrooms across 30 kindergartens, rated from transcripts by GPT-5.0 and compared with eight trained raters. Agreement was moderate overall (weighted κ = 0.681) but conditional on the construct: convergence held for the Emotional Support domain and for Quality of Feedback — the dimension carried by explicit, exchange-based verbal support — while Classroom Organization diverged entirely and raters rated the emotional dimensions higher than the model did, at a very large gap on the reverse-scored Negative Climate dimension (d = 2.732; raters gave the maximum of 7 in 63 of 71 observations, AI clustered at 6). The early-childhood lesson is developmentally specific: quality that lives in nonverbal, spatial and routine behavior — management, movement, warmth, tone — is invisible to a transcript-only pipeline, the model's error direction is construct-dependent rather than uniformly conservative, and the authors therefore position [[automated-assessment|AI scoring]] as a screening and reflection tool for [[teacher-role|teachers]] rather than a substitute for trained observers. ### Developmental and equity considerations Because young learners are more vulnerable and less able to self-regulate their use of AI, this cluster emphasizes **scaffolding by adults** (parents, guardians, and teachers), **age-appropriate design**, and the risk of [[cognitive-offloading|over-reliance]] and [[well-being|harm]] if AI substitutes for, rather than supports, the developmental work of play, discovery, and effortful learning. Equity is a live concern: access to AI-rich early learning (or to protective adult guidance) is uneven, connecting to [[digital-divide]] and [[equity-in-ai-education]]. Much of the evidence base remains exploratory or design-oriented rather than causal. ## Connected Concepts - [[k-12]] - [[ai-literacy]] - [[computational-thinking]] - [[educational-robotics]] - [[pedagogical-agent]] - [[game-based-learning]] - [[experiential-learning]] - [[ai-education]] - [[math-education]] - [[writing-education]] - [[well-being]] - [[agency]] - [[pedagogical-safety]] - [[trust-calibration]] - [[equity-in-ai-education]] - [[digital-divide]] - [[parents-and-families]] ## Connected Articles - [[preschool-teachers-ai-behavioral-intention-2026]] — Preschool teachers' behavioral intention to use AI in early childhood settings (Duan et al. 2026) - [[ai-play-framework-early-childhood-2026]] — AI-Play: unplugged AI concepts in early childhood - [[ai-toys-child-development-2026]] — AI-enabled toys and child development - [[tsingidou-ct-robotics-kindergarten-2026]] — Computational thinking through robotics in kindergarten - [[ai-powered-personalized-learning-elementary-fractions-2026]] — AI-powered personalized learning in elementary fractions - [[elementary-writing-genai-systematic-review-2026]] — Elementary writing instruction in the age of GenAI - [[awareness-technological-isomorphism]] — Technological isomorphism in elementary math - [[icub-humanoid-storytelling-llm-hri-2025]] — LLM humanoid storytelling with children - [[human-ai-complementarity-social-emotional-learning-2026]] — Human–AI complementarity in early social-emotional learning (Raave et al. 2026) - [[play-ai-pre-k-kindergarten-ai-literacy-2026]] — Play With AI (PL-AI): play-centered AI literacy curriculum for pre-K and kindergarten (Lee 2026) - [[demir-akar-ai-media-literacy-children-2026]] — AI-based critical media literacy program for children - [[luo-tahir-chatgpt-steam-lesson-planning-2026]] - [[vahedian-children-attitudes-ai-chatbot-2026]] - [[creative-project-approach-ai-early-childhood-2025]] — The Creative Project Approach: a framework for tailoring AI agents and robotics to early learning (Yang, Li & Lee 2025) - [[ai-rated-classroom-observation-scores-2026]] — I code or AI code: AI-rated CLASS Pre-K scores versus trained human raters in Hong Kong kindergartens (Fong et al. 2026) --- ## [Higher Education](https://edtechdev.github.io/aied/concepts/higher-ed/) > **Higher Education** — the integration of artificial intelligence into university [[teacher-role|teaching]], learning, assessment, and [[administrator|administration]]. Higher education is the most-studied context in the knowledge base, with over 100 articles examining how AI transforms college-level instruction, [[educational-policy-ai|institutional policy]], and [[student-experience|student experience]]. AI in higher education is both the dominant setting for [[ai-education|AIED]] [[research-methods-aied|research]] and the site where its tensions are most visible — between [[generative-ai|generative AI]]'s promise of scalable [[personalized-learning|personalization]] and its risks to [[academic-integrity|integrity]], [[cognitive-offloading|learning]], [[privacy]], and [[equity-in-ai-education|equity]]. It encompasses both undergraduate and graduate study, including professional programs such as [[medical-education|medicine]] and [[business-education|business]]. ## Questions to Consider - Higher education is where AI's tensions are most visible — scalable personalization versus risks to integrity, learning, privacy, and equity. Which of these tensions do you see playing out most concretely in your own institution? - Large studies reveal gaps between institutional policy and what students actually do with AI day to day. Why do you think official policy so often diverges from real student practice? - The page frames AI adoption as uneven — some institutions transform while others lag, shaped by organizational drivers and national policy. What factors do you think determine whether a university adapts or stalls? - As AI makes personalized one-on-one support nearly free and universal, how does that shift the value and role of the university itself? What becomes scarcer and therefore more valuable? - One finding notes a decline in study time among college students using AI. If students spend less time studying but report satisfaction, what does that suggest about what institutions should be measuring and guarding? - Given the gap between faculty self-assessed and actual AI readiness, what would meaningful faculty development look like in your context — and who should be responsible for it? ## Introduction AI in higher education research spans every function of the university: from [[intelligent-tutoring|AI tutoring]] and [[automated-assessment|automated grading]] to faculty development, academic integrity, [[governance|institutional governance]], and student support. The knowledge base's higher education articles cluster around several key themes — institutional transformation, student experience at scale, faculty and teaching, assessment and integrity, and the rapid shift in policy and practice. ### Institutional transformation [[institutional-change-framework-ai|Institutional change frameworks]] analyze how universities adapt to AI — not just at the classroom level but across policy, governance, and organizational structure. [[sangwa-epiq-ai-faculty-readiness-2026|The EPIQ-AI framework]] reframes faculty readiness as a sociotechnical alignment challenge involving epistemic, pedagogical, institutional, and quality domains. [[universities-ai-era-rethinking|Rethinking universities in the AI era]] examines whether current institutional models can accommodate AI-driven education. Adoption is not uniform: [[alrahmi-org-drivers-ai-adoption-he-2026|organizational-driver research]] identifies what enables or blocks institutional uptake, and [[ai-uk-higher-education-policy-2026|national policy analyses]] show how systemic context shapes university responses. A concrete institutional blueprint comes from [[ai-digital-transformation-liberal-arts-lingnan-2026|Qin (2026)]], who documents Lingnan University's repositioning as a "Research-Intensive Liberal Arts Institution in the Digital Era," mandating GenAI literacy for all undergraduates and embedding digital literacy across the Common Core while developing a human-in-the-loop model that foregrounds [[ethics|ethical]] reasoning and critical judgment — arguing the AI-for-education shift is an intellectual transformation, not technocentric augmentation. Macro-level mapping of 22 higher-education AI-integration interventions ([[alsheikh-mapping-ai-integration-higher-education-2026|AlSheikh et al., 2026]]) reinforces that transformation remains more aspiration than reality: graded on the [[samr-model|SAMR model]], most interventions sit at Substitution or Augmentation, integration is concentrated in [[medical-education|medicine]], engineering, and [[cs-education|computer science]] with little [[humanities-education|humanities]] presence, and evidence skews to North America and Asia with minimal [[global-south|Global South]] representation. ### Student experience at scale Large-scale studies of [[ai-in-the-wild-college|authentic student AI use]] and [[genai-availability-grades-satisfaction|GenAI availability and satisfaction]] document how students actually use AI — revealing gaps between institutional policy and everyday practice. [[genai-student-experiences-uk-he-survey-2026|Survey research]] captures how students navigate the AI landscape, and studies of [[generative-ai-reduced-study-time-math|study time]], [[ithaka-sr-ai-skills-college-graduates-2026|AI skills for graduates]], and [[student-perceptions-ai-study-productivity-2026|student perceptions of AI tools]] show that the student experience of AI is mixed — efficient but often shallower. Much of this university learning now happens online, where [[online-teaching-and-learning|online teaching and learning]] shapes both AI's benefits (scalable personalization, always-on support) and its risks ([[academic-integrity|academic integrity]], [[cognitive-offloading|cognitive offloading]]) for college students. An exploratory ML approach using SHAP analysis examined how students' perceptions and demographics relate to intended academic ChatGPT use, prioritizing [[explainable-ai|interpretability]] ([[determinants-chatgpt-use-higher-education-2026]]). In art and design education, where generative AI challenges the value of human [[creativity]], a [[project-based-learning|project-based learning]] model with [[storytelling-in-education|digital storytelling]] at its core cultivated the emotional, cultural, and narrative capacities AI lacks — evaluated through a 15-week embedded case study with 426 Chinese undergraduates who translated local cultural heritage into [[multimodal]] narratives ([[project-based-digital-storytelling-art-design-2026]]). Assistive and inclusive uses of GenAI are also reshaping the student experience for disabled learners: [[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]] — a [[qualitative-research|qualitative]] case study of 21 visually impaired undergraduates across three Palestinian universities — found GenAI tailors pace, content, and delivery to individual profiles, simplifies complex texts, and converts content across modalities, giving learners a sense of [[agency|autonomy]] and confidence in class participation while being viewed as complementing rather than replacing teachers. ### Graduate and professional education The knowledge base is weighted heavily toward undergraduate study, and it houses graduate and professional preparation inside [[discipline-specific-aied|discipline-specific]] pages ([[medical-education|medical and health-professions education]], [[business-education|business education]], [[teacher-education|teacher education]]) rather than in a separate graduate-education page. Two strands nonetheless bear specifically on master's, doctoral, and professional students: research training, and professional-degree programs. Postgraduate *research* practice is the more distinctive of the two. [[dai-chan-responsible-genai-research-ai-literacy-2026|Dai and Chan (2026)]] analyzed how 28 postgraduate research students across seven focus groups enacted [[ai-literacy]] in their GenAI-assisted research and found them to be calibrated users rather than passive adopters — matching tools to tasks by perceived stakes and disciplinary norms, retaining [[human-in-the-loop-ai|human oversight]], and drawing their own boundaries between acceptable drafting help and substitution of their reasoning. Their central finding is a policy gap: institutional GenAI guidance addresses [[teacher-role|teaching]], learning, and assessment but says almost nothing about the research process, so the authors propose researcher-oriented guidelines that [[scaffolding|scaffold]] each AI literacy dimension across the research workflow for [[research-methods-aied|researchers]] and their supervisors. [[engagement-intensity-learner-modeling|Oh, Talton and Bui (2026)]] approach the same population from the institutional side: among 93 bioscience graduate students and postdoctoral trainees in a required research-ethics course, three simple pre-instruction behavioral signals informed lightweight intake profiling for adaptive AI [[ethics]] instruction, while prior AI coursework predicted none of the five perception outcomes. AI is also reshaping the research team itself — [[ai-assisted-writing-research-teams|analysis of 147,074 publications]] since 2020 associates AI-assisted writing with smaller, junior-leaner teams producing highly cited work, reversing the decades-long "Big Science" drift toward larger collaborations — and [[persistent-ai-agents-academic-research|a single-investigator case study]] documents what sustained [[agentic-ai|agentic]] support looks like inside one researcher's own workflow. Professional-degree programs hold most of the graduate-level evidence, and it is organized by discipline. In [[medical-education|health-professions education]], [[ai-standardized-patient-scaffolding-medical-2026|a randomized controlled trial of 100 third-year medical students]] tested a multi-agent LLM system — simulated patient, [[socratic-method|Socratic]] tutor, turn-level evaluator — organized around explicit [[scaffolding]] functions for clinical interview training, while [[ai-teammate-task-distribution-medical-training-2026|the SCAN framework]] reframes trainee AI use as a problem of task distribution and real-time [[metacognition|metacognitive]] classification rather than learner *misuse*. Master's-level preparation appears in [[instructional-design-proficiency-masters-math-2026|Zhu et al. (2026)]], whose D–T–E model raised [[math-education|mathematics]]-education M.Ed. students' instructional-design competence (Cohen's *d* = 0.62) by having three LLMs act as "intelligent reviewers" in a design–feedback–reflection loop. [[genai-professionalization-metaphors-2026|How students conceptualize GenAI and their own professionalisation]] cuts across these programs, while workplace and continuing professional learning — adjacent to but distinct from graduate study — is treated under [[adult-learning]] and [[professional-training]]. ### Faculty and teaching [[educational-development]] research examines how instructors adopt, resist, or adapt to AI. [[teacher-ai-adoption-confidence|Teacher AI adoption studies]] identify confidence, support, and attitude as key predictors. [[ai-assistance-discretionary-feedback|AI-assisted discretionary feedback]] research explores whether AI increases instructor [[ai-feedback-quality|feedback quality]] and quantity, and [[luo-eaton-ai-student-feedback-ethics-2026|ethics research]] weighs whether teachers should use AI for feedback at all. [[enright-staff-perspectives-genai-2026|Staff-perspectives research]] and [[stenalt-good-education-teacher-ai-conceptions-2026|phenomenographic studies of teacher conceptions]] probe how educators understand their changing role. In large-class settings, AI can now shoulder a growing share of the feedback load: a semester-long field experiment in undergraduate macroeconomics tutorials ([[gpt4-feedback-student-activation-2026|Geschwind et al., 2026]]) found that individual GPT-4 feedback sustained the highest participation across eight open-ended tasks and produced the largest content [[learning-gains|learning gains]], positioning AI feedback as a scalable complement to lecturer feedback and a substitute for unreliable [[peer-assessment|peer feedback]]. A complementary route for large undergrad [[stem-education|STEM]] classes is a diagnostic teaching-assistant system: [[yin-arthur-ai-teaching-assistant-engineering-econ-2026|Arthur (Yin et al. 2026)]] uses a per-question [[machine-learning|ML]] backbone to give real-time, personalized feedback on Engineering Economics Calculated Formula Questions — a domain where handwritten solutions had previously blocked AI support — via a dialogue-based, question-bank web interface that balances feedback accuracy against collection efficiency. [[farazouli-navigating-uncertainty-teachers-genai-2026|Farazouli et al. (2026)]] capture the *emotional* dimension of that change: 24 Swedish university teachers experienced GAI's emergence as alarming and overwhelming, reporting a "state of vulnerability" (low confidence, insecurity, fear of "not being ahead of students") and feeling "stuck" between utopian and dystopian discourses — while rethinking [[assessment]] and re-evaluating their priorities toward [[critical-thinking|critical thinking]] and ethical GAI use. ### Assessment and integrity [[academic-integrity]] and [[ai-assessment-scale-reform|AI assessment reform]] research grapple with how universities should redesign evaluation for an AI-capable student body. [[ai-detection|Detection-centered]] approaches are giving way to [[authentic-assessment]] and process-based evaluation, including [[fenton-oral-exams-ai-authentic-assessment-2025|oral exams]], [[roe-assessment-twins-2026|assessment twins]], and [[beyond-detection-authentic-assessment-ai-2025|authentic assessment redesign]]. The shift reflects a deeper concern: [[ai-tools-academic-work-cheating-2026|how students and institutions define cheating]] with AI, and the [[assessing-quality-ai-generated-exams-field-2025|quality of AI-generated assessment materials]]. Automation of grading itself is maturing for open-ended university work: [[pecuchova-automated-grading-open-ended-genai-2026|Pecuchova, Benko & Drlik (2025)]] benchmarked eleven GenAI and sentence-embedding models on 1,885 responses to 24 software-engineering exam questions and found only GPTo1 reached almost-perfect agreement with two expert human graders (Fleiss' Kappa 0.82), with context-sensitive models outclassing reference-based ones — but its proprietary API costs led the authors to recommend hybrid human-in-the-loop deployment for resource-constrained institutions. At the professional-program level, [[olvet-genai-scoring-open-ended-medical-2026|Olvet et al. (2026)]] show that [[llm]] (GPT-4) scoring of pre-clerkship [[medical-education|medical]] open-ended exam questions can reach substantial-to-almost-perfect agreement with faculty graders (weighted kappa up to 0.94) after iterative human rubric refinement, while the authors still recommend keeping humans in the loop to arbitrate residual discrepancies. Complementing this model-level evidence, [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat, Das, Bhaumik & Thambi (2026)]] graded a 21-item university pharmacy exam with ChatGPT-5 and found near-perfect agreement with faculty on objective items (CCC 0.935–1.000) but unreliable agreement on short-answer and essay items, with rubric provision not consistently improving performance — reinforcing the case for [[human-in-the-loop-ai|hybrid, human-in-the-loop]] grading in university assessment. ## Implications for higher-education instructors - **Design assessment for an AI-capable student body.** Detection-centered integrity approaches are giving way to [[authentic-assessment|authentic]] and process-based evaluation ([[beyond-detection-authentic-assessment-ai-2025|beyond detection]], [[ai-assessment-scale-reform|assessment reform]]) — redesign what you assess, not just how you police it. - **Redesign the "what should students still learn by hand?" question.** Just as in computing, decide which skills must be preserved (verification, judgment, process) and make those the assessed core, rather than assuming AI skills transfer automatically. - **Treat faculty readiness as sociotechnical, not just technical.** [[sangwa-epiq-ai-faculty-readiness-2026|EPIQ-AI]] frames readiness across epistemic, pedagogical, institutional, and quality domains — align your [[pedagogy|teaching practice]] with institutional governance, not just tool fluency. - **Address the policy-vs-practice gap.** Large-scale studies ([[ai-in-the-wild-college|AI in the wild]]) show students use AI in ways institutional policy doesn't anticipate — align your expectations with real usage and teach [[ai-literacy]] explicitly. - **Use AI to raise feedback quality and quantity.** [[ai-assistance-discretionary-feedback|AI-assisted feedback]] can increase the feedback instructors deliver; pair it with [[human-in-the-loop-ai|human judgment]] so it improves learning rather than merely automating. - **Engage with grading and assessment innovation deliberately.** [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] argue that [[assessment]] and grading are among the most influential factors shaping learning in higher education, yet they remain under-researched in fields such as [[business-education|management education]] (only 58 articles over 20 years across four leading journals). Traditional, [[summative-assessment|summative]]-heavy and norm-referenced approaches undermine deep learning, [[well-being]], equity, and [[academic-integrity|integrity]] in the generative AI era; the authors urge instructors to engage more actively and reciprocally with five innovative practices (authentic assessment, self- and peer-assessment, reassessment, standards-based grading, ungrading), noting implementation demands institutional and cultural support and incremental experimentation. - **Treat research training and supervision as an AI-literacy site in its own right.** [[dai-chan-responsible-genai-research-ai-literacy-2026|Postgraduate researchers]] enact [[ai-literacy]] as a [[situated-learning|situated]] practice rather than a static skill set, and institutional guidance largely stops at teaching and assessment. Guidance for thesis, dissertation, and publication work should [[scaffolding|scaffold]] verification, attribution, and disclosure concretely for each AI literacy dimension — and [[engagement-intensity-learner-modeling|simple pre-instruction signals]] can target that support in research-ethics courses — instead of issuing binary rules. ## Connected Concepts - [[online-teaching-and-learning]] — Online Teaching and Learning - [[generative-ai]] — Generative AI technologies and models - [[llm]] — Large language models - [[student-experience]] — Student experience with AI in higher ed - [[educational-development]] — Faculty development and AI readiness - [[academic-integrity]] — Academic integrity in an AI-capable student body - [[ai-literacy]] — AI literacy for students and faculty - [[assessment-validity]] — Assessment validity in the age of GenAI - [[educational-policy-ai]] — Educational policy on AI - [[regulation]] — Regulation of AI in education - [[remote-proctoring]] — Remote proctoring and automated exam integrity - [[self-directed-learning]] — Self-directed learning with AI - [[problem-based-learning]] — Problem-based learning with AI - [[teacher-role]] — Evolving teacher role in AI classrooms - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) - [[arts-design-and-media-education]] ## Connected Articles - [[sangwa-epiq-ai-faculty-readiness-2026]] — EPIQ-AI Faculty Readiness Framework - [[ai-in-the-wild-college]] — AI in the Wild: College Student AI Use - [[institutional-change-framework-ai]] — Institutional Change in the Age of AI - [[genai-availability-grades-satisfaction]] — GenAI Availability and Student Satisfaction - [[ai-assessment-scale-reform]] — AI Assessment Scale and Reform - [[universities-ai-era-rethinking]] — Rethinking Universities in the AI Era - [[teacher-ai-adoption-confidence]] — AI Adoption Among Teachers - [[enright-staff-perspectives-genai-2026]] — Staff perspectives on GenAI in higher education - [[alrahmi-org-drivers-ai-adoption-he-2026]] — Organizational drivers of AI adoption in higher ed - [[luo-eaton-ai-student-feedback-ethics-2026]] — AI in student feedback: ethics - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond detection: redesigning authentic assessment in an AI-mediated world - [[genai-student-experiences-uk-he-survey-2026]] — GenAI experiences among UK higher-education students (survey) - [[ithaka-sr-ai-skills-college-graduates-2026]] — Instructors vs. employers on AI skills for college graduates - [[generative-ai-reduced-study-time-math]] — 26.9% study-time decline among college students - [[ai-uk-higher-education-policy-2026]] — UK higher-education AI policy - [[ai-tools-academic-work-cheating-2026]] — AI tools, academic work, and cheating - [[assessing-quality-ai-generated-exams-field-2025]] — Assessing the quality of AI-generated exams: a large-scale field study - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant assessment - [[roe-assessment-twins-2026]] — Assessment twins for strengthening assessment validity in the age of GenAI - [[student-perceptions-ai-study-productivity-2026]] — Students' Perceptions of Artificial Intelligence Tools for Study Productivity and Learning: An Exploratory Survey Study - [[stenalt-good-education-teacher-ai-conceptions-2026]] — phenomenographic study of university teachers' conceptions of AI - [[ai-digital-transformation-liberal-arts-lingnan-2026]] — Digital transformation of a liberal arts university toward a research-intensive model in the GenAI era (Qin 2026) - [[alsheikh-mapping-ai-integration-higher-education-2026]] — Systematic review mapping AI integration in higher ed via FACETS + SAMR frameworks (AlSheikh et al. 2026) - [[farazouli-navigating-uncertainty-teachers-genai-2026]] — University teachers' experiences and perceptions of GAI: vulnerability, rethinking assessment, student learning at risk (Farazouli et al. 2026) - [[gpt4-feedback-student-activation-2026]] - [[pecuchova-automated-grading-open-ended-genai-2026]] - [[mesny-innovative-assessment-grading-management-2026]] - [[dai-chan-responsible-genai-research-ai-literacy-2026]] — Postgraduate researchers' GenAI use in the research workflow and AI-literacy-oriented guidelines (Dai & Chan 2026) - [[engagement-intensity-learner-modeling]] — Engagement intensity as a learner-modeling signal for adaptive AI ethics instruction (Oh, Talton & Bui 2026) - [[ai-assisted-writing-research-teams]] — AI-assisted writing shifts research teams toward smaller, junior-leaner, highly cited collaborations (Wang et al. 2026) --- ## [Adult Learners](https://edtechdev.github.io/aied/concepts/adult-learning/) > **Adult learning** — the theory and practice of educating adults (andragogy), and how AI tools and technologies can be designed to support adult learners' [[agency|autonomy]], prior experience, and real-world relevance. Explored across 9 articles in this knowledge base. ## Questions to Consider - Adult learning theory assumes learners are self-directed, draw on life experience, and want real-world relevance. But if AI performs much of the cognitive work, does a learner completing a task without visible help actually prove they directed it? What would make you confident they did? - Behavioral independence from a tool no longer guarantees the learner directed the learning. If you were designing AI for adult learners, what would you look for to confirm genuine self-direction rather than quiet delegation? - Design guidelines for adult AI tools emphasize fitting into busy lives — mobile-friendly, offline-capable, connected to real problems. Which of these matter most for your own learning, and what does a tool that ignores them cost the learner? - Adult learners often study at work or at home, so online delivery dominates — bringing flexibility but also risks of offloading and integrity questions. How does the convenience of AI assistance interact with the goal of durable learning for a busy adult? - Research found no single adult-learning AI system satisfied all design guidelines — the full ecosystem was needed. What does that suggest about expecting one tool to meet every learner's needs? - For marginalized and neurodivergent adult learners, equity may be less about tool access than about the educator's relational care. How does positioning the human as the locus of care change how you would design or adopt an AI tool? ## Introduction Rooted in Knowles's andragogical model, adult learning assumes learners are self-directed, draw on life experience, are motivated by immediate and practical goals, and benefit most when learning connects to their real-world roles. These assumptions matter for AI design because generative AI can now participate in almost every stage of learning — identifying needs, setting goals, interpreting information, producing outputs, and evaluating performance. When AI performs so much of the cognitive work, behavioral independence from the tool no longer guarantees that the learner actually directed the learning. Research in this knowledge base accordingly reframes self-direction as an active design goal rather than an assumed default, and evaluates adult-learning AI against criteria like goal ownership, delegation control, and cognitive recoverability. ## Evidence from connected articles - **Andragogy and GenAI cognitive delegation.** [[andragogy-cognitive-delegation-genai-2026|Hyoung (2026)]] revisits Knowles's six andragogical assumptions under AI-mediated cognitive delegation, arguing that completing a task without visible AI help does not prove meaningful self-direction. It derives five analytical dimensions — need and goal ownership, delegation control, epistemic calibration, cognitive recoverability and transfer, and motivational autonomy — bridging adult learning with [[cognitive-offloading]] and [[self-regulated-learning]] to assess whether learners remain genuinely self-directed in the [[generative-ai]] era. - **Design guidelines for AI adult-learning tools.** Drawing on longitudinal deployment data from the National AI Institute for Adult Learning and Online Education (AI-ALOE), the DIS 2026 paper [[ai-adult-learning-guidelines-dis2026|synthesizes 19 empirically grounded design guidelines]] for AI-powered adult-learning technologies. Derived from ~1,600 stakeholder statements across seven deployed systems, the guidelines span cognitive, social, and teaching presence (a Community of Inquiry framing) and emphasize that tools should fit into busy adult lives (mobile-friendly, offline-capable), connect content to real-world problems, personalize meaningfully, provide substantive support and feedback, and be transparent about data. No single system satisfied all guidelines; the full AI-ALOE ecosystem was needed to cover them. - **Adult, distance, and lifelong learning contexts.** [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of|Rienties et al.]] show how the Open University designed and evaluated an embedded AI assistant (AIDA) through six design-based-research studies; students using it spent twice as long on the course, though the study warns that technical capability must be matched by [[governance]] and organizational readiness. [[ai-lifelong-learning-policy|Theodora and Tselios]] frame AI's dual role in adult and [[lifelong-learning]] as both an enabler of personalized, scalable education and a source of equity and governance risk, calling for inclusive, human-centered policy. [[community-centered-ai-education-adults|A Midwestern case study]] co-designed an AI-literacy program for 54 adults in an underserved community, finding that equity-oriented adult AI education must address foundational [[ai-literacy|digital literacy]] gaps, build [[trust]] around data privacy, and connect to lived experience. [[sovereign-hive-titl-further-education-2026|Herron's "Sovereign Hive" / Tutor-in-the-Loop framework]] treats GenAI equity in Further Education as atmospheric regulation rather than mere tool access, positioning the educator as the locus of relational and cognitive care for marginalized and [[neurodiversity|neurodivergent]] adult learners. ## Connections to related concepts Adult learning sits at the intersection of several closely linked concepts in this knowledge base. [[higher-ed]] supplies the institutional context in which much adult and distance learning occurs, while [[professional-training]] covers its workforce and [[lifelong-learning]] its continuous-education dimension. [[vocational-education|Vocational education and training]] is the neighbour that names an occupation rather than a learner: adult learning describes what learners bring to any context — self-direction, prior experience, immediate practical goals — whereas VET names initial, practice-proximal preparation for a named trade or technical role, judged by demonstrated competence with equipment and framed by qualification frameworks rather than by andragogical dispositions. [[online-teaching-and-learning|Online teaching and learning]] is the dominant delivery medium for adult learners — who often study at work or at home — so its affordances (24/7 access, asynchronous support) and risks ([[academic-integrity|integrity]], [[cognitive-offloading|offloading]]) are central to adult-learning design. [[self-regulated-learning]] and [[agency]] name the learner capacities that AI must protect rather than erode, and [[cognitive-offloading]] captures the mechanism by which AI can either support or undermine them. [[inclusive-learning]] and [[equity-in-ai-education]] frame the equity obligations of adult AI tools, [[human-in-the-loop-ai]] names the design pattern that keeps humans accountable, and [[scaffolding]] describes the graduated support such tools should provide. ## Implications for adult-education instructors and designers - **Design AI as a scaffold for self-direction, not a substitute.** Behavioral independence from the tool doesn't prove the learner directed the learning — protect goal ownership, delegation control, and cognitive recoverability ([[andragogy-cognitive-delegation-genai-2026|andragogy + cognitive delegation]]). - **Fit into busy adult lives.** Make tools asynchronous, mobile, and offline-capable, and connect content to real-world problems ([[ai-adult-learning-guidelines-dis2026|AI-ALOE guidelines]]). - **Keep a human in the loop.** Position the educator as the locus of relational and cognitive care, especially for marginalized and [[neurodiversity|neurodivergent]] adult learners ([[sovereign-hive-titl-further-education-2026|Tutor-in-the-Loop]]). - **Address foundational digital literacy and data trust.** Build [[ai-literacy]] and [[trust]] around data privacy before expecting adoption ([[community-centered-ai-education-adults|community AI education]]). - **Ground AI in learning science and andragogy, and prefer deep personalization.** Apply andragogical theory and connect content to real-world problems; favor deep personalization (task sequencing, difficulty calibration) over surface-level adaptation. - **Make transparency and community features first-class.** Data-practice transparency and social/community features are among the most neglected yet most valued dimensions of adult AI tools. - **Treat technical and structural reliability as a precondition.** Engagement depends as much on stable, inclusive infrastructure as on pedagogical quality — unstable or exclusionary platforms undermine otherwise sound design. - **AI design principles for andragogy.** [[kim-ai-andragogy-2026|Kim et al. (2026)]] find adult learners value AI as a collaborative learning agent and derive three AI design principles for andragogy: human-in-the-loop (shared mental models, human-AI co-creation), emotional design (calibrating AI reliance, empathetic communication), and adaptability (continuous adaptation, interoperability). Their eleven scenario prototypes also map each andragogical principle onto a concrete AI affordance: [[intelligent-tutoring|AI tutors]] and [[learning-by-teaching|teachable agents]] for involvement, monitoring and [[learning-analytics|analytics tools]] for autonomy and self-assessment, empathetic [[conversational-ai|chatbots]] and [[simulation|simulations]] for experience, case libraries and higher-order question generators for problem-centered work, and AI planners and career coaches for relevance. ## Connected Concepts - [[self-directed-learning]] - [[online-teaching-and-learning]] — Online Teaching and Learning - [[higher-ed]] - [[professional-training]] - [[vocational-education]] - [[lifelong-learning]] - [[inclusive-learning]] - [[agency]] - [[self-regulated-learning]] - [[generative-ai]] - [[human-in-the-loop-ai]] - [[cognitive-offloading]] - [[scaffolding]] - [[trust]] - [[ai-literacy]] - [[equity-in-ai-education]] - [[neurodiversity]] - [[governance]] - [[formative-assessment]] - [[rct]] - [[active-learning]] - [[discipline-specific-aied]] ## Connected Articles - [[ai-adult-learning-guidelines-dis2026]] — Guidelines for Designing AI Technologies to Support Adult Learning - [[andragogy-cognitive-delegation-genai-2026]] — What Remains Self-Directed? Revisiting Andragogy Through Cognitive Delegation in Generative AI-Mediated Adult Learning - [[ai-lifelong-learning-policy]] — Artificial Intelligence in Lifelong Learning: Opportunities and Challenges in Adult Education Policy - [[sovereign-hive-titl-further-education-2026]] — The Sovereign Hive and the Tutor-in-the-Loop (TITL) Framework for Equity in Further Education - [[community-centered-ai-education-adults]] — Co-Designing Community-Centered AI Education for Adults: A Midwestern Case Study - [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of]] — New Systems of Learning for Distance Learning Institutions? A Six-Study Review of Implementing AIDA - [[institutional-governance-ai-universities]] — Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools - [[generative-ai-enhanced-learning-experiences-for-computational-thinking-a-systema]] — Generative AI-enhanced learning experiences for computational thinking: A systematic scoping review and design guidelines - [[ai-assisted-se-curriculum-syllabus-analysis-2026]] — Mapping the Emerging Curriculum for AI-Assisted Software Engineering via Syllabus Analysis - [[learner-ai-interaction-patterns-oop]] — Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course - [[dot-framework-survey-2026]] - [[kim-ai-productive-failure-adult-2026]] — Designing AI Systems to Support Productive-Failure-Based Learning - [[kim-ai-andragogy-2026]] — AI Applications in Supporting Andragogy (Kim et al. 2026) --- ## [Vocational Education and Training](https://edtechdev.github.io/aied/concepts/vocational-education/) > **Vocational education and training** — the segment of education that prepares people for named occupations, trades and technical roles, organized around practice-proximal competence rather than disciplinary knowledge. Where [[professional-training|workplace learning]] describes upskilling for the already employed, VET includes initial preparation for a trade; where [[higher-ed|higher education]] names degree study, VET is often non-degree and framed by national qualification frameworks. Its defining features are that learners are assessed on what they can do with equipment, that instruction happens near the workshop, simulator or worksite, and that the human trainers who carry practical instruction are frequently the binding constraint. In AI research VET appears both as a distinct learner population — one whose academic confidence is tied to demonstrated skill and occupational identity — and as a distinct evidence base, thinner and more fragmented than the school or university literature. ## Questions to Consider - If a qualification certifies what a learner can do, how much AI-assisted learning counts as authentic practice and how much substitutes for the repetition that builds competence? - What is lost when an AI agent role-plays the counterpart — patient, pilot, client — that a human trainer used to play? Which judgments can a synthetic counterpart not model? - Oral assessment scales poorly with class size. If AI surfaces evidence but does not judge, which parts of assessment capacity are relieved and which are merely displaced to the teacher? - No study in the AI-in-VET evidence base is set in a workplace, yet VET is defined by work-based learning. What would credible research where the learning happens require? - The EU AI Act treats AI evaluating learning outcomes in vocational training as high risk, while some jurisdictions have none. Should procurement follow the strictest standard available? ## Introduction Vocational education and training prepares people for specific occupations — automotive technicians, air traffic controllers, interior designers, care workers — and its currency is demonstrated competence rather than accumulated credit. Assessment tends to be performance-based, instruction is tied to equipment, and the skilled trainers who supervise practice are scarce. Its neighbours in this knowledge base differ chiefly in scope. [[professional-training]] covers workplace and corporate upskilling, mostly for people already employed; VET covers initial occupational preparation as well. [[adult-learning]] names learner characteristics rather than occupational specificity. [[higher-ed]] names degree-granting study, whereas much VET is organized by qualification frameworks such as NZQA unit standards or the EQF. [[stem-education]] and VET overlap in technical domains, but STEM education aims at conceptual understanding while VET aims at usable procedure. [[career-development-and-readiness]] names the employability dispositions VET programs are judged by; VET names the instructional system held accountable for producing them. What is distinctive about AI in VET is a gap between what the sector says it wants and what it builds: constructivist theory is widely espoused while behaviorist drill-and-practice systems dominate and learner-agency designs remain rare. The recurring design question is not whether AI can deliver instruction, but whether it can absorb the parts of vocational learning that are expensive to staff — realistic scenarios, timely feedback, role-played counterparts, oral evidence — without displacing the practice that produces competence. ### How AI appears in vocational education and training - **A young, fragmented and geographically concentrated evidence base.** The first systematic review of AI in VET ([[ai-vocational-education-training-review]]) identified 26 empirical studies published 2015–2026 through ERIC, Web of Science and Elicit under PRISMA guidelines: nine technical-domain, nine domain-general, five in business administration and three in health. Settings were six classroom, eight online, four blended and eight simulation-based — and none in a workplace, despite VET's work-based character. Seventeen of the 26 originated in Asia; only nine studies shared at least one reference and none cited each other. Five were randomized experiments and 21 used pre-experimental or quasi-experimental designs, mostly measuring outcomes immediately after intervention, and only three gave learners an active role in an AI-empowered design. The authors warn of an educational "Turing Trap" — using AI to replicate human instruction rather than augment [[human-in-the-loop-ai|human judgment]] — and call for failure cases and boundary conditions in place of the prevailing success narrative. - **Simulation absorbs the scarce role-player.** [[astra-atco-training-simulator]] targets a capacity constraint in air traffic control training: *simpilots*, specialized human trainers who role-play both pilots and controllers in a simulated airspace. ASTRA substitutes autonomous LLM-driven sim-pilots, keeping scenario complexity while removing the staffing bottleneck and enabling [[adaptive-learning|adaptive]] practice at scale. It is a system description rather than an efficacy trial, but it names a mechanism that recurs across AI in VET: where the scarce input is a skilled human playing a counterpart, an agent can hold the role and let practice expand. - **Immersive, agent-supported project work can raise design ability — selectively.** [[ai-ive-pbl-vocational-design-creativity-2026]] specifies AI-IVE-PBL, a four-dimension, five-phase model (discovery, envisioning, modeling, communication, refinement) run in VR with an LLM-backed digital-human assistant in a first-year interior design course at a Chinese vocational college. In a 12-week two-group quasi-experiment (63 valid responses; 31 vs 32), the immersive-agent condition scored higher on design ability (η²p = .138) and creative ability (η²p = .111) under ANCOVA, with cognitive engagement d = 0.90, behavioral engagement d = 0.75, motivation d = 0.74, satisfaction d = 0.69, and cognitive load lower (d = −0.52). Innovative thinking and affective engagement did not reach significance, attributed by the authors to short-term ceilings on entrenched cognitive patterns. Every outcome is [[self-report-measures|self-report]], with no design artifacts or expert ratings. Read against [[genai-xr-architectural-design-education-2026]], where a GenAI-plus-XR pipeline produced declining design [[self-efficacy]] and no blinded-panel advantage, the difference looks less like hardware than like who holds the phase structure and the rubric. - **Assessment is where the authenticity problem is sharpest.** [[ai-supported-oral-assessment-tvet-2026]] documents AkoVoice, trialled in four Level 3 Automotive and one Level 3 Engineering class, designed so AI surfaces rubric evidence while the human assessor judges. Of 33 surveyed learners, 21 (64%) agreed the voice task was realistic and the same proportion said it gave a clear way to communicate what they knew; none disagreed that speaking in real time suited this [[authentic-assessment|authentic assessment]] better than a written portfolio. Word counts for identical questions varied five- to eight-fold between learners without improving accuracy on fact-based questions, and nine learners answering in 2 to 13 words were all marked correctly, the shortest being two words, "3500 kgs", matched on value. A full cycle of capture, storage, AI judgment drafting and teacher reporting ran offline on one Windows laptop with 8 GB of graphics memory (Mistral 7B via Ollama, faster-whisper, Chatterbox), assessing up to 12 learners at once in a workshop where steel framing defeats wifi, with recordings encrypted and deleted after 90 days. The paper notes the EU AI Act treats AI evaluating learning outcomes in vocational training as high risk, while New Zealand has no sector-specific TVET framework. - **Bounded accompaniment rather than substitution.** [[ai-pedagogical-accompaniment-amico]] argues that the value of AI in technical and vocational settings depends on accountable [[pedagogy|pedagogical]] mediation rather than human-likeness. Its Amico prototype pairs AmicoMio, oriented to technical clarity and step-by-step task guidance, with AmicoTuo, oriented to reflective dialogue and maieutic questioning. The design principle is a *relational bridge*: interaction deliberately temporary, directional toward human contact, and bounded by safeguards, with adults retaining human-in-command responsibility. Exploratory pilots (N = 30, Italy and China, 20 bounded sessions) found participants treated the system as a bounded support tool, with no reported expectation of substitution or dependency. - **AI-assisted learning carries a psychological cost when it replaces effort.** [[ai-autonomous-learning-accomplishment-2026]] surveyed 1,264 vocational college students in China using structural equation modeling and found AI-assisted autonomous learning negatively associated with hardiness (commitment, control, challenge) and positively associated with reduced academic accomplishment, a burnout dimension of negative self-evaluation. Hardiness partially mediated the relationship. The design is cross-sectional, self-report and single-institution, so causality is not established, but the framing matters for VET: where confidence is built through repeated practice, AI as a substitute may reduce both the disposition to persist and the felt experience of mastery. - **Accelerating curriculum production without surrendering verification.** [[crewscaler-ai-upskilling-framework]] applies AI across all five stages of professional upskilling — knowledge acquisition, content development, review and verification, [[intelligent-tutoring|tutor coaching]] and assessment development — while keeping blueprint design, subject-matter review and misconception authoring with humans. Its external validation includes NASBA CPE accreditation, three of three learners passing an NVIDIA certification exam using only the framework's knowledge base, and a 530-question bank tagged to a 53-skill blueprint. It belongs here because it treats verification as a first-class stage: [[hallucination-risk|hallucination]] detection is largely absent from education pipelines, and the paper reports default LLM tutoring reaching only 52–70% correct pedagogical actions. ## Connected Concepts - [[professional-training]] - [[adult-learning]] - [[higher-ed]] - [[career-development-and-readiness]] - [[simulation]] - [[authentic-assessment]] - [[intelligent-tutoring]] - [[human-in-the-loop-ai]] - [[cognitive-offloading]] - [[self-efficacy]] ## Connected Articles - [[ai-vocational-education-training-review]] — First systematic review of AI in VET: purposes, theory and empirical effectiveness - [[ai-ive-pbl-vocational-design-creativity-2026]] — AI-IVE-PBL: immersive project-based learning and design creativity - [[ai-supported-oral-assessment-tvet-2026]] — AkoVoice: offline AI-supported oral assessment in TVET - [[ai-pedagogical-accompaniment-amico]] — Amico dual-mode prototype: design principles and observable indicators - [[ai-autonomous-learning-accomplishment-2026]] — AI-assisted autonomous learning and reduced accomplishment, mediated by hardiness - [[astra-atco-training-simulator]] — Autonomous sim-pilots for scalable ATCO training - [[crewscaler-ai-upskilling-framework]] — AI-accelerated end-to-end framework for rapid professional upskilling - [[genai-xr-architectural-design-education-2026]] — Counter-case: GenAI plus multi-user XR with declining design self-efficacy --- ## [Special Education](https://edtechdev.github.io/aied/concepts/special-education/) > **Special Education** — the design and delivery of instruction for learners with disabilities, spanning cognitive, physical, sensory, and neurodevelopmental differences. [[ai-education|AI in education]] [[research-methods-aied|research]] in this knowledge base explores how AI tools can support diverse learner needs through [[personalized-learning|personalization]], [[scaffolding|adaptive scaffolding]], and accessible interfaces — while also examining the risks of [[ai-technologies|AI systems]] that overlook or marginalize disabled learners. > ⚠️ **Special Education is primarily a [[k-12]] term.** It is rooted in the U.S. Individuals with Disabilities Education Act (IDEA) and the entitlement-based system of Individualized Education Programs (IEPs) that governs special-education services in primary and secondary schooling. In [[higher-ed|higher education]] — and increasingly in K-12 as well — the more common framing is **[[universal-design-for-learning|Universal Design for Learning]]** (a proactive design framework that benefits all learners) alongside [[accessibility]] and [[assistive-technology]] rather than "special education." A K-12 special-education article and a college UDL piece are about overlapping but distinct contexts; the knowledge base keeps both because the research literature spans both. When a source concerns higher education and disabled learners, it is usually better linked to [[universal-design-for-learning]], [[accessibility]], or [[inclusive-learning]] than to special-education. ## Questions to Consider - The page emphasizes that 'special education' is primarily a K-12, entitlement-based term rooted in IDEA and IEPs, while higher education more often speaks of Universal Design for Learning and accessibility. Why do you think these contexts differ, and what does that difference reveal? - AI's promise of personalization seems tailor-made for learners with diverse needs. But if a system can adapt to 'individual cognitive profiles,' what could go wrong when the model of a disability is too coarse or absent altogether? - The research includes AI tools designed for specific disability profiles (e.g., dyslexic or Deaf and Hard of Hearing learners). What risks do you see in designing for narrow profiles versus designing universally for all learners from the start? - How might AI systems that are built for the 'average' learner end up overlooking or marginalizing disabled learners, even unintentionally — and whose responsibility is it to prevent that? - What would it mean for an AI tool to genuinely include, rather than merely accommodate, a learner with a disability — and how would you recognize the difference in practice? ## Introduction Special education is a domain where AI's capacity for personalization and adaptation offers particular promise. Unlike one-size-fits-all instruction, [[intelligent-tutoring|AI tutors]] can theoretically adapt to individual cognitive profiles, communication needs, and learning paces. The articles in this knowledge base span AI for specific disability profiles, neurodivergent learner experiences, and critical perspectives on AI and disability. **Disability-specific AI tutoring** tailors AI to particular learner needs. **[[special-r1-rl-special-education|Special-R1]]** extends [[reinforcement-learning|reinforcement learning]] to model cognitive and communicative diversity across five disability profiles, using persona-aware prompts and thinking rewards to shape tutor responses for each learner. **[[dyslexlens-dyslexic-learners-ai|DysLexLens]]** analyzed how dyslexic learners experience AI tools, revealing both the value of AI for literacy support and persistent accessibility barriers. **[[llm-question-generation-deaf-hard-of-hearing-2026|Chen et al.]]** designed an [[llm]]-powered question-generation system for [[accessibility|Deaf and Hard of Hearing learners]], introducing Visual and Emotion question strategies and iteratively refining questions with the target community to overcome the mismatch between text-based AI prompts and sign-based first languages. **[[embodied-string-learning-blindness-low-vision-musicians]]** developed non-visual learning strategies with blind and low-vision musicians, centering disability-led [[embodied-learning|embodied]] design. These connect to [[inclusive-learning]] and [[neurodiversity]]. **Neurodivergent learner experiences** center autistic and ADHD students. **[[neurodivergent-computing-students|Zastudil et al.]]** found neurodivergent computing students need structured assignments, small consistent teams, and explicit role definitions — design requirements that [[collaborative-learning]] tools must address. **[[adhd-video-segmentation-computing-education]]** demonstrated that AI-segmented videos eliminated the ADHD performance gap. Both connect to [[learning-design]] and [[universal-design-for-learning]]. **Critical perspectives** examine how AI can marginalize disabled learners. **[[genai-minoritized-knowledges-disability|Tali-Otmani]]** argues that AI systems actively marginalize disability-centered knowledge due to Western-centric training data — connecting to [[equity-in-ai-education]] concerns about epistemic justice. **AI for dyslexia across detection, support, and personalized learning.** A 2026 interdisciplinary [[meta-analysis-systematic-review|systematic review]] (Dabaghi, D'Urso & Sciarrone, PRISMA-guided, 2018–2024, n=72) maps how AI supports students with dyslexia in education, finding AI used for detection, assistive support, and personalized learning — but with these strands evolving in parallel rather than in integration, driven more by technological opportunity than by consolidated educational theory. ML-based help-education tools fall into five areas (specific applications, engagement, personalization, recommendation, generic support) yet emphasize technical performance and classification accuracy while overlooking ecological validity and practical classroom deployment. Detection research (EEG, eye-tracking, ML models) prioritizes early intervention and shows diagnostic promise, but often requires specialized equipment and controlled environments, limiting scalability and accessibility in typical school settings. Open challenges include limited experimental validation, scalability, [[ethics]]/privacy concerns with sensitive student data, limited teacher support and training, and language/cultural barriers (most research targets English-speaking populations). **Cognitive offloading for students with learning disabilities (SWLDs).** [[seung-basham-cognitive-offloading-swld-2026|Seung & Basham (2026)]], a conceptual review in a *Learning Disability Quarterly* special series on AI for students with LD, reframe [[generative-ai|GenAI]] use for SWLDs through the [[cognitive-offloading]] lens. They argue that GenAI can be a **compensatory aid or a shortcut** depending on how offloading decisions interact with SWLDs' cognitive and [[motivation|motivational]] profiles (executive-function and working-memory challenges, heightened cognitive load, effort-avoidant performance goals, lower academic self-efficacy, and inflated expectations toward GenAI) and with instructional design. For reading and writing, GenAI can scaffold access (text leveling, summarizing, [[multimodal]] outputs, planning, drafting, revision feedback) while preserving higher-order [[student-engagement|engagement]] — but excessive offloading risks bypassing the comprehension, planning, and monitoring processes that are already fragile for these learners, fostering "[[metacognition|metacognitive]] laziness" and compounding literacy difficulties across domains. The paper positions **instructional [[guardrails]]** as the key moderating factor and recommends [[teacher-role|teaching]] strategic offloading, building [[ai-literacy]] to calibrate tool trust, sequencing mastery experiences to build [[self-efficacy]], and aligning tasks and assessment with IEP goals that prioritize skill development over substitution. This extends the knowledge base's special-education coverage to the equity dimension of offloading: the same tool that lowers barriers to access can, if unguarded, substitute for the practice SWLDs need most. ## Implications for special-education instructors - **Co-design AI with the target learners and community.** [[llm-question-generation-deaf-hard-of-hearing-2026|Question generation for Deaf/Hard-of-Hearing learners]] shows the value of iteratively refining AI with the community to bridge the gap between text-based prompts and sign-based first languages — involve learners and their communities in design rather than assuming AI fits them. - **Match AI to specific disability profiles, not generic accessibility.** [[special-r1-rl-special-education|Special-R1]] models cognitive and communicative diversity across disability profiles; [[dyslexlens-dyslexic-learners-ai|DysLexLens]] documents both the literacy value and the persistent accessibility barriers dyslexic learners face — choose tools aligned to each learner's profile and be alert to unmet barriers. - **Structure collaboration for neurodivergent learners.** [[neurodivergent-computing-students|Neurodivergent computing students]] need structured assignments, small consistent teams, and explicit roles — apply these design requirements to any AI-mediated collaborative activity. - **Use AI to close (not widen) performance gaps.** [[adhd-video-segmentation-computing-education|AI-segmented videos]] eliminated the ADHD performance gap — deploy adaptive AI where evidence shows it equalizes outcomes, not where it merely automates. - **Center disability-led embodied design.** [[embodied-string-learning-blindness-low-vision-musicians|Blind/low-vision musicians]] research shows non-visual, disability-led strategies outperform default visual interfaces — build and adapt AI with disabled learners' expertise. - **Guard against epistemic marginalization.** [[genai-minoritized-knowledges-disability|Critical perspectives]] warn that Western-centric training data can marginalize disability-centered knowledge — audit AI content and tools for epistemic justice alongside [[equity-in-ai-education]]. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[inclusive-learning]] - [[equity-in-ai-education]] - [[neurodiversity]] - [[universal-design-for-learning]] - [[learning-design]] - [[student-experience]] - [[ai-literacy]] - [[k-12]] - [[higher-ed]] - [[cs-education]] - [[generative-ai]] - [[discipline-specific-aied]] ## Connected Articles - [[seung-basham-cognitive-offloading-swld-2026]] — GenAI cognitive offloading for students with learning disabilities - [[special-r1-rl-special-education]] - [[dyslexlens-dyslexic-learners-ai]] - [[llm-question-generation-deaf-hard-of-hearing-2026]] — LLM-powered question generation for Deaf and Hard of Hearing learners - [[neurodivergent-computing-students]] - [[adhd-video-segmentation-computing-education]] - [[genai-minoritized-knowledges-disability]] - [[embodied-string-learning-blindness-low-vision-musicians]] - [[gemini-lualatex-physics-video-transcription-2026]] — Gemini+LuaLaTeX math-accessible physics video transcription - [[dabaghi-ai-dyslexia-education-review-2026]] — AI to help people with dyslexia in education - [[assistive-tech-neurodivergent-higher-ed-review-2026]] — Generative AI, virtual reality, and beyond: A scoping review of digital assistive technologies for neurodivergent students in higher education - [[adapted-stories-social-story-intervention-2026]] — AI-Assisted Social Story Intervention for Special Education: The Design of AdaptED Stories --- ## [Professional Development](https://edtechdev.github.io/aied/concepts/teacher-education/) > **[[educational-development|Professional development]]** — the preparation and ongoing professional development of teachers, spanning pre-service teacher training (initial certification programs) and in-service professional development. In AI-in-education [[research-methods-aied|research]], teacher education has become a central concern because teachers' AI literacy, technological-[[pedagogy|pedagogical]] knowledge, [[ethics|ethical]] fluency, and readiness to integrate AI into instruction determine whether AI adoption in classrooms succeeds. This concept organizes the knowledge base's substantial coverage of how AI reshapes the preparation, knowledge, beliefs, and practice of both prospective and practicing teachers. ## Questions to Consider - Teacher education here spans pre-service training (future teachers in certification programs) and in-service professional development (practicing teachers). Which of the two do you think faces the bigger challenge preparing for AI classrooms, and why? - The page suggests teacher education must prepare teachers not just to use AI tools but to understand, evaluate, and ethically integrate them — a shift that 'redefines what it means to be a teacher.' How do you think that redefinition should change what a credentialing program teaches? - Frameworks like AI-TPACK extend the classic TPACK model by adding an AI and ethics dimension. What do you think an 'AI dimension' of technological-pedagogical-content knowledge should actually contain beyond how to operate a [[conversational-ai|chatbot]]? - A [[meta-analysis-systematic-review|scoping review]] of 55 studies cited on the page finds AI enhances pre-service teachers' instructional design, reflection, and critical thinking. But if AI can do some of this for them, when does its use in teacher training build skill versus substitute for the very judgment they'll need? - If you were redesigning a teacher-education program for the AI era, what would you require every future teacher to experience — and what would you deliberately keep them from outsourcing to AI? ## Introduction Teacher education sits at the intersection of several knowledge base strands: it is a discipline/domain (like [[medical-education]] and [[humanities-education]]), but it also draws on the general concepts of [[teacher-role]], [[tpack]], [[teacher-ai-competency]], and [[ai-literacy]]. In the AI era, teacher education must prepare teachers not only to use AI tools but to understand, evaluate, and ethically integrate them — a shift that redefines what it means to be a teacher. The [[learning-sciences|learning sciences]] supply the knowledge this formation has to carry: where that field studies how people learn and how learning environments should be designed, teacher education is the professional formation in which such evidence has to land as a teacher's own capacity to design, teach and assess. Cross-level evidence sharpens the design priority: [[pedagogy-first-technology-second-teacher-knowledge-2026|a multilevel study of 46 teachers and 2,832 secondary students]] found that pedagogical AI knowledge (TPAIK) — not technical AI knowledge — drove students' perceptions of AI for social good and their intention to learn AI. Teacher education should therefore foreground *how to teach with and about AI* over tool proficiency, consistent with the "pedagogy first, technology second" guideline. ### Pre-service teacher education Pre-service (initial) teacher education prepares future teachers during their certification programs. AI research in this strand includes: - **AI-TPACK and intelligent-TPACK readiness.** Instruments and frameworks measure and build pre-service teachers' readiness to integrate AI, extending the [[tpack]] framework with an AI/ethics dimension.([[conceptualizing-preservice-teachers-ai-readiness-2026]])([[ai-tpack-mathematics-teacher-education-2026]]) - **Applications and benefits.** A scoping review of 55 studies shows AI enhances pre-service teachers' instructional design, subject instruction, practical teaching skills, evaluation, reflective practice, [[critical-thinking|critical thinking]], technology integration, and pedagogical innovation.([[harnessing-ai-preservice-teachers-scoping-2026]]) - **[[educational-robotics|Educational robotics]] and ML.** Initial teacher training embeds coding, robotics, and [[machine-learning]] activities (e.g., micro:bit) to build [[computational-thinking|computational thinking]] in future teachers.([[microbit-robotics-machine-learning-teacher-training-2026]]) - **Authentic assessment and metacognition.** AI-mediated assessment models (e.g., AAIWA) integrate [[authentic-assessment|authentic rubric-based assessment]], condition-responsive [[ai-feedback-quality|AI feedback]], and [[metacognition|metacognitive reflection]] in pre-service programs.([[aaiwa-ai-authentic-assessment-metacognition-2026]]) - **AI-supported inquiry in [[stem-education|science education]].** A quasi-experiment with 48 pre-service science teachers in Türkiye ([[ai-supported-inquiry-photosynthesis-respiration-2026|Aydın]]) integrated problem- and design-based learning into an 8-week AI-supported guided inquiry program on photosynthesis and respiration. It produced significant gains in conceptual understanding of the two [[biology-education|biological]] processes, but *no* significant effect on AI literacy or self-perceived [[computational-thinking|computational thinking]] — a reminder that AI-IBL can deepen subject-matter understanding in teacher candidates without automatically building their AI/CT competencies, which require explicit, targeted design. - **[[generative-ai|Generative AI]] for constructivist instructional design.** The [[sahab-model-genai-constructivist-id-2026|SAHAB model]] is a quasi-experimental professional-development intervention in which generative AI empowers teachers to design [[constructivist]] instruction, with large gains (d = 1.18). It shows GenAI can [[scaffolding|scaffold]] [[learning-design|instructional design]] for teacher candidates and practicing teachers alike — not just deliver content, but support the *design* of student-centered learning activities. - **Smart-classroom training for graduate [[math-education|mathematics]] M.Ed. students.** Zhu, Liang, Mao, and Wang (2026) show how a mathematics M.Ed. course can be enhanced with intelligent educational [[ai-technologies|technologies]] ([[automated-assessment|automated scoring]], personalized recommendations, multi-AI feedback) across pre-, in-, and post-class stages, yielding statistically significant gains in instructional-objective design proficiency. The transferable D-T-E Model (Disciplinary Demand–Technological Empowerment–Evaluation Loop) offers teacher educators a [[discipline-specific-aied|discipline-specific]] framework for integrating smart education into graduate teacher preparation. - **Adoption under a permissive policy is thinner than surveys suggest, and assessment literacy does not transfer.** [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al. (2026)]] surveyed 85 student teachers across three Hong Kong teacher-education courses where generative AI was explicitly permitted in assessment: 62.4% (53) chose not to use it at all, and adopters' use was shallow and corrective (proofreading 43.8%, clarity checks 34.4%, text generation 18.8%, brainstorming 12.5%). Fear of being wrongly accused of plagiarism (41.5% of non-adopters) and a preference for solo work (77.4%) outweighed missing skills (13.2%), and nine of eleven interviewees read the permissive policy as a possible "trap" — evidence that pre-service teachers bring [[academic-integrity|integrity]] anxiety, not just tool readiness, into AI-permitted coursework. - **Classroom AI policy writing as a lens on preservice thinking.** [[nash-preservice-teachers-classroom-ai-policies-2026|Nash & Burriss (2026)]] had 27 preservice secondary English teachers in the capstone year of a secondary English preparation program write the [[educational-policy-ai|AI policies]] they would use in their own 6–12 classrooms, in a state with no state-level AI guidance. 26 of the 27 permitted some AI use, almost always restricted to teacher-specified tasks, times, and places, and only one banned it outright on ethical grounds tied to labor practices and copyright; 22 permitted AI for generating ideas while 22 disallowed or left unclear the use of AI to compose sentences, paragraphs, or papers, and the accompanying reflections read as negotiations between pedagogical commitments and a felt obligation to integrate [[generative-ai|AI]]. Asking preservice teachers to author a policy — rather than debate AI in the abstract — surfaced their philosophies, beliefs, and curricular assumptions, making policy writing a high-value activity for methods coursework. ### In-service professional development In-service professional development supports practicing teachers in integrating AI. AI research here includes: - **Intelligent-TPACK-based PD frameworks.** A research-informed framework aligns the five i-TPACK knowledge domains with four evidence-based PD pathways ([[active-learning|active learning]], models/examples, coaching, feedback/reflection).([[designing-ai-professional-development-itpack-2026]]) - **GenAI-specific technological pedagogical knowledge (TPK).** Teacher educators themselves need GenAI-TPK — pedagogical reasoning, ethical awareness, and AI-augmented instructional design — to prepare teachers.([[teaching-the-teachers-genai-tpk-review-2026]]) - **Human-centered and critical AI literacy.** [[design-based-research|Design-based research]] produces professional-learning curricula that operationalize critical AI literacy through human-centered AI activities, including educator-in-the-loop tasks.([[human-centered-ai-teacher-educators-2026]]) - **Collaborative [[ai-ed-evaluation|evaluation of AI]]-generated content as PD.** [[teachers-collaborative-evaluation-ai-content-2026|Gat, Usher, and Barak (2026)]] report on a workshop in which 60 [[k-12|middle-school]] science teachers rated ChatGPT-generated assessment questions through their disciplinary, pedagogical, and curricular judgment. Collaborative evaluation itself functioned as professional development, helping teachers apply conceptual-precision criteria to AI output and surface the risk that AI content reinforces [[misconceptions]] — positioning teachers as critical evaluators of [[generative-ai|GenAI]] material rather than passive consumers. - **Structured PD for language educators is rare but effective.** A [[li-language-educators-genai-review-2026|systematic review of 23 studies]] (Li et al. 2026) found only three included studies reported structured professional development — an embedded grammar-course module, a government EMI program, and embedded chatbot inquiry — yet all converged on gains in knowledge, confidence, and identity reframing, shifting educators' views of GenAI from "replacement risk" to assistant/augmenter. The review argues PD should pair technical skill-building with practical wisdom, moving from [[ai-literacy|awareness-raising]] and ethics through hands-on tool mastery to co-design of AI-enhanced lessons, and recommends a two-phase "back-end then classroom" implementation strategy. - **Post-qualification programs.** In-service science educators' AI literacy and usage inform the design of AI-related post-qualification programs.([[science-educators-ai-literacy-postqualification-2026]]) - **AI support and guidance in teacher design work.** [[pishtari-teacher-ai-training-learning-design-2026|Pishtari, Gnadlinger & Ley (2026)]] had 13 higher-education teachers design activities across no-AI, AI-chatbot, and AI-plus-training conditions in a half-day program. AI access raised higher-order (Bloom) task attainment and cut perceived cognitive effort, while the subsequent interaction-training session plateaued quality but slightly raised effort (germane vs. extraneous load unresolved). It frames PD for AI-era teaching as needing both pedagogy-grounded design frameworks (Bloom, [[icap-framework|ICAP]]) and structured chatbot-interaction strategies, while cautioning that quality gains may mask [[cognitive-offloading|offloading]] of pedagogical decisions to AI. ### Knowledge, beliefs, and practice A key finding across teacher-education research is the gap between what teachers *articulate* and what they *enact*: teachers often claim operational AI skills but struggle to apply pedagogically meaningful knowledge in practice.([[teachers-ai-knowledge-genai-lesson-planning-2026]]) The distinction runs even deeper when experience is crossed with AI proficiency: [[choi-teacher-ai-interaction-lesson-design-2026|Choi et al. (2026)]] found that technically fluent *novices* still accept AI output passively during lesson design, whereas experienced teachers — even those with lower measured AI proficiency — critically re-prompt and adapt AI suggestions to their students and context. The implication is that in-service PD cannot treat "building AI skill" as sufficient; it must be differentiated by the teacher's experience and proficiency profile — response-evaluation and critical-adaptation scaffolds for novices, hands-on AI skill-building for experienced teachers with weaker AI proficiency. Expert judgment of AI-generated lesson plans marks that critical-adaptation stance as the norm, not the exception: [[karaismailoglu-ai-lesson-plans-science-experts-2026|Karaismailoglu, Surmeli and Yildirim (2026)]] found eleven Turkish science-education specialists rated ChatGPT-4 and Teacher's Buddy plans for a sixth-grade unit as usable drafts — 7 of 11 judging them "applicable by correction," only 3 directly "Applicable" — and their preferences diverged from raw scores, favoring affective and contextual qualities over structural fidelity. Teacher education should therefore build the ability to evaluate, adapt, and localize AI-generated instructional materials as a core professional competence, in both pre-service and in-service programs. Psychological factors also matter — [[self-efficacy]] positively predicts AI-TPACK, while strong traditional teaching beliefs can act as a cognitive barrier.([[ai-tpack-mathematics-teacher-education-2026]]) [[trust]] in AI is shaped by both technical knowledge and ethical perceptions (transparency, fairness, accountability, inclusiveness).([[intelligent-tpack-ethics-teachers-trust-distrust-2026]]) Measuring teachers' AI literacy is itself an active front: because most AI-literacy assessments target students or general users, the Teachers' AI Literacy Scale (TAILS) — grounded in the six-dimension ED-AI framework — was developed and validated specifically for [[language-learning|preservice language teachers]] through EFA and CFA, filling the measurement gap within teacher education. A comparably neglected question is whether the *mode* of preparation changes AI readiness. [[ai-training-science-teacher-tpack-distance-2026|Mnguni et al. (2026)]] surveyed 186 final-year science student teachers in South Africa and found [[self-report-measures|self-reported]] TPACK for AI-integrated teaching higher at a campus-based university (64.0%) than at a [[online-teaching-and-learning|distance education]] university (47.4%), with Pedagogical Knowledge weakest in both, and AI training associated with self-reported TPACK only at the distance institution, where a five-credit short course predicted weaker reported TPACK than no training. Distance preparation therefore needs its own evidence base rather than assuming that training designed for campus programs transfers. ### Simulated instructional practice Beyond content and beliefs, teacher preparation increasingly uses **simulated classrooms** for hands-on practice that scales. [[educasim-cs1-instructional-practice|EducaSim]] uses generative student agents (with personas, course-grounded memories, and an [[llm]]-as-judge speech oracle) to simulate a small-group section for teachers-in-training in a CS1 course supporting ~20,000 students. Deployed as an optional prep tool across 254 sessions (mean ~16 min), it provides low-cost (\$0.05–\$0.10/session) role-play practice with structured post-session feedback (talk-time statistics, LLM-identified instructional behaviors) and self-reflection prompts. This complements the [[simulating-students]] paradigm: simulated learners serve not just evaluation but experiential, high-frequency teacher preparation, especially for massive online courses where live coaching cannot scale. A subject-specific instance is **Student GPT** ([[zhuang-zhang-chatgpt-math-teacher-education-2026|Zhuang & Zhang 2025]]), a custom ChatGPT chatbot that role-plays a middle school student holding common ratio-reasoning misconceptions; preservice secondary mathematics teachers practiced diagnosing and remediating those misconceptions in low-risk, personalized interactions within a methods course, showing how lightweight [[generative-ai|GenAI]] role-play can support practice-based teaching. ## Implications for teacher educators - **Build both AI literacy and critical AI-TPACK.** Pre-service and in-service teachers need readiness to integrate AI — instruments extend TPACK with an AI/ethics dimension ([[ai-tpack-mathematics-teacher-education-2026|AI-TPACK]], [[designing-ai-professional-development-itpack-2026|i-TPACK PD]]); teach ethical reasoning, transparency, and fairness alongside tool use. - **Close the articulate-vs-enact gap.** Teachers often claim operational AI skills but struggle to apply them pedagogically; design PD that moves from knowledge to enacted practice ([[teachers-ai-knowledge-genai-lesson-planning-2026|lesson planning]]). - **Use simulated practice to scale preparation.** [[educasim-cs1-instructional-practice|EducaSim]]-style simulated classrooms give teachers-in-training low-cost, high-frequency practice with feedback — a complement to limited live coaching. - **Address beliefs and trust.** [[self-efficacy]] predicts AI-TPACK while strong traditional-teaching beliefs can be a barrier, and [[trust]] is shaped by transparency/fairness — attend to these psychological factors, not just skills. - **Backtracking from classroom dilemmas.** [[preservice-teachers-responsible-genai-2026|Kohnke et al. (2026)]] interviewed **17 pre-service teachers** at a Hong Kong university and worked backward from anticipated classroom dilemmas to teacher-education design needs. Pre-service teachers raised concerns about **[[academic-integrity|academic integrity]], privacy, and preserving essential human skills** in a GenAI-driven environment, pointing to [[curriculum-design|curriculum]] implications that weave **ethics, [[privacy]], and AI literacy** (with critical-thinking skills) throughout teacher preparation rather than treating responsible use as an add-on. - **AI as cognitive mediator in teacher reflection (2026):** A practice-based Reflective Triangle Model uses AI to mediate between individual teacher reflection and shared professional knowledge in learning communities, addressing the common failure of reflection to transform into collective professional learning ([[reflective-triangle-model-teacher-ai-2026]]). - **Make GenAI [[assessment|assessment design]] part of professional curriculum work.** Student teachers in [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al. (2026)]] showed little intention of redesigning their own future school assessments in response to generative AI, and the [[early-childhood-elementary-ai-education|kindergarten]] and secondary participants largely judged the technology irrelevant to their own teaching context — an underdeveloped dimension of assessment literacy rather than a knowledge gap. Teacher-education programs should therefore embed ethical GenAI use in assessment design as part of professional [[curriculum-design|curriculum]], pairing [[educational-policy-ai|program-level policy]] consistency with concrete exemplars of acceptable and unacceptable use and timely [[feedback]] on students' actual AI use. - **Make classroom policy writing a core teacher-education activity, and treat ambiguity as a policy defect.** [[nash-preservice-teachers-classroom-ai-policies-2026|Nash & Burriss (2026)]] found that asking preservice teachers to write their own classroom AI policies functions as rhetorical world-building about what writing is for — a window into emergent beliefs that abstract discussion does not open. Because AI separates writing-as-product from writing-as-process, programs should help teachers locate where thinking, learning, and value actually live across brainstorming, drafting, revising, and editing rather than locating thinking only in final text, and should support them in translating beliefs into operational, unambiguous rules paired with [[ai-literacy]] instruction: contradictions such as barring AI-generated text while holding students responsible for the AI-generated content they submit leave students unable to comply. ## Connected Concepts - [[teacher-role]] - [[tpack]] - [[teacher-ai-competency]] - [[ai-literacy]] - [[educational-development]] - [[k-12]] - [[ethics]] - [[ai-education]] - [[learning-sciences]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation ## Connected Articles - [[nash-preservice-teachers-classroom-ai-policies-2026]] — Preservice English teachers authoring classroom AI policies: permitted, limited, and banned uses (Nash & Burriss 2026) - [[zou-is-this-a-trap-student-teachers-genai-2026]] — 85 student teachers: shallow GenAI adoption, integrity anxiety, and assessment literacy that doesn't transfer - [[pishtari-teacher-ai-training-learning-design-2026]] — AI chatbot support and training in teachers' learning design - [[generative-ai-k12-teaching-learning-systematic-review-2026]] — Systematic review of generative AI in K-12 teaching and learning (Marzano 2026) - [[pedagogy-first-technology-second-teacher-knowledge-2026]] — Pedagogical AI knowledge as the priority lever for teacher professional learning (Shen et al. 2026) - [[preservice-teachers-responsible-genai-2026]] — Pre-service teachers' responsible GenAI use: curriculum implications (Kohnke et al. 2026) - [[ai-adaptation-gap-higher-education-2026]] — The AI Adaptation Gap in Higher Education - [[harnessing-ai-preservice-teachers-scoping-2026]] — Scoping review of AI in preservice teacher development - [[designing-ai-professional-development-itpack-2026]] — Intelligent-TPACK-based professional development framework - [[human-centered-ai-teacher-educators-2026]] — Professional learning for critical AI literacy in teacher educators - [[teaching-the-teachers-genai-tpk-review-2026]] — GenAI-specific TPK in teacher education - [[teachers-ai-knowledge-genai-lesson-planning-2026]] — Teachers' AI knowledge in GenAI lesson planning - [[choi-teacher-ai-interaction-lesson-design-2026]] — Teacher-AI interaction patterns across teaching experience and AI proficiency in lesson design (Choi et al. 2026) - [[conceptualizing-preservice-teachers-ai-readiness-2026]] — Pre-service intelligent-TPACK readiness - [[ai-tpack-mathematics-teacher-education-2026]] — AI-TPACK in mathematics teacher education - [[intelligent-tpack-ethics-teachers-trust-distrust-2026]] — Ethics domain and in-service teachers' trust - [[microbit-robotics-machine-learning-teacher-training-2026]] — Micro:bit robotics in initial teacher training - [[aaiwa-ai-authentic-assessment-metacognition-2026]] — AI-mediated authentic assessment in pre-service education - [[science-educators-ai-literacy-postqualification-2026]] — Science educators' AI literacy and post-qualification programs - [[ithaka-sr-ai-skills-college-graduates-2026]] — Instructors report institutional AI-skills consensus and assessment gaps - [[young-people-learning-generative-ai-rapid-review-2026]] — Teachers essential for relational/higher-order work in hybrid arrangements - [[ai-science-chemistry-education-systematic-review-2025]] — Systematic review of AI in science/chemistry education - [[context-based-ai-secondary-chemistry-2026]] — Context-based 7E + AI instruction in secondary chemistry - [[instructor-ai-roles-chatgpt-formative-assessment-2026]] — Instructor and AI roles in ChatGPT-enhanced formative assessment - [[educasim-cs1-instructional-practice]] — EducaSim: interactive simulacra for CS1 instructional practice - [[chen-preservice-teachers-chatgpt-lpa-2026]] — Pre-service teacher ChatGPT acceptance profiles - [[motivation-shape-future-education-ai-switzerland-china]] — Motivation to shape the future of education with AI - [[ai-supported-inquiry-photosynthesis-respiration-2026]] — AI-supported guided inquiry in science teacher education - [[reflective-triangle-model-teacher-ai-2026]] — Reflective Triangle Model: AI as cognitive mediator - [[pre-service-science-teachers-ai-perceptions-2026]] — Ghanaian pre-service science teachers' AI perceptions (UTAUT/TPB) - [[sahab-model-genai-constructivist-id-2026]] — SAHAB model: GenAI constructivist instructional design - [[caruana-pre-university-ai-education-slr-2026]] — Preparing learners and teachers for an AI-driven future: SLR of pre-university AI education (Caruana et al. 2026) - [[riandi-teacher-ai-green-energy-education-2026]] — Teacher involvement in AI integration for green energy education (Riandi et al. 2026) - [[utility-value-intervention-teach-responsibly-genai-2026]] — Utility-value intervention effects in learning to teach responsibly with GenAI (Boos, Eder & Lachner 2026) - [[preservice-teacher-agency-genai-design-learning-2026]] — Pre-service teacher agency during GenAI interactions in design for learning (Krushinskaia, Elen & Raes 2026) - [[longitudinal-ai-usage-ethics-policy-teacher-education-2026]] — Longitudinal GenAI usage, ethics, and policy in teacher education (Parker et al. 2026) - [[instructional-design-proficiency-masters-math-2026]] — Smart-classroom model and D-T-E loop improving M.Ed. instructional design proficiency in mathematics (Zhu et al. 2026) - [[language-teachers-ai-literacy-edai-2026]] — Teachers' AI Literacy Scale (TAILS) psychometric study (ED-AI framework) - [[li-language-educators-genai-review-2026]] — Language educators' practices and development with GenAI - [[zhuang-zhang-chatgpt-math-teacher-education-2026]] - [[karaismailoglu-ai-lesson-plans-science-experts-2026]] - [[teachers-collaborative-evaluation-ai-content-2026]] — Teachers' collaborative evaluation of AI-generated content as professional development (Gat, Usher & Barak 2026) - [[ai-training-science-teacher-tpack-distance-2026]] — Campus versus distance comparison of 186 science student teachers' AI-related TPACK and training levels - [[obyrne-co-constructing-ai-boundaries-agency-judgment-2026]] — traces pre-service teachers building AI boundaries in a semester-long literacy ethnography --- ## [Levels of Education](https://edtechdev.github.io/aied/concepts/education-levels/) > **Levels of education** — the bands that organize this knowledge base's `level` metadata: **preschool**, **primary education**, **middle school**, **secondary**, **k 12**, **higher ed**, **undergraduate**, **graduate**, **adult learning**, **special education** and **[[teacher-role|teacher]] education**. This page is the umbrella for that field rather than a duplicate of any single band page: it explains what changes as you move across the bands, why the school/university break matters more than the subject being taught, and where the AI evidence is dense and where it is thin. ## Questions to Consider - A graduate seminar and a second-grade mathematics lesson are both "education," yet a tool that helps one can harm the other. What about the band — not the subject — changes what good AI use looks like? - One trial with 132 second-graders found an adaptive mathematics tutor no better than a fixed sequence. Would you expect the same null for a university student, and what learner capacity does adaptation silently assume? - 'K-12' and 'higher education' each compress several distinct settings into one label. Which comparisons do those labels hide? - If a band's evidence base is thin, is the honest response to say nothing, to borrow from an adjacent band, or to design a study? ## Introduction The `level` metadata field names the band of education a source is about — bands, not ages: **preschool**, **primary education**, **middle school**, **secondary**, **k 12**, **higher ed**, **undergraduate**, **graduate**, **adult learning**, **special education**, **teacher education**. Some name a stage of schooling, two name the very different halves of university study, two name a [[learners|learner population]] rather than a stage, and one names the people who teach. A source can carry several bands at once. This page is the umbrella for that field; the band pages do the deep work and should be read with it: [[k-12]], [[early-childhood-elementary-ai-education]], [[higher-ed]], [[adult-learning]], [[special-education]], [[teacher-education]] and [[vocational-education]]. ## What distinguishes one band from another Bands differ along several axes at once, and AI research usually varies only one of them. - **Capacity for self-regulation.** Two versions of the same second-grade mathematics tutor — identical content, interface, feedback and spoken hints, differing only in whether task selection followed a [[knowledge-tracing|Bayesian Knowledge Tracing]] mastery estimate — produced no posttest difference among 132 seven-year-olds (F(1, 124) = 0.32, p = .574). The authors' first explanation is developmental: [[adaptive-learning|adaptive systems]] presuppose learners who can engage feedback, regulate effort and stay focused, and young children have limited [[self-regulated-learning|self-regulation]] ([[adaptive-intelligent-tutoring-primary-mathematics-2026]]). - **Who mediates the interaction.** In the school bands an adult stands between learner and tool. A survey of 270 preschool teachers found adoption intention driven by [[technology-acceptance-model|perceived usefulness]], ease of use, AI [[self-efficacy]] and [[anxiety-and-stress|AI anxiety]] — and the study deliberately excluded AI directed at children ([[preschool-teachers-ai-behavioral-intention-2026]]). - **What the assessment is for.** Secondary assessment feeds external stakes; university assessment is coursework and credentials; graduate assessment is the making of a researcher. - **[[prior-knowledge|Prior knowledge]].** Prior knowledge dominated that trial's posttest performance (F(1, 124) = 206.99, p < .001, η²p = .63), so a band label correlates with, but does not equal, a knowledge level. ## The school years and the university years The most consequential divide is the break between school and university, and the two commonest labels each hide it or cross it. Inside the school years the measured outcome changes sharply by band. In primary it is subject learning: 97 Chinese third-graders using [[conversational-ai|GenAI chatbots]] in [[science-education|science inquiry]] posed better problems than a search-engine control (t = 2.47, p = 0.015) ([[dai-chatbots-problem-posing-primary-2026]]). In secondary it often becomes attitude rather than achievement: a survey of 508 Taiwanese junior high students found experiential appeal working through enjoyment to shape intention to use [[generative-ai|ChatGPT]] for lyric learning (XM → PEOU β = 0.630; PE → ATU β = 0.369), with 81.5% on the free tier — an access-[[equity-in-ai-education|equity]] fact disguised as a technology-acceptance finding ([[chatgpt-music-education-junior-high-2026]]). Secondary also holds the corpus's largest learning warning: across 26,811 Chinese students in grades 7–12, homework scores rose 18% and completion time fell 30% while [[summative-assessment|closed-book exam]] scores dropped about 20% within six months, concentrated among the ~81% whose behavior indicated homework outsourcing ([[stromberg-generative-ai-learning-penalty-secondary-2026]]). The university years split again. Undergraduate study is coursework under external judgment: among undergraduate writers at a minority-serving R1 university, [[ai-literacy]] predicted *which type* of [[llm]] reliance a student occupied rather than how much they used it ([[llm-reliance-types-undergrad]]). Graduate study is research training, where the outcome shifts from performance to formation. Among 420 astronomy doctoral and postdoctoral researchers, AI dependence was negatively associated with research [[agency|autonomy]] (r = −.355) and [[self-efficacy]] (r = −.321), and the indirect path to innovative behavior ran mostly through autonomy (−.115) rather than self-efficacy (−.069); supervisory support weakened the negative link to autonomy (B = .077, p = .020) ([[ai-mediated-research-agency-formation-2026]]). A doctoral student is judged on the judgment that AI dependence appears to erode; an undergraduate is not. Lumping both under 'higher education' hides that. ## Where the evidence is concentrated, and where it is thin The corpus is unevenly populated, and the level field makes the imbalance visible. **Dense.** Higher education is the most-covered band: the pages consulted here include an eight-week platform trial with 60 engineering students ([[ai-assisted-seminar-learning-information-literacy-2026]]), a survey of 395 education managers ([[ai-adoption-readiness-ukraine-education-managers-2026]]), a meta-synthesis of 18 African higher-education studies ([[data-privacy-ai-african-higher-education-2026]]) and a 420-researcher study of doctoral training ([[ai-mediated-research-agency-formation-2026]]). Primary and secondary also carry outcome evidence. **Thin, and in places absent.** The pages consulted here report no child-outcome study at the preschool band at all; the nearest evidence is a survey of teachers' intentions that excluded child-facing tools ([[preschool-teachers-ai-behavioral-intention-2026]]). Middle school appears mainly as a *proposed* design and longitudinal study rather than reported outcomes ([[ai-lms-middle-school-longitudinal]]). The vocational band's strongest evidence is 63 [[self-report-measures|self-report]] responses from one course, which its authors say requires replication ([[ai-ive-pbl-vocational-design-creativity-2026]]). The graduate evidence is cross-sectional, with a reverse-path model fitting slightly better than the developmental one, so the direction from dependence to reduced autonomy is inferred rather than confirmed ([[ai-mediated-research-agency-formation-2026]]). Nothing consulted here reports AI outcome evidence for **adult learning** or **special education** as bands; those have their own coverage on [[adult-learning]] and [[special-education]], and this page does not generalize from school-age findings to fill the gap. One pattern holds at every level surveyed: the human layer absorbs the hardest cases. A human expert stayed at the credibility decision point for engineering students even when algorithmic recommendation reached F1 = 0.64 ([[ai-assisted-seminar-learning-information-literacy-2026]]); supervisory support was the one condition that weakened AI dependence's negative links in doctoral training ([[ai-mediated-research-agency-formation-2026]]); and it is preschool teachers, not children, who are the adopters ([[preschool-teachers-ai-behavioral-intention-2026]]). ## What level-appropriate design actually changes - **Autonomy and support.** With young learners, adapt the *level of support* — more [[scaffolding]], guidance or hints — rather than task difficulty; difficulty adaptation bought nothing in early primary, while mastery gating may have held adaptive students back ([[adaptive-intelligent-tutoring-primary-mathematics-2026]]). - **Reading level.** An age-tailored chatbot used prompt variables for ages 7–9, 9–11 and 12–14, and 63 children from first through eighth grade treated it as a credible information source ([[vahedian-children-attitudes-ai-chatbot-2026]]). - **Mediation.** [[parents-and-families|Parents]] and teachers mediate in the school years, [[librarians]] and supervisors in the university years — and the university counterpart is expertise rather than guardianship: the engineering platform kept a human where the algorithm was least confident ([[ai-assisted-seminar-learning-information-literacy-2026]]). - **Assessment stakes.** The middle-school LMS design gates AI by activity, keeping bounded hints in practice mode and switching AI off on graded items ([[ai-lms-middle-school-longitudinal]]) — a precaution the secondary learning-penalty evidence makes concrete ([[stromberg-generative-ai-learning-penalty-secondary-2026]]). - **Data-protection duties for minors.** Obligations scale with age: for minors the corpus's proposals are structural — data minimization, age-appropriate response constraints, role-based access control, auditable logs ([[ai-lms-middle-school-longitudinal]]) — and children's digital-safety awareness cannot be assumed, since some were willing to confide secrets in a chatbot ([[vahedian-children-attitudes-ai-chatbot-2026]]). For adults the duties shift toward consent, control and cross-border data flow ([[data-privacy-ai-african-higher-education-2026]]). - **Governance that fits the band.** Readiness is layer-specific: 395 Ukrainian managers scored personal readiness 0.68 points above system readiness (d = 0.73), cited regulatory absence most often (58.5%), and rated [[personalized-learning|personalization]] — the benefit vendors promise most — lowest of all applications (29.4%) ([[ai-adoption-readiness-ukraine-education-managers-2026]]). ## Implications for AI in education - **Read the level field before the topic field.** A finding from one band is a hypothesis for another, not a transferable result. - **For the youngest bands, adapt support, not task difficulty** ([[adaptive-intelligent-tutoring-primary-mathematics-2026]]). - **Don't port a tool across the school/university break without re-specifying who holds epistemic authority.** Dependence that lowers autonomy in doctoral training is a research-formation risk ([[ai-mediated-research-agency-formation-2026]]). - **Gate AI by activity, not by enthusiasm**, and scale privacy duties with age ([[ai-lms-middle-school-longitudinal]], [[data-privacy-ai-african-higher-education-2026]]). - **Design for the mediating adult of that band** — parent, teacher, librarian or supervisor — and measure their readiness separately from the institution's ([[ai-adoption-readiness-ukraine-education-managers-2026]], [[ai-assisted-seminar-learning-information-literacy-2026]]). - **Say plainly when a band has no evidence**, and **don't mistake intention for achievement**: several level-specific studies report attitude or [[self-report-measures|self-report]] outcomes rather than learning. ## Connected Concepts - [[k-12]] - [[early-childhood-elementary-ai-education]] - [[higher-ed]] - [[adult-learning]] - [[special-education]] - [[teacher-education]] - [[vocational-education]] - [[learners]] - [[parents-and-families]] - [[differential-effects-across-learner-groups]] ## Connected Articles - [[adaptive-intelligent-tutoring-primary-mathematics-2026]] — Adaptive versus non-adaptive tutoring in second-grade mathematics (Sibley et al. 2026) - [[dai-chatbots-problem-posing-primary-2026]] — GenAI chatbots and problem posing with third-graders in primary science - [[chatgpt-music-education-junior-high-2026]] — Junior high students' attitudes toward ChatGPT for lyric learning (Weng & Chiang 2026) - [[ai-lms-middle-school-longitudinal]] — AI-integrated LMS designed for middle school, with a proposed longitudinal study - [[stromberg-generative-ai-learning-penalty-secondary-2026]] — The generative AI learning penalty in Chinese secondary education (Strömberg et al. 2026) - [[llm-reliance-types-undergrad]] — Four types of LLM reliance among undergraduate writers (Hossain 2026) - [[ai-assisted-seminar-learning-information-literacy-2026]] — AI-assisted seminar platform with embedded librarian support for engineering students - [[ai-mediated-research-agency-formation-2026]] — AI dependence, autonomy and innovation in doctoral and postdoctoral training (Han & Liu 2026) - [[data-privacy-ai-african-higher-education-2026]] — Stakeholder views on data privacy in African higher education (Duncan 2026) - [[ai-adoption-readiness-ukraine-education-managers-2026]] — Education managers' AI adoption readiness across Ukraine (Kremen et al. 2026) - [[ai-ive-pbl-vocational-design-creativity-2026]] — AI-enabled immersive PBL for vocational design students (Jin et al. 2027) - [[preschool-teachers-ai-behavioral-intention-2026]] — Preschool teachers' intention to use AI in early childhood settings - [[vahedian-children-attitudes-ai-chatbot-2026]] — Children's attitudes toward an age-tailored AI chatbot --- ## [Assessment](https://edtechdev.github.io/aied/concepts/assessment/) > **Assessment** — the process of gathering and interpreting evidence about what learners know and can do, and the methods used to evaluate learning. [[ai-education]] has fundamentally reshaped assessment: it powers [[automated-assessment|automated grading and scoring]], generates and adapts assessment items, and raises deep questions about what assessments actually measure when students can use AI. Assessment is the umbrella concept that organizes the knowledge base's coverage of [[formative-assessment]], [[automated-assessment|automated grading]], [[assessment-validity|validity]], and [[educational-measurement]]. ## Questions to Consider - Assessment here is defined as gathering and interpreting evidence about what learners know and can do. Before reading, how would you finish the sentence 'an assessment is valid if…' — and what would change about that answer if every student could secretly use AI to produce their work? - The page claims that how a student uses GenAI (evaluative integration versus uncritical shortcut uptake) predicted performance, while how often they used it predicted nothing. What does this suggest about a common instinct to regulate AI by limiting its use? - A learner may deliver professional-standard work they cannot reproduce without the tool — severing the inference between performance and underlying ability. Have you encountered a 'competency' being awarded for work that the student couldn't actually do unaided? What did that reveal? - The constructive question the page offers is not 'how do we stop students using AI?' but 'how do we enable thoughtful use in contexts that mirror their future work?' What would assessment look like in your field if that were the goal? - The DRIVE framework suggests assessing the quality of a student's [[student-engagement|engagement]] with GenAI — not just the artifact — by looking at whether they steer prompts strategically and integrate their own ideas. What would you look at to tell a deep, reflective [[student-ai-interaction|AI interaction]] from surface consumption? - AI-mediated assessment is diversifying into [[oral-assessment|oral exams]], portfolios, and conversational formats that reduce anxiety and feel professionally relevant. Which assessment format from your own experience do you think is most 'AI-resistant' — and is resistance the same thing as educational value? ## Introduction Assessment is central to AI in education for two reasons. First, AI itself is used to assess students — grading essays, code, short answers, and exams at scale. Second, AI in the classroom changes what assessments can validly measure, since students may use [[generative-ai]] to produce work. The field therefore spans both the *tools* that automate assessment and the *validity and integrity* questions that AI raises. ## How AI is used in assessment - **Automated assessment:** [[automated-assessment|AI-based assessment]] spans multiple modalities — multiple-choice, short answer, essay, code, and performance-based evaluation — through [[automated-assessment|automated grading]], [[automated-essay-scoring]], and [[automated-question-generation]]. - **Formative assessment:** [[formative-assessment|AI systems]] generate, validate, and adapt formative assessment items at scale, informing ongoing instruction rather than only [[summative-assessment|summative]] evaluation. - **Learning analytics and measurement:** [[learning-analytics]] and [[educational-measurement]] connect assessment data to learning processes, using [[item-response-theory]], [[knowledge-tracing]], and [[student-modeling]] to interpret performance. [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] demonstrate LLM-based item-difficulty estimation as a measurement input: across 5,170 K-5 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but varied by grade, while a feature-based approach (LLM-extracted features into tree-based models) reached correlations up to r = 0.87. The study offers a practical seven-step workflow for testing professionals, while cautioning that generalizability beyond K-5 math and reading is unclear. - **Feedback loops:** AI assessment increasingly feeds into [[feedback|feedback loops]] that close the cycle from assessment to learning. - **E-portfolio assessment:** [[eportfolio|e-portfolios]] assemble student work and reflections over time as a process-based, AI-robust assessment form. Generative AI can assist the portfolio *process* — generating feedback, [[scaffolding]] reflection, and (with appropriate rubric design) supporting evaluation — while the portfolio's emphasis on reasoning traces and drafts resists AI fabrication. [[ni-lam-multiliteracies-ai-portfolio-2026|AI-assisted portfolio assessment]] and [[sutama-chatgpt-eportfolio-speaking-2026|ChatGPT + e-portfolio for EFL speaking]] show AI can enhance both the portfolio experience and learner [[feedback-literacy]]. - **Open-ended grading reliability is model-dependent:** [[pecuchova-automated-grading-open-ended-genai-2026|Pecuchova, Benko & Drlik (2025)]] benchmarked eleven GenAI and sentence-embedding models against two expert graders on 1,885 open-ended software-engineering responses: only GPTo1 reached almost-perfect agreement (Fleiss' Kappa 0.82), while reference-based models penalized correct-but-divergently-phrased answers. Because GPTo1 was the only model deemed deployable without oversight but carries proprietary API costs, the authors recommend hybrid strategies pairing advanced models with affordable options or [[human-in-the-loop-ai|human oversight]] in resource-constrained settings. A PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 empirical studies (2023–2025) corroborates this conditional-reliability picture across the whole grading field: LLMs match human raters on closed-ended and short-answer tasks but cannot fully replace human judgment on complex, open-ended, or subjective work, and no uniform grading bias emerged — models were sometimes more lenient, sometimes stricter, and often avoided extreme scores ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). Iterative rubric refinement can push LLM open-ended scoring toward near-human reliability in high-stakes [[medical-education|medical]] settings: [[olvet-genai-scoring-open-ended-medical-2026|Olvet et al. (2026)]] found that once faculty repeatedly revised analytic and holistic rubrics on error-pattern analysis, GPT-4 reached substantial-to-almost-perfect agreement with faculty graders on three of four pre-clerkship questions (weighted kappa up to 0.94), while the residual holistic-rubric item stayed at only moderate agreement (κw = 0.54) — showing both the payoff of human-in-the-loop rubric engineering and its limits on synthetic, holistic tasks. - **AI-supported oral and performance assessment.** A [[vocational-education|TVET]] design study moves assessment away from text-heavy formats toward interactive [[oral-assessment]] supported by an LLM: across four cohorts, 21 of 33 learners rated the voice task realistic and none disagreed that speaking in real time reflected competence better than a written portfolio, while word counts for identical questions varied five- to eight-fold between cohorts and nine learners answered in two to thirteen words and were all correct. The system ran offline on a single laptop serving up to 12 simultaneous learners with Mistral 7B and faster-whisper, deleted recordings after 90 days, and left scoring judgments with assessors — an instance of [[human-in-the-loop-ai]] assessment design for workplace-facing qualifications. ([[ai-supported-oral-assessment-tvet-2026]]) ## Validity and measurement challenges AI raises fundamental [[assessment-validity|validity]] questions: do AI-graded assessments measure student learning or [[prompt-engineering|AI-prompting skill]]? Does student use of AI invalidate traditional assessments? Key challenges include: - **Construct validity:** [[competency-based-education-genai-production-2026|Research on competency-based education]] shows generative AI has severed the inference between performance and underlying ability — a learner may deliver professional-standard work they cannot reproduce without the tool. This motivates reconceptualizing what competencies are assessed. - **Coauthorship and integrity:** [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene|Coauthorship integrity]] proposes a new source of validity evidence violated when students submit AI-generated content they do not understand, and explores conversational "AI Vivas" as a response. - **Psychometric quality:** [[psychometrically-aware-ai]] and [[automated-assessment|confidence-aware assessment]] work to keep AI scoring reliable, unbiased, and interpretable. - **Evaluation of AI assessors:** [[ai-ed-evaluation]] provides the methods and [[benchmark|benchmarks]] for determining whether automated assessors actually work — on reliability, [[pedagogy]], and [[equity-in-ai-education]] — rather than headline accuracy alone. - **Quality-dependent reliability of AI and peer grading (2026):** Comparing ChatGPT, peer, and [[teacher-role|instructor]] grading of the same [[higher-ed|undergraduate]] [[group-work|group projects]], [[usher-faraon-who-grades-best-2026|Usher & Faraon (2026)]] showed grading alignment with the instructor is *conditional on the quality of student work*: ChatGPT inflated low-quality submissions most and aligned better with the instructor on high-quality work, while peers aligned best on weaker work and under-graded strong projects. The finding challenges the binary "reliable vs unreliable" framing, pointing toward conditional-reliability models in which alternative assessors are matched to task and performance level. - **Human grading is itself value-laden (2026):** [[luo-dawson-value-judgments-grading-2026|Luo & Dawson (2026)]] show that even human grading of GenAI-assisted work is not a neutral, criteria-based act. Scenario-based interviews with 33 university teachers revealed grading decisions driven by person-oriented (honesty, diligence), capability-oriented (independence, GenAI skill, disciplinary mastery), relation-oriented (trust), and justice-oriented (fairness, beneficence) values — extending beyond the assignment to teachers' conjecture about the student. The study reframes the assessment question from "is GenAI use cheating?" to "how do teachers' value judgments shape grades, and are those values relevant to the outcomes being assessed?" — foregrounding [[assessment-validity|validity]] and "two-way [[explainable-ai|transparency]]" about how GenAI use will affect grades. ## Integrity and the debate over detection AI in assessment has intensified the [[academic-integrity|integrity]] conversation. One strand focuses on [[ai-detection|detecting AI-generated text]], while a growing body of [[research-methods-aied|research]] argues that detection is a limited, situational tool — not a strategy of first resort. [[beyond-detection-authentic-assessment-ai-2025|Beyond Detection]] and [[responsible-assessment-ai-era-stanford-2026|Responsible Assessment]] argue that authenticity cannot be policed into existence; it must be redesigned, positioning AI as a declared collaborator and prioritizing [[authentic-assessment|authentic, process-based assessment]] over surveillance. **[[walton-bearman-assessment-judgment-2025|Walton et al. (2025)]]** ground this in evidence of **how students actually judge** their way through assessment with GenAI: scroll-back interviews with 26 students revealed a spectrum of six judgment events — from critically evaluating AI knowledge and learning through AI's limitations, to adopting ideas uncritically and misjudging AI contributions as their own. **[[stamatoulis-genai-use-patterns-2026|Stamatoulis et al. (2026)]]** add a [[quantitative-research|quantitative]] counterpart: across 157 students, *how* GenAI is used (evaluative integration to support understanding vs. low-verification shortcut uptake) predicted performance, while simple usage **frequency predicted neither** performance nor academic [[self-efficacy]]. Together these studies reframe the assessment question from *whether* students use AI to *how they judge and pattern that use*. ## Assessment redesign in the AI era The constructive question in the knowledge base's assessment literature is not "how do we prevent students from using AI?" but "how do we enable them to use it thoughtfully in contexts that mirror their future work?" This reframes assessment around: - **Authentic and process-based tasks** that make AI use visible and assessed. [[authentic-assessment]] — examining student performance on worthy, realistic tasks — is the leading response to AI's challenge: any task an [[llm]] can credibly simulate loses its validity, so authenticity must be redesigned around real-time [[collaborative-learning|collaboration]], digital and social contribution, and individual meaning-making. This connects to [[zhan-boud-du-authentic-assessment-scoping-review-2025|design frameworks for authentic assessment]], [[authentic-products-authenticated-processes-2026|authentic products and authenticated processes]], and [[tool-invariant-framework-agentic-ai|tool-invariant assessment of process]]. - **Responsible assessment design** grounded in validity evidence ([[responsible-assessment-ai-era-stanford-2026]]) - **Task-level AI permissions derived from assessment targets:** [[mccorkle-aligned-genai-course-policy-2025|McCorkle (2025)]] shows the assessment-design work that precedes an [[educational-policy-ai|AI policy]] — inventorying every task in a project, specifying what is being assessed and against which objective, and permitting or prohibiting AI per task on that basis (brainstorming and image curation allowed; composing learning objectives and slide design not). The same alignment exercise doubles as a check on the inference the assessment supports, because it forces the instructor to name the performance that a grade is meant to warrant ([[assessment-validity]]). - **Coauthorship and declaration** as part of the assessment contract - **Production as a competency** — evaluating learners' ability to direct tools and produce professional-standard work ([[competency-based-education-genai-production-2026]]) - **Assessing the interaction process, not just the artifact** — the [[assessing-student-drive-framework-2025|DRIVE framework]] (Directive Reasoning Interaction + Visible Expertise) treats the quality of a student's *engagement with GenAI* as the assessed construct. It distinguishes surface consumption from deep, reflective interaction by looking at whether students steer prompts strategically (DRI) and integrate and develop their own disciplinary ideas through the exchange (VE), grounding process-focused criteria in theories of [[self-directed-learning]] and cognitive engagement along the lines of the [[icap-framework|ICAP]] hierarchy. This makes DRIVE an example of *AI-mediated authentic assessment* — a rubric for evaluating how learners partner with GenAI rather than a detection tool. A proposal in this literature pushes past redesign-within-the-current-frame. [[ai-agents-joyful-assessment-third-space-2026|El Khoury and Ma (2026)]] argue that reform organized around preventing misconduct or detecting AI use narrows the educational imagination to control and compliance, and propose **joyful assessment** instead: assessment that is safe, emotionally responsive, empowering and supportive of student [[agency|student agency]], with safety as the load-bearing condition because without it emotional attunement becomes performance, empowerment becomes pressure and agency becomes risk. Their framing inverts the detection agenda — integrity becomes a consequence of designing assessment students want to engage in rather than its starting point — and they position instructor-built [[agentic-ai|AI agents]] (custom GPTs, Gems, Copilot Studio agents) as a low-stakes rehearsal space where students practice before judgment, with the claim that the AI organizes evidence while the instructor interprets it. ## Implications for AI in education - **Assessment and learning are inseparable:** good AI assessment should support learning ([[feedback|formative feedback]]) as much as it evaluates it. - **Validity must be reconceptualised:** when AI can produce student work, assessments must measure processes, judgment, and authentic production, not just outputs. - **Automation must be evaluated rigorously:** automated assessors need psychometric and [[bias-mitigation|fairness]] evaluation, not just accuracy claims. - **Integrity shifts from detection to design:** the most robust response to AI in assessment is designing tasks where AI use is expected, declared, and scrutinized. - **Innovative practices can address multiple problems at once.** [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] argue that a set of five mutually reinforcing practices — [[authentic-assessment]], self- and [[peer-assessment]], reassessment, [[mastery-learning|standards-based grading]], and ungrading — aligned with the "assessment for learning" paradigm can counter the harms of traditional, summative-heavy, norm-referenced grading (superficial learning, eroded [[motivation|intrinsic motivation]], stress and anxiety, inequity, and compromised [[academic-integrity|integrity]]) in the generative AI era. They find uptake is uneven across fields — self- and peer-assessment dominate while the grading-focused innovations remain marginal — and urge educators to engage more actively and reciprocally with assessment and grading innovation, backed by incremental experimentation and [[governance|institutional]] support. - **AI-mediated assessment is diversifying.** [[aivaluate-anxiety-assessment-2026|AIvaluate]] shows an LLM-augmented [[conversational-ai|conversational agent]] reduced student anxiety during performance-based assessments; [[asynchronous-oral-assessment-2026|Pentland (2026)]] finds asynchronous oral assessments offered higher engagement and were perceived as professionally relevant; [[graph-its-adaptive-algorithms-2026|graph-based ITS]] uses adaptive knowledge-state tracking to inform assessment. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[formative-assessment]] — Formative assessment: AI-generated, validated, adaptive items at scale - [[automated-assessment]] — Automated grading and scoring across assessment modalities - [[authentic-assessment]] — Process-based, AI-robust authentic tasks - [[assessment-validity]] — Validity of assessments under generative AI - [[educational-measurement]] — Measurement theory underpinning AI assessment - [[automated-essay-scoring]] — Automated essay scoring - [[automated-question-generation]] — Automated question generation - [[oral-assessment]] — Oral assessment: live and recorded formats that resist AI substitution - [[summative-assessment]] — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams) - [[item-response-theory]] — IRT for interpreting AI-era assessment responses - [[psychometrically-aware-ai]] — Psychometrically aware AI scoring - [[learning-analytics]] — Analytics connecting assessment data to learning - [[feedback]] — Feedback loops closing the assessment-to-learning cycle - [[feedback-literacy]] — Learner feedback literacy - [[academic-integrity]] — Integrity and the debate over AI detection - [[ai-detection]] — Detecting AI-generated text - [[ai-ed-evaluation]] — Methods and benchmarks for evaluating automated assessors - [[eportfolio]] — Process-based e-portfolio assessment - [[speech-and-voice-technologies]] - [[peer-assessment]] ## Connected Articles - [[ai-agents-joyful-assessment-third-space-2026]] — AI agents, joyful assessment, and third space - [[mccorkle-aligned-genai-course-policy-2025]] — Task-level AI permissions derived from what is assessed (McCorkle 2025) - [[usher-faraon-who-grades-best-2026]] — Comparing ChatGPT, peer, and instructor grading across project quality levels (Usher & Faraon 2026) - [[responsible-assessment-ai-era-stanford-2026]] — Responsible assessment in the AI era - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond detection: authentic assessment - [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene]] — Coauthorship integrity and assessment validity - [[competency-based-education-genai-production-2026]] — Competency-based education after generative AI - [[genai-assessment-governance]] — Evidence-centered governance of generative AI in assessment - [[cong-confidence-asag-2026]] — Automatic short-answer grading - [[ssaho-ai-academic-integrity-review-2025]] — AI integrity review: detection must pair with assessment redesign - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant assessment - [[bassett-ai-detectors-education-2026]] — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026) - [[aivaluate-anxiety-assessment-2026]] — AIvaluate: LLM-Augmented Assessment of Student Anxiety (2026) - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[asynchronous-oral-assessment-2026]] — Asynchronous Oral Assessments in the AI Era (Pentland 2026) - [[assessing-student-drive-framework-2025]] — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise) - [[xiong-ai-educational-measurement-review-2026]] — AI reshaping assessment practice - [[walton-bearman-assessment-judgment-2025]] — Judgment in students' work with GenAI on assessment tasks (26 students, scroll-back) - [[stamatoulis-genai-use-patterns-2026]] — Patterns of GenAI use (evaluative integration vs low-verification uptake) and outcomes - [[luo-dawson-value-judgments-grading-2026]] — Value judgments in grading GenAI-assisted work: honesty, trust, validity, and two-way transparency (Luo & Dawson 2026) - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[pecuchova-automated-grading-open-ended-genai-2026]] - [[mesny-innovative-assessment-grading-management-2026]] - [[olvet-genai-scoring-open-ended-medical-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] --- ## [Evaluative Judgment](https://edtechdev.github.io/aied/concepts/evaluative-judgment/) > **Evaluative judgment** — the capacity to make sound judgments about the quality of one's own work and the work of others, against criteria one can reason about rather than recite. In the [[generative-ai|generative AI]] era it has moved from a desirable graduate attribute to a load-bearing capability: when a tool can produce plausible finished work, the capability that distinguishes a competent learner is the ability to appraise that output — to judge what is good, what is wrong, what to accept, what to reject, and why. This makes evaluative judgment a legitimate object of [[assessment]] in its own right, and the practical pivot for moving institutions from [[ai-detection|detection]] toward [[assessment-validity|validity-centered]] and [[authentic-assessment|authentic]] assessment design. ## Questions to Consider - Think of something you can judge well. How did you learn that judgment — explicitly taught, or absorbed by making and comparing — and what does that imply for how it should be assessed? - A rubric can be applied mechanically or understood. If a student follows criteria correctly but cannot say why one piece of work beats another, which capability has the assessment actually measured? - When a [[llm|language model]] drafts a plausible solution, the learner's remaining work is largely judgment: what to keep, what to fix, what to discard. Is that a diminished task, or a more demanding one than producing the draft? - A student can accept an AI suggestion because it reads well, or because it survives their own scrutiny. From the outside the submission may look identical — what would let an assessor see the difference? - A multisite experiment found that direct [[ai-feedback-quality|AI feedback]] produced the largest revision gains but risked passive outsourcing of judgment. What does that trade-off suggest about how you would design feedback to build rather than bypass judgment? - Peer and AI review can train calibration, but both can become performative. How would you design assessment so that exercising judgment is required rather than performed? ## Introduction Assessments that ask students to produce work presume they can tell good work from poor work, first in others and eventually in their own. That presumption is no longer reliable: when [[generative-ai|generative AI]] can generate competent prose, code, and analysis on demand, the ability to *produce* a product is a weak signal of learning, while the ability to *judge* a product is a strong one. Evaluative judgment names that ability, and the knowledge base's assessment scholarship keeps arriving at it when asking what remains assessable and worth assessing. It sits where [[feedback]], [[formative-assessment]], and [[self-regulated-learning]] meet, because judgment is developed by comparing work against standards and acting on the difference, and it is what makes genuinely [[critical-thinking|critical]] [[student-engagement|engagement]] with AI output possible rather than aspirational. ## What the construct is Evaluative judgment is the capacity to appraise quality — one's own work, peers' work, and increasingly the work of an AI system — using criteria one can reason about rather than recite. Three features distinguish it from adjacent ideas. - **It is about quality, not correctness.** Answering correctly requires knowledge; judging whether an answer is good requires standards a generator cannot supply on the learner's behalf. - **It is developed, not transmitted.** Judgment grows through repeated acts of comparison — against exemplars, explicit criteria, and peers' differing approaches — which is why exemplars, calibration exercises, and [[peer-assessment|peer assessment]] are its natural pedagogies. - **It is domain-[[situated-learning|situated]].** It is exercised inside a discipline's standards of evidence and argument, so it cannot be assessed generically any more than [[transfer-of-learning|transfer]] can be assumed. It is closely related to, but narrower than, authenticity in assessment: [[authentic-assessment|authentic assessment]] asks whether a task resembles worthwhile real-world work; evaluative judgment asks whether the learner can tell good work from poor work. [[self-assessment]] appraises the learner's own work, competence, or progress against criteria, whereas evaluative judgment extends that appraisal to peers' work and to the output of an AI system as well. ## Why generative AI made it central Four arguments converge. - **The product stops being evidence.** Completion of a task no longer demonstrates the capability the task was designed to certify, because a tool can complete much of it. What remains is the reasoning that produces, checks, and accounts for the product — judgment, in short. - **It is the capability that governs AI use.** Deciding whether an AI suggestion is worth accepting, adapting, or rejecting *is* evaluative judgment applied to machine output. Courses that cannot make it visible cannot tell responsible use from substitution, which is why [[academic-integrity|integrity]] scholars increasingly treat critical AI use as an assessable outcome rather than a rule to be policed. - **It is the pivot from enforcement to design.** Where prohibited use cannot be detected, [[teichmann-detecting-undetectable-misconduct-2026|Teichmann (2026)]] names cultivating evaluative judgment among the three design strands that relocate institutions from [[ai-detection|detection-led enforcement]] to validity-centered assessment — because where the institution cannot detect, the capability that survives inspection is whether the student can account for the work and its quality. - **It converts disclosure from confession to evidence.** When a task asks students to explain what they used AI for, which suggestions they accepted or rejected, and how the final submission reflects their own judgment, the [[ai-use-disclosure|declaration]] becomes a demonstration of judgment rather than an admission to be interpreted punitively — which is what lowers the [[academic-integrity|cost of honesty]] that otherwise drives disclosure underground. A further layer concerns the *source* of a judgment, not only its content. [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]] gives that layer a sharper vocabulary: his 13 Saudi computing students rendered two independent judgments within a single evaluation event — whether the feedback was useful, and whether its source held authority to grade. They all accepted ChatGPT's rubric-based feedback as clear and specific while overwhelmingly reserving the grade for the instructor ("I agree with it, but with one condition, that the doctor checks the feedback"), suggesting evaluative judgment in AI-mediated assessment must now include reasoning about evaluative authority itself, not only about the quality of the feedback received. [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] gives the construct its strongest placement in the integrity literature. Extending Eaton's (2023) postplagiarism framing from an [[ethics|ethical]] orientation into assessment design, he argues that detection- and verification-based models of [[academic-integrity|integrity]] are misaligned with work in which human judgment and machine generation are entangled, and reframes integrity as a pedagogical practice enacted through evaluative judgment — the learner's capacity to weigh options, justify academic choices and assume responsibility under epistemic uncertainty. His four practices (annotated decision trails, verification and accountability, oral defense and dialogic accountability, draft differences with version history) make that judgment visible as evidence rather than inferring it from the artifact, and he separates judgment from the constructs it is often collapsed into: reflective practice considers one's thinking retrospectively, [[metacognition]] is the awareness and regulation of cognitive processes, whereas judgment evaluates options against standards and accepts consequences under uncertainty. He also names the failure mode this page tracks — requiring documented reasoning risks "the replacement of one compliance regime with another", since judgment as evidence "remains relational and situated rather than mechanically verifiable". ## Evidence: how judgment develops, and what AI does to it The knowledge base's [[collaborative-learning|collaborative]] and assessment literature supplies convergent evidence that evaluative judgment is *learnable*, *displaceable*, and *designable* — and that AI cuts both ways. - **Comparison-based training measurably builds it.** In a 14-week study with 28 pre-service [[teacher-role|teachers]], [[ai-internal-feedback-evaluative-judgments|AI-generated strong, average, and weak exemplars]] anchored iterative comparison of students' drafts against reference points; evaluation focus expanded from surface language features to content, organization, and coherence — but the reasoning stayed thin, with example-based justification doubling while the most sophisticated comparative reasoning remained rare. Judgment can be scaffolded into existence; it does not appear on demand, and the quality of the reference points matters. - **Feedback design decides whether agency survives.** A multisite, cluster-randomized field experiment with 1,176 first-year undergraduates across 48 sections [[genai-feedback-design-multisite-experiment|compared four feedback conditions]] for scientific argumentation: peer-only, direct GenAI, reflective GenAI (self-evaluate then critique), and hybrid (self-evaluate + peer + GenAI). The **hybrid condition produced the largest argument-quality gain**; direct GenAI feedback risked passive uptake — students outsourcing evaluative judgment to the system — while reflective and hybrid designs preserved [[agency|epistemic agency]] by forcing the student to evaluate their own work first. The authors' core finding is that GenAI's educational value depends less on AI access than on whether the feedback environment preserves agency, judgment, and ownership during revision. - **Structured peer-and-AI review scales calibration.** The PAIRR model [[pairr-ai-peer-review-2025|(Sperber et al., 2025)]] combines peer review with AI review and structured reflection, and was tested in the largest study of students' use of AI feedback to date (654 students, ten writing courses). It treats the *comparison* between one's own judgment, peers', and AI's as the training ground, positioning evaluative judgment as an explicit learning outcome rather than an implicit by-product. - **AI feedback helps only when literacy already exists.** A conceptual framework from feedback-literacy leaders [[zhan-boud-dawson-genai-feedback-engagement|(Boud, Dawson & Yan)]] analyses feedback across eliciting, processing, and enacting, using two contrasting IELTS-with-ChatGPT cases: a student with low [[feedback-literacy|feedback literacy]] used a vague prompt, received generic output, and trusted or over-copied it, while a literate student used AI critically and learned. GenAI lowers the cognitive and emotional barriers to seeking feedback but can itself be hallucinated, biased, or generic — which is why [[teacher-role|teachers]] are argued to model evaluative planning and train judgment, not just [[prompt-engineering|prompting]]. - **Sustainable judgment, not momentary feedback, is the gap.** A [[meta-analysis-systematic-review|scoping review]] of authentic assessment [[zhan-boud-du-authentic-assessment-scoping-review-2025|(Zhan, Boud & Du)]] found AI-formative feedback abundant but **sustainable feedback** — transferable to future contexts — present in only 4 of 23 formative studies. Feedback that closes the current task while leaving the learner unable to judge future work is the evaluative-judgment failure in its clearest form. - **Assessment can be designed to reward judgment instead of output.** The [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|response-region model]] formalizes why: redesign lowers the payoff from hidden outsourcing and raises the value of visible reasoning, explanation, and critique. Asking students to explain which AI suggestions they accepted and rejected turns the decision itself into assessable evidence. - **The institution's own judgment lacks the criteria it asks students to meet.** [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] rated the probative value of 1,855 evidence items across 1,162 generative-AI misconduct allegations and found that evidence quality bore no reliable relationship to case outcomes, because no stage of the procedure set a minimum evidentiary threshold or required investigators to weigh probative value before progressing an allegation — panels weighed evidence through unstructured professional judgment. Detector output was the weakest category in that corpus (standalone detector outputs wholly low in credibility), and [[hadra-ai-detector-accuracy-efl-2026|Hadra et al. (2026)]] show the instrument cannot establish the fact a finding asserts — macro accuracy of 0.69 and 0.61 across 192 texts, near-total failure on hybrid human–AI writing, and a borderline bias concern for EFL writers — which is why the authors' own recommendation is human judgment, clearer [[educational-policy-ai|policy]] and better-specified tasks rather than surveillance. Better-specified matters twice over: [[wright-transcription-not-generation-2026|Wright (2026)]] argues that rules written by platform identity rather than function are over-inclusive by definitional accident, sanctioning a student who did not do the thing the rule was designed to prevent, so the assessor's judgment has to be applied to what a tool did rather than to which platform was used. Judgment exercised without shared criteria cannot be taught, calibrated or audited — the consistency problem this page reports for students, arriving at the level of the institution. Evaluator judgment also carries a measurable direction of its own: five admissions officers separating 50 human from 50 AI admissions essays reached AUC 0.70 (95% CI [0.65, 0.75]), and their quality ratings fell as perceived AI-likelihood rose, an association that survived adjustment for machine-rated quality, actual provenance and applicant characteristics (−0.109, p < .05). Because nobody screened the submissions, that unaided judgment was the entire enforcement mechanism, and each flagged essay carried 1.5 percentage points less admission probability, 2.6 once essay quality was held constant ([[ai-written-admissions-essays-penalized-2026|Isley, Gaebler and Goel 2026]]). - **Classroom-level proxies for judgment are being proposed.** [[austin-ai-agents-assignment-redesign-2026|Austin (2026)]] adds two behavioral signals to a redesigned assignment sequence: the UnBlooms™ *Discernment Rate*, the share of AI outputs a learner interrogates, challenges or revises rather than accepts, and the *First-Pass Acceptance Rate*. Neither is validated — both are offered as cheap per-assignment diagnostics — but they make the construct observable inside ordinary grading, where a near-zero discernment rate across several assignments reads as a task-design signal rather than an individual deficit. ## How misuse displaces judgment The clearest risk is displacement, and it is evidenced in the feedback literature rather than only theorized. When the tool supplies the rubric, the feedback, and the monitoring, the learner's [[metacognition|metacognitive]] and evaluative practice is displaced rather than supported — the assessment becomes a performance of judgment the student never exercised, the same failure mode as [[cognitive-offloading|cognitive offloading]] expressed in the assessment register. Direct GenAI feedback in the multisite experiment pushed toward exactly this passive outsourcing. Grading pressure compounds it: where self-assessment or reflection on AI use is itself graded, students can produce the reflection the rubric wants rather than the reasoning, so genuine judgment requires making the exercise consequential — a decision that shows up in later work, an oral explanation, or a documented rationale that is itself examined. ## Implications for practice - **State it as an outcome.** If appraising work quality — including AI output — is what matters, write it into the [[learning-gains|learning outcomes]] and assess it, rather than leaving it as a by-product of the task. - **Assess the decision, not only the artifact.** Ask students to justify what they kept, changed, or rejected; this turns the thinking into evidence and makes substitution visible without surveillance. - **Use exemplars and calibration deliberately.** Comparative judgment against strong, average, and weak exemplars is the mechanism the evidence supports, and AI makes generating those exemplars cheap — one of its clearest [[pedagogy|pedagogical]] uses. - **Prefer hybrid and reflective feedback over direct output.** Self-evaluation before AI critique, or self + peer + AI, preserves the [[agency]] that direct AI feedback erodes; design the feedback environment rather than just adding a tool. - **Build toward sustainable judgment, not momentary fixes.** Feedback that cannot transfer to the next task trains nothing durable. - **Pair it with process and orals.** Process artifacts and short oral explanations show judgment in action where a final product cannot, which is a core reason AI-era assessment redesign relies on them. ## Connected Concepts - [[ai-education]] — AI in education (umbrella) - [[assessment-validity]] - [[feedback]] - [[feedback-literacy]] - [[formative-assessment]] - [[self-assessment]] - [[authentic-assessment]] - [[academic-integrity]] - [[self-regulated-learning]] - [[metacognition]] - [[critical-thinking]] - [[assessment]] - [[ai-detection]] - [[agency]] - [[higher-ed]] ## Connected Articles - [[ai-internal-feedback-evaluative-judgments]] — How AI-supported internal feedback develops evaluative judgments, and where reasoning stays thin - [[genai-feedback-design-multisite-experiment]] — Hybrid self/peer/AI feedback preserves agency and judgment better than direct AI feedback - [[pairr-ai-peer-review-2025]] — Peer-and-AI review with structured reflection as calibration at scale - [[zhan-boud-dawson-genai-feedback-engagement]] — How feedback literacy governs whether AI feedback teaches or substitutes - [[zhan-boud-du-authentic-assessment-scoping-review-2025]] — The sustainability gap in AI feedback - [[teichmann-detecting-undetectable-misconduct-2026]] — Evaluative judgment as one of the design strands replacing detection-led enforcement - [[mohamed-temimi-assessment-imperfect-information-disclosure-2026]] — Making visible reasoning and judgment the thing the rubric rewards - [[learner-centered-feedback-ai]] — Teacher evaluative judgment with AI feedback tools - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond detection: authenticity redesigned rather than policed - [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene]] — Reconceptualizing assessment validity for the age of generative AI - [[du-yuan-epistemic-dependence-2026]] — Contestability, recoverability, and the criteria separating reliance from dependence - [[rethinking-ai-writing-feedback-literacy]] — Feedback literacy for AI-assisted writing - [[student-perspectives-ai-writing-grading-2026]] — Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education - [[sharma-judgment-visible-genai-assessment-2026]] — Judgment as the evaluative core of integrity: annotated decision trails, oral defense, version history (Sharma 2026) - [[munoz-misconduct-allegation-evidence-2026]] — Panels weighing evidence without codified standards: the institution's own judgment problem (Munoz et al. 2026) - [[hadra-ai-detector-accuracy-efl-2026]] — Detector inaccuracy and hybrid-writing failure: human judgment as the recommended replacement for the verdict (Hadra et al. 2026) - [[wright-transcription-not-generation-2026]] — Function, not platform: judging what a tool did rather than what it is (Wright 2026) - [[austin-ai-agents-assignment-redesign-2026]] — UnBlooms and the Discernment Rate: grading the reasoning trail behind AI-assisted work (Austin 2026) - [[llm-grade-bands-calibration-bias-2026]] — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing - [[ai-written-admissions-essays-penalized-2026]] — AI-written admissions essays are widespread but penalized - [[authentic-assessments-generative-ai-pilot-2026]] — Designing Authentic Assessments with Generative AI: A Pilot Study of Assessment Authentifire in Higher Education - [[obyrne-co-constructing-ai-boundaries-agency-judgment-2026]] — the Agency Check (credible, relevant, acceptable, nuanced) that structures each AI interaction --- ## [Feedback](https://edtechdev.github.io/aied/concepts/feedback/) > **Feedback** — information provided to a learner about their performance or understanding that is intended to close the gap between current and desired performance. In [[ai-education|AI in education]], feedback has become a central and rapidly transforming theme: [[ai-technologies|AI systems]] now generate, deliver, and even teach students how to use feedback, reshaping every stage of the feedback process. ## Questions to Consider - Feedback is often imagined as information handed to a learner, but [[research-methods-aied|research]] argues it only 'counts' as feedback when the learner makes sense of and acts on it. If feedback is a relationship rather than a message, what changes about how we should deliver it? - Students' perception of whether feedback came from an AI or a human can change how much they learn from it. Why do you think the *source* matters so much — and what does that imply for AI-generated feedback in your context? - Immediate feedback can be a trap: the 'correct-answer trap' shows that students who simply copy corrections learn less. When does instant feedback help, and when does it short-circuit the very learning it's meant to support? - Research even found that a well-designed sequence of feedback — encouragement, then hints, then the answer — designed to build autonomy actually *harmed* learning despite boosting [[student-engagement|engagement]]. If students feel good about feedback that makes them learn less, what should we optimize? - Feedback is described as a system with parts: the feedback loop, quality of provision, the learner's uptake capacity, and the assessment context. If you were trying to improve feedback in your own course, which part would you fix first, and why? ## Introduction This is the umbrella concept for the knowledge base's feedback-related ideas. Feedback sits at the intersection of assessment and learning: without feedback, assessment measures performance but does not improve it; with effective feedback, assessment becomes a learning event. The knowledge base treats feedback as a **system** with multiple facets — the quality of the feedback itself ([[ai-feedback-quality]]), the loop through which it closes the learning gap, the learner's capacity to use it ([[feedback-literacy]]), and the assessment contexts in which it operates ([[formative-assessment]], [[peer-assessment]], [[automated-assessment]]). - **[[yilmaz-genai-feedback-srl-online-higher-ed-2026|Yilmaz et al.]]** show that students' perception of the feedback *source* (AI vs human) shapes self-regulated learning from [[generative-ai|GenAI]] feedback — a crucial qualifier for feedback effectiveness claims. ## The feedback system Feedback is best understood not as a single event but as a connected system of interacting parts, each of which the knowledge base documents as its own concept: - **The feedback loop** — the mechanism through which feedback closes the gap between current and desired performance; the cycle of performance → feedback → revision → improved performance that drives learning. - **[[ai-feedback-quality]]** — the accuracy, usefulness, timeliness, and [[pedagogy|pedagogical]] value of feedback generated by AI; the provision side of the system. - **[[feedback-literacy]]** — the capabilities and dispositions students need to understand, evaluate, and act on feedback; the uptake side of the system. - **[[formative-assessment]]** — the assessment context in which feedback is used to improve learning while it is still in progress, rather than merely to judge it. - **[[peer-assessment]]** — feedback exchanged between [[learners]], increasingly augmented by AI and a site for developing feedback literacy. - **[[automated-assessment]]** — AI-driven scoring and feedback, from [[automated-essay-scoring|automated essay scoring]] to confidence-aware short-answer grading. ### The feedback loop The feedback loop is the cyclical process where AI systems assess student work, deliver feedback, observe the student's response, and adapt subsequent instruction. AI-mediated feedback loops operate at multiple timescales: - **Immediate feedback:** [[automated-assessment|Automated grading systems]] and [[intelligent-tutoring|AI tutors]] provide real-time correction during [[problem-solving]]. [[correct-answer-trap-ai-tutor|Correct-answer trap research]] shows that immediate feedback can short-circuit learning if students simply copy corrections. Generating that immediate corrective feedback with LLMs inside a tutor is feasible but imperfect: across 6,926 logged transactions in the Apprentice Tutor College Algebra platform, [[reddig-maclellan-personalized-feedback-llm-2026|Reddig, Arora & MacLellan (2025)]] found GPT-4 produced hints targeted to the student's specific error ~66% of the time, yet ~35% of hints were too general, incorrect, or prematurely gave away the answer, and LLM-based automated quality checks misaligned with human judgment — reinforcing that unvetted AI feedback in the immediate loop carries real short-circuiting risk. - **Assignment-level feedback:** [[formative-assessment]] systems and [[ai-feedback-quality|AI feedback quality research]] examine whether AI-generated assignment feedback improves subsequent work. [[sequenced-ai-feedback-learning|Sequenced feedback studies]] test whether the order of feedback matters. - **Course-level loops:** [[learning-analytics|Learning analytics dashboards]] and [[edtech-platform|educational platforms]] aggregate feedback across assignments to identify patterns and recommend interventions. The effectiveness of a feedback loop depends on [[ai-feedback-quality|feedback quality]] — accuracy, specificity, timeliness, and actionability. [[becerra-aicofe-feedback-2026|AI peer feedback systems]] add a social dimension to the loop. **The human in the loop:** feedback loops are not purely automated — teachers often mediate AI-generated feedback before it reaches learners. [[learner-centered-feedback-ai|Studies of AI feedback tools for teachers]] (e.g., the PolyFeed tool combining an ML detector with an [[llm]] rephraser) find teachers use professional judgment to **accept, edit, or reject** AI suggestions — an "assist but verify" pattern — and systematically moderate exaggerated praise and generic suggestions to protect authenticity and voice. The **relational/[[affective-computing|affective]] dimension** of feedback (student–teacher relationship, encouragement) most strongly resists AI delegation, suggesting this part of the loop remains inherently human. This human-in-the-loop mediation connects to [[human-in-the-loop-ai]] and to [[teacher-role]]. The relational reading cuts deeper than substitution. In workshops with 12 students and 18 educators, [[ai-feedback-ecosystem-higher-education-2026|Bearman et al. (2026)]] found students treated AI as an additional but flawed source, weighing it against rubrics, lecture notes, discussion boards and educator comments, and using it as a holding pattern when human responses were slow; one pasted thin educator comments into a chatbot to work out what to do next. AI's presence changed relations rather than merely adding comments, to the point that one student stopped trusting peer discussion-board feedback as probably AI-written. The authors therefore propose humans-as-the-loop: people helping people build stronger feedback relationships over time, in place of oversight of machine output. ### How AI transforms feedback AI changes feedback in three consequential directions, each raising the stakes of the other facets: - **Volume and immediacy.** AI can generate feedback instantly and at scale, dramatically increasing how much feedback students receive. [[genai-feedback-design-multisite-experiment|Multi-site GenAI feedback studies]] and [[ai-generated-feedback-higher-ed|AI feedback in higher education]] document AI feedback experienced as comparable to teacher feedback in acceptability and supportiveness. A PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 empirical studies (2023–2025) corroborates this scale benefit — LLMs can accelerate grading and deliver rapid, personalized feedback at scale, especially in large or [[higher-ed|higher-education]] cohorts — while cautioning that such feedback is sometimes too generic or misaligned with the assigned grade, and that reliability slips on longer, [[multilingual-learning|multilingual]], or nuanced tasks ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). - **Evaluation demands on the learner.** Because AI feedback can be inaccurate or hallucinated, students must judge whether to [[trust]] and act on it. This is where [[feedback-literacy]] and [[trust-calibration]] become decisive: [[mendoza-ai-feedback-feedback-literacy-srl|Mendoza et al. (2026)]] show that only feedback-literate students convert AI feedback into [[self-regulated-learning]] gains, while [[llm-fallacy-misattribution|the LLM fallacy]] captures how students may over-credit AI feedback to their own competence. - **Task-level suitability: sorting feedback by epistemic status.** [[tripartite-feedback-framework-ai-assessment-2026|Venetsanos (2026)]] argues that the frameworks the field relies on — Hattie and Timperley's levels, Boud and Molloy's design principles — explain *where* feedback operates in learning but not whether a given feedback task involves rule-checking, factual verification, or interpretive judgment, and that it is the second distinction that decides whether automation is defensible. The proposed tripartite split is low-level structural and presentational feedback (formatting, referencing, grammar, sentence-level clarity — rule-based against explicit standards and safely automatable), intermediate-level factual and procedural verification (whether a calculation used a specified method correctly, whether a stated fact matches authoritative references — comparison against established knowledge, not judgment of reasoning), and high-level [[critical-thinking|critical evaluation]] and synthesis, where judgment is interpretive and must stay human. The levels are dimensions rather than a ladder: one calculation can be correctly presented, arithmetically wrong, and methodologically inappropriate at the same time, and can attract verification, process feedback and self-regulation feedback together. The framework's sharpest boundary case is clarity: sentence-level clarity is surface-level, but *argumentative* unclarity often signals incomplete conceptual grasp, so flagged instances should escalate to human review rather than being corrected as prose — fixing the writing would leave the actual problem untouched. - **Role-aware feedback.** [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] showed that [[prompt-engineering|prompting]] the same LLM to evaluate the same student artifact under different roles — instructor, peer reviewer, grant reviewer — produced qualitatively distinct feedback: instructors were encouraging and process-oriented, peers supportive and conversational, grant reviewers formal and outcomes-oriented. These differences were epistemic, not merely stylistic, foregrounding different aspects of design practice — evidence that role-aware prompting can generate role-sensitive evaluative feedback rather than a single generic response. - **Feedback literacy as a training goal.**single generic response. - **Feedback utility vs. evaluative authority.** [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]] shows students can simultaneously accept AI-generated feedback as useful and reject AI as the appropriate grader — separating *feedback utility* (the informational value of surface-level feedback) from *evaluative authority* (who may decide the grade). In a transparent post-[[assessment|assessment design]], Saudi computing students treated these as analytically distinct judgments, valuing ChatGPT's clarity for revision while reserving grading authority for the human instructor and articulating the dialogical, institutionally-weighted reasons why. The finding extends [[feedback-literacy]] beyond judging feedback quality to reasoning about evaluative authority itself, and suggests [[explainable-ai|transparency]] about AI involvement can make feedback an object of critical reflection rather than passive acceptance. - **Feedback literacy as a training goal.** Rather than only providing feedback, AI tools are increasingly designed to *teach* students to use feedback — [[tubino-adachi-ai-automated-feedback-literacy|reframing automated feedback tools as literacy-building instruments]], [[zhan-boud-dawson-genai-feedback-engagement|framing GenAI as an enabler of feedback engagement]], and [[richmond-nicholls-genai-psych-feedback-ai-literacies|using GenAI critique tasks to build psychological, feedback, and AI literacies together]]. ### What makes feedback effective: calibration evidence [[metacognitive-training-optimal-cognitive-offloading-2026|Ngai & Gilbert (2026)]] provide a clean experimental result on feedback design with direct relevance to AIED: **veridical, immediate, trial-by-trial feedback tied to a learner's own prior prediction is what changes behavior — prediction or beliefs alone are not enough**. In their four-group design, feedback combined with a preceding performance prediction improved calibration and behavior, while predictions without feedback did nothing, and adding an explicit over-/underconfidence label added nothing further. LLMs can also serve as calibration partners in this process: [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] found that after rubric refinement, LLMs demonstrated greater consistency than some human raters in applying performance thresholds — useful for norming sessions and [[formative-assessment|formative peer feedback]] environments, where the model helps calibrate [[evaluative-judgment|evaluative judgment]] rather than replace it. This aligns with the knowledge base's feedback-system view: feedback's power lies in closing the gap against a learner's *own* estimate (the provision–uptake pairing), and effective feedback should be immediate, accurate, and explicitly connected to what the learner predicted — a design principle for AI feedback systems and [[feedback-literacy]] training alike. Effort, however, is not automatically the missing ingredient. In a preregistered between-subjects experiment with 302 US adults learning introductory Python, [[structured-reflection-ai-explanatory-feedback-2026|Asher, Gold & Carvalho (2025)]] paired personalized [[llm|Claude Sonnet 3.5]] explanatory feedback with a structured three-step self-explanation prompt — for each incorrect code section, what the correct code does and why the learner's own attempt failed — and the added reflection bought nothing. Reflective Practice participants spent 4.1 minutes reviewing feedback against 2.1 minutes in plain Practice (t(267) = 11.30, p < .001), which left room for only 2.0 practice problems instead of 3.4 (t(267) = 8.20, p < .001); each problem was worth the same either way (Reflection × problem-number OR = 1.03, z = 0.23, p = .486), so Practice ended with higher mastery, 79% versus 65% (d = .41), and never fell behind on near-[[transfer-of-learning|transfer]], far-transfer, or code-evaluation items. Reflection also failed to deliver the metacognitive benefit it targeted: judgments of learning did not differ across arms, and the practice conditions were better calibrated (underconfident by 11 points, against video watchers 20 points overconfident, d = −.93) and reported less [[cognitive-offloading|distraction]] (d = −.80) than reflective learners, whose reflections averaged only 35 words. The authors' reading is that the AI feedback had already supplied the [[scaffolding]] that self-explanation normally adds — a message that names the error, explains the correct approach and describes the underlying concept leaves little for a reflection prompt to do — and that writing reflections while still learning unfamiliar syntax competed directly with the varied repetition that builds transferable skill. This locates a boundary condition for [[desirable-difficulties]] and [[productive-failure]] on the processing-depth axis: effort pays when it drives additional retrieval, application or comparison, and is redundant when the learner is already receiving a personalized explanation of their specific error. It also bears on how [[ai-feedback-quality|feedback quality]] interacts with [[self-regulated-learning|uptake]]: the more elaborated and targeted the feedback, the narrower the space for an added layer to add value, which shifts design attention from deepening one episode to spending scarce practice minutes on volume, problem variability, and timing reflection to moments such as a detected plateau — the study tested a single unvetted enhancement, and the authors note that these variants remain open. ### Feedback across assessment contexts The knowledge base's feedback research spans the full range of assessment contexts, each with distinct feedback dynamics: - **Formative feedback** is the canonical site of feedback-for-learning — [[formative-assessment|formative assessment]] feeds [[self-regulated-learning|self-regulated learning]] through feedback that arrives while learning is still in progress ([[automated-formative-assessments-a-level-sciences|automated formative assessments]], [[ai-feedback-enactment-workflow-2026|feedback enactment]]). - **Summative feedback** is the feedback attached to [[summative-assessment|summative assessment]] — end-of-unit tests, oral exams, and proctored/closed-book examinations. In the AI era, the summative setting is where AI resistance matters most (see [[summative-assessment]]): feedback on a proctored oral exam or closed-book assessment tests genuine learning, whereas feedback on AI-assisted homework can be inflated. [[fenton-oral-exams-ai-authentic-assessment-2025|Oral assessments]] reframe the feedback moment as a live, interactive dialogue — feedback becomes immediate, conversational, and inseparable from the assessment itself, which is precisely why they resist AI substitution. - **Peer feedback** adds a social layer — [[peer-assessment|peer feedback]] develops feedback literacy and is increasingly AI-augmented ([[becerra-aicofe-feedback-2026|AI peer feedback]], [[irwin-muller-efl-peer-feedback-literacy|GenAI in EFL peer feedback]]). Yet peer feedback is only as good as its actual delivery: in a semester-long field experiment, under two-thirds of students assigned peer feedback received no textual feedback at all, and much of what was delivered was non-targeted praise — so when individual GPT-4 feedback replaced unreliable peers, students sustained the highest participation and posted the largest content [[learning-gains|learning gains]] ([[gpt4-feedback-student-activation-2026|Geschwind et al., 2026]]), evidence that the AI's edge is a *reliability* advantage rather than inherent superiority. - **Automated feedback** scales delivery — [[automated-assessment|automated scoring and feedback]] from essay scoring to short-answer grading. [[yin-arthur-ai-teaching-assistant-engineering-econ-2026|Arthur (Yin et al. 2026)]] extends this to unstructured [[quantitative-research|quantitative]] problems: a per-question [[machine-learning|XGBoost]] diagnosis backbone infers rubric-labeled mistakes from students' submitted numerical answers to Engineering Economics Calculated Formula Questions and turns them into natural-language feedback, while a dialogue-based scheme requests intermediate answers only when confidence is low — balancing feedback accuracy against the collection burden on students. A key cross-context insight is that **the reliability of feedback depends on the integrity of the assessment it is attached to**: feedback is only as trustworthy as the measure it responds to. In the AI era this pushes educators toward [[authentic-assessment|authentic]] and AI-resistant [[summative-assessment|summative formats]] where the feedback a student receives reflects genuine learning rather than AI-assisted output. The "assessment for learning" paradigm reframes feedback as an overarching philosophy rather than a component. [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] synthesize the wider higher-education literature to frame [[formative-assessment|formative]], ongoing, and individualized feedback as central to that philosophy, balancing formative with summative purposes and foregrounding [[agency|student agency]], self-[[regulation]], and [[metacognition|metacognitive]] skill. They note, however, that in [[business-education|management education]] feedback-rich practices remain unevenly adopted: self- and peer-assessment (a site of peer feedback) dominate the literature, while reassessment — which uses feedback-driven second chances to improve learning — is virtually absent. ### The provision-uptake pairing The knowledge base's core feedback insight is that **feedback quality and feedback literacy are two sides of one system**: high-quality feedback is inert without a literate recipient, and a literate student gains little from poor feedback. [[ai-feedback-quality]] covers the provision side (is the feedback accurate, timely, actionable?), while [[feedback-literacy]] covers the uptake side (can the student judge and act on it?). The feedback loop is what connects them — the mechanism by which quality feedback, received by a literate learner, closes the gap. Designing effective AI feedback therefore means designing both the system and the student. A two-layer model shows how the pairing can be organized around a machine without delegating judgment to it. In [[ai-agents-joyful-assessment-third-space-2026|El Khoury and Ma's worked example]], an [[agentic-ai|agent]] produces a draft rubric-based evidence report on each transcript with every rating tied to quoted excerpts, and the instructor then reviews the report, leads a debrief and decides what the evidence means for that student — a division the authors summarize as the AI organizing evidence while the instructor interprets it. Their claim about uptake is appraisal-based: feedback and iteration outside the social hierarchies students navigate with peers and instructors are easier to attempt, and rehearsal at the student's own pace is what they argue turns occasional confidence into a settled habit of engaging with feedback. ### Why feedback matters for AI in education Feedback is one of the most consequential and best-evidenced mechanisms in education, and AI both amplifies and complicates it. Well-architected AI feedback can match or exceed human feedback and scale across cohorts, but it demands new learner capabilities ([[feedback-literacy]], [[ai-literacy]]) and carries risks (uncritical acceptance, [[cognitive-offloading|Over-Reliance]]). As AI-generated feedback becomes ubiquitous, the knowledge base frames feedback as a whole system — quality, loop, literacy, and assessment context working together — rather than as any single component. ## Connected Concepts - [[eportfolio]] - [[ai-feedback-quality]] - [[feedback-literacy]] - [[formative-assessment]] - [[summative-assessment]] - [[peer-assessment]] - [[self-assessment]] - [[automated-assessment]] - [[assessment]] - [[authentic-assessment]] - [[assessment-validity]] - [[self-regulated-learning]] - [[ai-literacy]] - [[scaffolding]] - [[writing-education]] ## Connected Articles - [[ai-agents-joyful-assessment-third-space-2026]] — AI agents, joyful assessment, and third space - [[student-perspectives-ai-writing-grading-2026]] — Student perspectives on transparent AI-assisted writing assessment (AlGhamdi 2026) - [[usher-faraon-who-grades-best-2026]] — Comparing ChatGPT, peer, and instructor grading across project quality levels (Usher & Faraon 2026) - [[layer-sensitive-cognitive-offloading-writing-2026]] — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026) - [[deceptive-overgeneralization-adaptive-learning-2026]] — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026) - [[mejeh-fromm-srl-adaptive-learning-feedback-2026]] - [[farrokhnia-genai-feedback-student-revisions-2026]] — Teacher vs. GenAI feedback: students revise less with AI - [[yilmaz-genai-feedback-srl-online-higher-ed-2026]] — GenAI feedback and self-regulated learning: perceived source matters - [[chang-genai-peer-feedback-collaborative-argumentation-2026]] — GenAI-assisted peer feedback in collaborative argumentation - [[luo-eaton-ai-student-feedback-ethics-2026]] — Ethics of AI in student feedback - [[metacognitive-training-optimal-cognitive-offloading-2026]] — Metacognitive training facilitates optimal cognitive offloading (Ngai & Gilbert 2026) - [[liu-deris-ai-feedback-literacy-uptake]] — AI Feedback Literacy scale and uptake prediction (Liu & Deris 2025) - [[zhan-boud-dawson-genai-feedback-engagement]] — GenAI as enabler of student feedback engagement (Zhan, Boud, Dawson & Yan 2025) - [[mendoza-ai-feedback-feedback-literacy-srl]] — Feedback literacy moderates AI feedback → SRL (Mendoza et al. 2026) - [[hawkins-feedback-literacy-ai-essay-writing]] — Feedback literacy predicts essay grade in AI writing (Hawkins et al. 2026) - [[rethinking-ai-writing-feedback-literacy]] — Feedback literacy scripts for AI-assisted writing (Dai 2026) - [[feedback-literacy-scripts-eap-writing]] — Feedback literacy scripts + second-rater in EAP writing (Yao 2026) - [[jin-genai-learning-analytics-feedback-literacy]] — GenAI learning analytics in feedback (Jin et al. 2025) - [[tubino-adachi-ai-automated-feedback-literacy]] — AI automated feedback tool for feedback literacy (Tubino & Adachi 2025) - [[irwin-muller-efl-peer-feedback-literacy]] — GenAI in EFL peer feedback for feedback literacy (Irwin & Muller 2026) - [[scaffolding-srl-feedback-genai-human-peers]] — Scaffolding self-regulated feedback: GenAI vs. human peers (Gu, Chen & Yan 2026) - [[ai-generated-feedback-higher-ed]] — AI-generated feedback in higher education - [[care-full-feedback-genai]] — Care-full feedback design with GenAI - [[feedback-futures-genai]] — Feedback futures with GenAI - [[ai-feedback-ecosystem-higher-education-2026]] — The feedback ecosystem in higher education: AI reworks feedback relations, and humans as the loop (Bearman et al. 2026) - [[learner-centered-feedback-ai]] — Learner-centered AI feedback practices - [[genai-feedback-design-multisite-experiment]] — Multi-site GenAI feedback design - [[ai-feedback-critical-thinking-writing-2026]] — AI feedback and critical thinking in writing - [[richmond-nicholls-genai-psych-feedback-ai-literacies]] — GenAI assessment builds psychological, feedback, and AI literacies - [[zhao-learnlens-feedback-educators-loop]] — LearnLens: curriculum-grounded feedback with educator oversight (Zhao et al. 2025) - [[sequenced-ai-feedback-learning]] — Sequenced AI feedback studies - [[becerra-aicofe-feedback-2026]] — AI peer feedback systems - [[automated-formative-assessments-a-level-sciences]] — Automated formative assessments in A-level sciences - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant assessment - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Marked Pedagogies: bias in automated writing feedback - [[shap-llm-rationales-teaching-quality-assessment]] — SHAP vs LLM rationales for rubric-based teaching feedback - [[llm-adaptive-programming-error-explanations-2026]] — LLM adaptive explanations of programming errors - [[student-perceptions-ai-study-productivity-2026]] — Students' Perceptions of Artificial Intelligence Tools for Study Productivity and Learning: An Exploratory Survey Study - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design - [[gpt4-feedback-student-activation-2026]] - [[reddig-maclellan-personalized-feedback-llm-2026]] - [[yin-arthur-ai-teaching-assistant-engineering-econ-2026]] - [[mesny-innovative-assessment-grading-management-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] - [[tripartite-feedback-framework-ai-assessment-2026]] — Tripartite framework: sorting feedback by epistemic status and the five boundary principles for AI involvement (Venetsanos 2026) - [[structured-reflection-ai-explanatory-feedback-2026]] — Benefit or bottleneck? Structured reflection on AI explanatory feedback (Asher, Gold & Carvalho 2025) - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback - [[peer-group-vs-ai-feedback-2026]] — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education - [[pivot-generative-video-tutors-stem-2026]] — From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning --- ## [Feedback Literacy](https://edtechdev.github.io/aied/concepts/feedback-literacy/) > **Feedback literacy** — the capabilities and dispositions students need to understand, evaluate, and act on feedback to improve their learning. It is the learner-side counterpart to feedback provision: whereas [[ai-feedback-quality]] and [[feedback|Feedback Loop]] concern the quality and mechanics of the feedback system, feedback literacy concerns the learner's capacity to seek, make sense of, judge, and use feedback productively. ## Questions to Consider - Well-designed feedback only helps students who can interpret and act on it. If two students receive identical feedback and learn different amounts, where does the difference live — and is it the student's fault or the system's? - [[research-methods-aied|Research]] finds that students with stronger feedback literacy benefit more from AI feedback, while weaker-literacy students show minimal or even negative effects. What does that suggest about simply adding AI feedback to a course without also building students' capacity to use it? - Feedback literacy includes seeking feedback, making judgments, managing affect, and acting on feedback — not just receiving it. When did you last *seek out* feedback rather than wait for it, and what made you brave enough (or not) to do so? - If AI can now generate abundant, instant feedback, has the bottleneck shifted from feedback provision to the learner's ability to use it? What might change about how you design feedback if you saw students' feedback literacy as the real constraint? ## Introduction Feedback literacy matters because well-designed feedback only helps students who can interpret and act on it. A student who cannot evaluate whether AI-generated feedback is accurate, or who does not know how to turn feedback into a concrete revision, learns far less from the same feedback than a more feedback-literate peer. As AI reshapes feedback provision, feedback literacy has become a central boundary condition for whether AI feedback improves learning. ### What feedback literacy is Feedback literacy is widely framed as a set of interrelated capabilities — the capacity to appreciate feedback, make judgments, manage affect, and take action (after Carless & Boud). The knowledge base's articles cluster feedback literacy around several capabilities: - **Seeking and eliciting feedback** — proactively requesting feedback rather than passively receiving it. - **Making judgments** — evaluating the accuracy and usefulness of feedback, including feedback produced by AI. - **Sense-making** — interpreting feedback in relation to task goals and criteria, and understanding what it implies for improvement. - **Managing affect** — engaging productively with feedback without being discouraged or over-inflated by it. - **Acting on feedback** — translating feedback into concrete revisions or changes in approach ([[feedback|Feedback Loop]], [[self-regulated-learning]]). - **Judging evaluative authority** — a further capacity proposed by [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]]: alongside appreciating feedback, making judgments, managing affect and taking action, students reasoning about AI-produced feedback needed to judge *which source holds evaluative authority*, separately from judging feedback quality. Because participants were told ChatGPT generated both score and feedback, their feedback literacy was exercised on the source and legitimacy of the evaluation, not only its content, and it produced demands for instructor validation rather than passive acceptance. ### How feedback literacy appears in the research - **Feedback literacy as a moderator of AI feedback value:** [[mendoza-ai-feedback-feedback-literacy-srl|Mendoza et al. (2026)]] show that feedback literacy moderates the link between ChatGPT acceptance and [[self-regulated-learning]]: students with stronger literacy perceive greater SRL benefit from AI feedback, while weaker-literacy students show minimal or even negative ([[cognitive-offloading|Over-Reliance]]) effects. Feedback literacy is a boundary condition for whether students can "make sense of" AI feedback. - **Feedback literacy predicts learning from AI-assisted writing:** [[hawkins-feedback-literacy-ai-essay-writing|Hawkins et al. (2026)]] find that feedback literacy was the only significant positive predictor of essay grade in an AI-enhanced essay-writing task, while [[liu-deris-ai-feedback-literacy-uptake|Liu & Deris (2025)]] develop and validate an AI Feedback Literacy (AIFL) scale and show it predicts feedback uptake. - **Frameworks for [[generative-ai|GenAI]]-enabled feedback [[student-engagement|engagement]]:** [[zhan-boud-dawson-genai-feedback-engagement|Zhan, Boud, Dawson & Yan (2025)]] (Boud and Dawson are leading feedback-literacy scholars) argue GenAI can *enable* student feedback engagement, mapping a cyclical self-[[regulation]] feedback model onto the eliciting/processing/enacting phases. - **AI feedback arrives without the teacher's framing:** [[brunnstrom-ai-interaction-literacy-srl-2026|Brunnström and Palmqvist (2026)]] argue that GenAI makes feedback literacy *more* demanding, because responses are generated interactively and without a teacher's immediate framing: students must interpret, evaluate, engage with and use them — including deciding when a response is too abstract, when to persist, and when to return to conventional resources. Across their eight-exchange demonstration, simplification and usable structure appeared only after learner interventions, making the eliciting/processing/enacting cycle of feedback engagement something the student has to drive ([[zhan-boud-dawson-genai-feedback-engagement|Zhan, Boud, Dawson & Yan]] mapped that cycle for GenAI; [[self-regulated-learning]]). - **Feedback literacy in AI-assisted writing and EAP:** [[rethinking-ai-writing-feedback-literacy|Feedback literacy scripts]] and [[feedback-literacy-scripts-eap-writing|second-rater mechanisms]] train students to engage critically with AI feedback during writing revision, shifting revision toward argument-level improvement rather than surface edits. - **Automated feedback tools for literacy development:** [[tubino-adachi-ai-automated-feedback-literacy|Tubino & Adachi (2025)]] argue AI automated feedback tools should be reframed as instruments for *developing* students' feedback literacy, not just providing more feedback. - **Peer feedback and feedback literacy:** [[irwin-muller-efl-peer-feedback-literacy|Irwin & Muller (2025)]] position GenAI within EFL peer feedback to train feedback literacy and enable uptake in speaking classes, and [[scaffolding-srl-feedback-genai-human-peers|scaffolding studies]] compare GenAI vs. human peers in fostering self-regulated feedback. - **Feedback literacy in learning analytics and GenAI dashboards:** [[jin-genai-learning-analytics-feedback-literacy|Jin et al. (2025)]] examine how students perceive GenAI-powered [[learning-analytics]] feedback from a feedback-literacy perspective. - **Feedback literacy as a goal of AI-literacy and assessment design:** [[richmond-nicholls-genai-psych-feedback-ai-literacies|Richmond & Nicholls (2025)]] use a process-over-artifact assessment in which students critique ChatGPT output against a rubric to build feedback, psychological, and AI literacies together; [[learner-centered-feedback-ai|learner-centered AI feedback]] and [[care-full-feedback-genai|care-full feedback design]] link feedback quality to the learner's capacity to engage. ### Why feedback literacy matters for AI in education AI changes feedback in two directions that both raise the stakes of feedback literacy. First, AI dramatically increases the *volume and immediacy* of feedback ([[ai-feedback-quality]], [[feedback|Feedback Loop]]), so students confront far more feedback they must triage and evaluate. Second, AI-generated feedback carries distinct risks — inaccuracy, [[hallucination-risk|hallucination]], and the "illusion of mastery" — that demand [[critical-thinking|critical evaluation]] skills [[cognitive-offloading|Over-Reliance]] [[llm-fallacy-misattribution]]. Feedback literacy therefore becomes a core component of [[ai-literacy]]: knowing not only how to prompt an AI for feedback, but how to judge whether the feedback is worth acting on and how to convert it into genuine learning rather than task completion. [[ai-feedback-ecosystem-higher-education-2026|Bearman and colleagues (2026)]] add a further demand to that list. In their workshop study, students were not managing one feedback source but a network: they weighed educator comments against AI output, rubrics, course materials and peers, and when human feedback was slow they kept working with AI as a holding pattern, revising later if it had misled them. The authors suggest students can act as their own human-as-the-loop by managing those relations from the granular task to the broader trajectory, which is itself a feedback-literacy capability. Their warning is that thin or late educator comments are what push students toward unverified sources. ### Connections to related concepts Feedback literacy connects to [[ai-feedback-quality]] and [[feedback|Feedback Loop]] (the provision side it complements), [[formative-assessment]] (the assessment cycle it feeds), and [[self-regulated-learning]] (the [[self-assessment]] and adaptation it supports). It is a subset of [[ai-literacy]] when applied to AI-generated feedback, intersects with [[peer-assessment]] in collaborative contexts, and is particularly consequential for [[writing-education]]. It also connects to [[metacognition]] and [[trust-calibration]] — the ability to judge whether feedback is trustworthy. ## Connected Concepts - [[eportfolio]] - [[ai-feedback-quality]] - [[feedback]] - [[formative-assessment]] - [[self-regulated-learning]] - [[ai-literacy]] - [[peer-assessment]] - [[self-assessment]] - [[writing-education]] - [[metacognition]] - [[trust-calibration]] - [[cognitive-offloading]] - [[higher-ed]] ## Connected Articles - [[brunnstrom-ai-interaction-literacy-srl-2026]] — AI feedback without teacher framing raises the feedback-literacy bar (Brunnström & Palmqvist 2026) - [[mejeh-fromm-srl-adaptive-learning-feedback-2026]] - [[sutama-chatgpt-eportfolio-speaking-2026]] - [[ni-lam-multiliteracies-ai-portfolio-2026]] - [[mendoza-ai-feedback-feedback-literacy-srl]] — Feedback literacy moderates AI feedback → self-regulated learning (Mendoza et al. 2026) - [[hawkins-feedback-literacy-ai-essay-writing]] — Feedback literacy predicts essay grade in AI-enhanced writing (Hawkins et al. 2026) - [[liu-deris-ai-feedback-literacy-uptake]] — AI Feedback Literacy scale and uptake prediction (Liu & Deris 2025) - [[zhan-boud-dawson-genai-feedback-engagement]] — GenAI as enabler of feedback engagement framework (Zhan, Boud, Dawson & Yan 2025) - [[rethinking-ai-writing-feedback-literacy]] — Feedback literacy scripts and calibration training for AI-assisted writing (Dai 2026) - [[feedback-literacy-scripts-eap-writing]] — Feedback literacy scripts + second-rater in EAP writing revision (Yao 2026) - [[tubino-adachi-ai-automated-feedback-literacy]] — AI automated feedback tool for developing feedback literacy (Tubino & Adachi 2025) - [[irwin-muller-efl-peer-feedback-literacy]] — Positioning GenAI in EFL peer feedback to train feedback literacy (Irwin & Muller 2026) - [[scaffolding-srl-feedback-genai-human-peers]] — Scaffolding self-regulated feedback: GenAI vs. human peers (Gu, Chen & Yan 2026) - [[jin-genai-learning-analytics-feedback-literacy]] — GenAI learning analytics in feedback, feedback literacy perspective (Jin et al. 2025) - [[richmond-nicholls-genai-psych-feedback-ai-literacies]] — GenAI assessment builds psychological, feedback, and AI literacies (Richmond & Nicholls 2025) - [[learner-centered-feedback-ai]] — Teachers' practices and perceptions of AI learner-centered feedback - [[care-full-feedback-genai]] — Care-full feedback design with GenAI - [[ai-generated-feedback-higher-ed]] — AI-generated feedback in higher education - [[feedback-futures-genai]] — Feedback futures with GenAI - [[ai-feedback-ecosystem-higher-education-2026]] — AI reworks the relations among students, educators, peers and materials in the feedback ecosystem (Bearman et al. 2026) - [[ai-feedback-critical-thinking-writing-2026]] — AI feedback and critical thinking in writing - [[repeated-ai-writing-feedback-semester]] — Repeated AI writing feedback across a semester - [[student-perspectives-ai-writing-grading-2026]] — Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback --- ## [AI Feedback Quality](https://edtechdev.github.io/aied/concepts/ai-feedback-quality/) > **AI feedback quality** — the accuracy, usefulness, timeliness, and [[pedagogy|pedagogical]] value of feedback generated by AI systems for learners. As AI-generated feedback becomes ubiquitous in education, understanding what makes feedback effective — and when it falls short — is critical to ensuring AI supports rather than undermines learning. ## Questions to Consider - What makes feedback 'good' — accuracy alone, or also being timely, specific, actionable, and calibrated to what you already know? Which of these would you notice missing first? - [[research-methods-aied|Research]] finds students experience AI-generated feedback as comparable to a teacher's — but acceptability does not guarantee [[learning-gains|learning effectiveness]]. Why might feedback that feels fine still fail to help you improve? - Even high-quality AI feedback is inert without a 'feedback-literate' recipient — studies show low feedback literacy can make AI feedback minimally useful or even negative. Whose responsibility is it to build that literacy? - Feedback quality and revision depth are linked: scaffolding how students engage with AI feedback shifts them toward argument-level improvement rather than surface edits. How does the way you receive feedback change what you do with it? - Sycophantic feedback conflates support with agreement — an AI that validates your answer rather than challenging it undermines feedback's corrective function. When does feedback need to challenge you rather than comfort you? - Confidence calibration matters: a system that knows when it's uncertain gives better feedback than one that's confidently wrong. How should a tool signal its uncertainty to you, and how would you use that signal? ## Introduction AI feedback quality is not simply about correctness. Effective feedback must be timely, specific, actionable, and calibrated to the learner's current understanding. Research in this knowledge base examines AI feedback quality across multiple dimensions: accuracy (is the feedback correct?), usefulness (does it help the student improve?), and pedagogical alignment (does it promote learning rather than just task completion?). ### How AI feedback quality appears in the research - **Comparability to human feedback:** [[ai-generated-feedback-higher-ed|Studies in higher education]] find that AI-generated feedback is experienced as acceptable and supportive — comparable to teacher feedback. But acceptability does not guarantee learning effectiveness. A PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 empirical studies (2023–2025) likewise finds [[llm]] grading and feedback quality to be task-contingent — matching human raters on short, well-structured answers with detailed rubrics but degrading on complex, open-ended, or [[multilingual-learning|multilingual]] work, with feedback sometimes too generic or misaligned with the grade — and identifies prompt quality, rubric detail, model version, and assessment language as the dominant determinants of grading and feedback quality ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). - **Feedback classification [[benchmark|benchmarks]]:** [[teaching-feedback-classification-benchmark|Cross-language feedback benchmarks]] assess whether feedback quality classification transfers across languages and educational contexts, connecting to [[ai-ed-evaluation]]. - **Collaborative feedback systems:** [[becerra-aicofe-feedback-2026|AICoFE]] implements and deploys AI-based collaborative feedback in [[higher-ed|higher education]], evaluating both system performance and student reception. - **Automated grading feedback:** [[automated-assessment|Automated Grading]] and [[formative-assessment]] research examine whether AI-scored [[assessment|assessments]] provide feedback that matches or exceeds human grading quality. - **Essay scoring feedback:** [[cong-confidence-asag-2026|Confidence-aware ASAG]] and [[choi-anchor-aes-prompting-2025|anchor-based AES]] explore how confidence calibration and [[prompt-engineering|prompting]] design affect feedback quality for writing assessment. - **Discretionary feedback provision:** [[ai-assistance-discretionary-feedback|Research on AI-assisted feedback in higher education]] examines whether AI increases the quantity and quality of feedback instructors provide. - **Reliable provision, not inherent superiority, drives AI's edge:** [[gpt4-feedback-student-activation-2026|Geschwind et al. (2026)]]'s semester-long field experiment found GPT-4 feedback beat [[peer-assessment|peer feedback]] because it was *consistently* delivered — nearly all AIF respondents (369/398) received textual plus numeric feedback while under two-thirds of peer-feedback students received none, and much peer feedback was non-targeted praise. When high-quality textual peer feedback *was* received, peer outcomes matched AI's — indicating the AI advantage is reliability, not quality at the margin. Students also rated peer feedback slightly higher on perceived validity and emotion (mild [[trust|algorithm aversion]]) yet still activated and learned more from AI, showing perceived quality and behavioral outcomes can diverge. - **Source label versus feedback quality:** [[perceptions-teacher-vs-ai-feedback-bias-2026|Mertens et al. (2026)]] separate the *source* of feedback from its quality by holding content constant — 401 teachers rated feedback messages that were all generated with GPT-4-turbo but randomly labeled as teacher- or [[generative-ai|ChatGPT]]-generated, and the label alone shifted credibility (b = 0.21), usefulness (b = 0.47) and fairness (b = 0.48, all p < .001), with 73% preferring the teacher-labeled message, t(400) = 15.55, d = 0.78. Provider-directed items moved furthest — perceived effort (b = 1.34) and willingness to rely (b = 1.45) — and a written explanation of how ChatGPT works changed nothing (interaction ps ≥ .399), pointing to [[trust]], [[teacher-role|professional identity]] and [[bias-mitigation|ingroup bias]] in the rater rather than any deficiency in the feedback itself. Perceived quality can therefore be discounted by provenance even when the content is identical, which means a genuinely good tool can still stall at the point of use. - **Feedback literacy as the uptake-side boundary:** [[liu-deris-ai-feedback-literacy-uptake|Liu & Deris (2025)]] validate an **AI Feedback Literacy (AIFL) scale** (16 items, Attitudes/Practices factors) that predicts students' actual uptake of AI feedback, and [[mendoza-ai-feedback-feedback-literacy-srl|Mendoza et al. (2026)]] show [[feedback-literacy|feedback literacy]] moderates whether AI feedback improves [[self-regulated-learning]] — high literacy yields benefit, low literacy yields minimal or even negative effects. Feedback quality and feedback literacy are two sides of one system: even high-quality AI feedback is inert without a literate recipient. - **Feedback quality drives revision depth:** [[rethinking-ai-writing-feedback-literacy|Feedback literacy scripts]] and [[feedback-literacy-scripts-eap-writing|second-rater mechanisms]] show that [[scaffolding]] how students engage with AI feedback shifts revision toward argument-level improvement rather than surface edits — feedback that promotes deeper revision is higher-quality feedback. - **Sycophancy as a feedback failure:** [[ai-sycophancy|Sycophantic]] feedback conflates support with agreement — an AI that validates a student's answer rather than challenging it undermines feedback's corrective function, degrading quality even when it feels affirming. [[eduframetrap-llm-sycophancy-educational-safety|EduFrameTrap]] shows [[intelligent-tutoring|tutors]] withholding corrective feedback under social pressure, and [[contextual-sycophancy-ai-literacy|contextual sycophancy]] propagates errors into subsequent advice; effective feedback must sometimes challenge the learner. - **Diagnostic accuracy only partly determines feedback quality — and LLM self-evaluation misaligns with [[human-in-the-loop-ai|human judgment]]:** [[reddig-maclellan-personalized-feedback-llm-2026|Reddig, Arora & MacLellan (2025)]] found GPT-4 produced error-targeted hints ~66% of the time in a College Algebra tutor, yet ~35% were too general, incorrect, or gave away the answer; even when diagnosis was wrong the model often recovered with relevant, general-but-correct feedback, though almost all incorrect feedback followed a misdiagnosis. Their simulated-student automated quality checks passed only 21.4% of hints on both tests, rejected targeted feedback ~70% of the time, and favored hints that simply revealed the answer — a stark demonstration that automated [[ai-ed-evaluation|evaluation]] can diverge sharply from human judgments of helpfulness and must be calibrated against them. - **Linguistic and perceptual quality of AI instructional comments:** [[wang-chatgpt-comments-video-learning-scaffolding-2026|Wang, Du and Jin (2026)]] assess ChatGPT-generated in-video scaffolding comments against instructor comments on part-of-speech composition, 3-gram diversity, Zipf's law conformity, readability, topical relevance (TF-IDF and BERTScore), and learner ratings. The generated comments were *more* complex and adjective-rich but *less* varied and less readable, and trailed human comments on topical alignment (0.607 vs. 0.747 BERTScore for knowledge support) and on perceived timing and helpfulness — evidence that AI feedback quality must be judged on linguistic [[accessibility]] and [[affective-computing|affective]] fit, not relevance alone. Their analytic battery is offered as a reusable [[learning-analytics]] pipeline for auditing AI-generated instructional content. - **Prompt design and model choice as measured predictors of quality:** [[teacher-ai-literacy-prompt-feedback-quality-2026|Jacobsen et al. (2026)]] decompose the sources of AI feedback quality with hierarchical regression on feedback generated for 153 pre-service teachers' lesson-planning goals. Across 240 feedbacks from three models under four systematically varied prompts, the model alone explained 26.9% of the variance in nine-category quality ratings and adding the prompt lifted the model to 42.8% (ΔR² = 15.9%); in a replication with the strongest model-prompt combinations (345 feedbacks) the model explained 18.4% and the prompt a further 5.7%. The largest single prompt effect was negative: replacing domain-specific technical terminology with everyday paraphrases lowered feedback quality significantly (β = −0.412), while adding concrete examples and removing the chain-of-thought instruction were not significant in the first study. Quality is therefore not a fixed property of "the AI" — it is jointly produced by which model is chosen and how the task is phrased, and both are teachable. ### Quality dimensions AI feedback quality spans multiple dimensions captured in the knowledge base: - **Accuracy:** Does the feedback correctly identify errors and strengths? ([[automated-assessment|Automated Grading]], [[automated-essay-scoring]]) - **Retrievability of evidence:** Can the system actually reach the evidence it is judging? [[ai-assisted-physics-lab-report-assessment-2026|Abreu, Stari and Martí (2026)]] separate evidence present in a submission from evidence available after processing — an equation, graph or unit may be included in a report yet never retrieved, so feedback about that criterion rests on nothing — making retrieval a dimension of quality distinct from accuracy or calibration. Their response is to require each score to cite concrete evidence from the report, so an observation that cannot be traced back to the text is visible as unsupported. - **Helpfulness:** Does the feedback guide improvement? ([[feedback|Feedback Loop]], [[becerra-aicofe-feedback-2026]]) - **Timeliness:** Is feedback delivered when the learner can act on it? ([[formative-assessment]]) - **Bias:** Is feedback equitable across student populations? ([[bias-mitigation]], [[equity-in-ai-education]]) - **Calibration:** Does the system know when it's uncertain? ([[automated-assessment|Confidence Aware AI Assessment]]) ### Connection to broader concepts AI feedback quality connects fundamentally to [[formative-assessment]] and [[feedback|Feedback Loop]] — quality feedback closes the gap between current and desired performance. It intersects with [[automated-assessment|Automated Grading]] (which generates the scores feedback is based on), [[ai-literacy]] (students must evaluate feedback quality critically), and [[cognitive-offloading|Over-Reliance]] (uncritical acceptance of AI feedback can displace learning). For [[writing-education]], feedback quality is particularly consequential given AI's growing role in writing assessment. ## Connected Concepts - [[formative-assessment]] - [[automated-assessment]] - [[feedback]] - [[ai-literacy]] - [[cognitive-offloading]] - [[bias-mitigation]] - [[automated-essay-scoring]] - [[assessment-validity]] - [[writing-education]] - [[higher-ed]] - [[teacher-role]] - [[feedback-literacy]] - [[ai-sycophancy]] - [[trust-calibration]] ## Connected Articles - [[wang-chatgpt-comments-video-learning-scaffolding-2026]] — Assessing ChatGPT in-video comments: linguistic, semantic, and perceptual quality (Wang, Du & Jin 2026) - [[luo-eaton-ai-student-feedback-ethics-2026]] - [[ai-vs-human-assessment-efl-tpck-2026]] — AI-generated vs human-developed assessment tasks in EFL - [[coach-not-crutch-ai-writing]] — AI writing feedback outperformed human editors on practice letters (Lira et al. 2025) - [[zhao-learnlens-feedback-educators-loop]] — LearnLens: LLM feedback generation with educators in the loop (Zhao et al. 2025) - [[richmond-nicholls-genai-psych-feedback-ai-literacies]] — Critiquing ChatGPT output against a rubric builds feedback literacy (Richmond & Nicholls 2025) - [[yasir-llm-tutoring-agents-2026]] — LLM tutoring feedback: accurate diagnosis ≠ actionable feedback (Yasir et al. 2026) - [[melo-llm-classroom-observation-teach-2026]] — LLM classroom observation feedback reliability and limits (Melo et al. 2026) - [[learner-centered-feedback-ai]] — Teachers' practices and perceptions of AI learner-centered feedback (PolyFeed) - [[ai-generated-feedback-higher-ed]] — AI-Generated Feedback in Higher Education - [[teaching-feedback-classification-benchmark]] — Teaching Feedback Classification Benchmark - [[becerra-aicofe-feedback-2026]] — AICoFE: AI-Powered Collaborative Feedback - [[cong-confidence-asag-2026]] — Confidence-Aware Short Answer Grading - [[choi-anchor-aes-prompting-2025]] — Anchor-Based AES Prompting - [[ai-assistance-discretionary-feedback]] — AI Assistance for Discretionary Feedback - [[sequenced-ai-feedback-learning]] — Sequenced AI Feedback and Learning - [[eduframetrap-llm-sycophancy-educational-safety]] — Sycophancy is an educational safety risk: Why LLM tutors need sycophancy benchmarks - [[contextual-sycophancy-ai-literacy]] — The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention - [[genai-educational-outcomes-meta-analysis]] - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Marked Pedagogies: bias in automated writing feedback - [[can-ai-evaluate-assessment-llm-meta-assessment-2026]] - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[gpt4-feedback-student-activation-2026]] - [[reddig-maclellan-personalized-feedback-llm-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] - [[teacher-ai-literacy-prompt-feedback-quality-2026]] — Prompt engineering and model selection as predictors of AI-feedback quality (Jacobsen et al. 2026) - [[perceptions-teacher-vs-ai-feedback-bias-2026]] — Randomized source labels on identical GPT-4 feedback: teachers discount AI-attributed feedback (Mertens et al. 2026) - [[peer-group-vs-ai-feedback-2026]] — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education - [[ai-assisted-physics-lab-report-assessment-2026]] — AI-Assisted Assessment of Experimental Physics Laboratory Reports: Potential, Limitations, and Support for Teaching Practice --- ## [Formative Assessment](https://edtechdev.github.io/aied/concepts/formative-assessment/) > **Formative assessment** — assessment designed to inform ongoing instruction and learning, as opposed to [[summative-assessment|summative]] evaluation. In AI education, formative assessment is both transformed by AI and essential to it: AI systems can generate, validate, and adapt formative items and feedback at scale, while formative feedback is a primary mechanism through which [[intelligent-tutoring|AI tutors]] and [[adaptive-learning|adaptive systems]] support learning. The knowledge base's [[research-methods-aied|research]] examines AI-generated formative items, AI-generated feedback, and the design and evaluation of these systems. ## Questions to Consider - Formative assessment is meant to close the loop — surface what students don't know and give feedback they can act on *while learning is still in progress*. How is that fundamentally different from summative evaluation, and when might the two get confused in practice? - AI can generate multiple-choice questions with impressive accuracy on verifiable dimensions, but is weakest on instructional-judgment dimensions. If a machine is good at correct answers but weaker at [[pedagogy|pedagogical]] judgment, what should it be trusted to do — and what should humans keep doing? - AI-generated feedback only helps when students actually enact it — the 'enacted feedback' condition, where students select, evaluate, and apply suggestions, outperformed simply being given feedback. If enactment matters more than the feedback itself, what does that mean for how formative feedback should be designed? - Feedback is described not as information transfer but as an [[ethics|ethical]], relational practice. What gets lost when formative assessment is mass-produced as 'AI slop' — and what can human comment banks and relational care preserve? ## Introduction Formative assessment is central to [[ai-education|AI in education]] because it sits at the junction of [[assessment]] and [[feedback|learning feedback]]. Its purpose is to close the loop: surface what students know and don't know, and provide feedback they can act on to improve. AI makes this feasible at scale — generating items, scoring responses, and delivering individualized feedback — but the knowledge base's research shows that quality varies dramatically across item types, and that feedback only helps when students actually enact it. ## AI-generated formative items AI systems generate formative assessment items across modalities, with reliability varying by type: - **Multiple-choice questions:** [[code-gen]] shows [[agentic-ai|agentic AI]] can reliably generate MCQs for coding comprehension when validated across seven pedagogical dimensions — success rates reach **98.6%** for concept alignment and **79.9%** for feedback quality — suggesting AI is strongest on verifiable dimensions and weakest on instructional-judgment dimensions. This connects to [[automated-question-generation|automated question generation]] more broadly. - **Automated essay scoring:** multi-agent frameworks (e.g., MASS) improve consistency over stand-alone [[llm|LLMs]] for [[automated-essay-scoring|essay scoring]], though [[explainable-ai|interpretability]] of multi-agent scoring decisions remains an open challenge. - **Formative scoring pipelines:** [[cotal-formative-assessment-scoring-2026|CoTAL]] couples Chain-of-Thought prompting with [[active-learning|active learning]] and Evidence-Centered Design to produce generalizable formative-assessment scoring with human-in-the-loop [[prompt-engineering|prompt engineering]]. - **High-frequency, [[automated-assessment|automatically-marked assessments]]:** [[automated-formative-assessments-a-level-sciences|automated formative assessments in A-level sciences]] examines the effect of high-frequency, automatically-marked formative assessment on [[learning-gains|learning outcomes]]. A scoping review of short-answer auto-marking in [[science-education|science]] (2017–early 2024) confirms this formative short-answer use case is a field with real traction: BERT-family models dominated auto-marking through 2021 before prompting larger [[llm|LLMs]] from ~2022, and domain-augmented, rubric-aware, and chain-of-thought models performed best — yet the review's calls for comprehensive evaluation and unresolved [[bias-mitigation|fairness]] and explainability gaps caution against extending such systems to unmediated [[summative-assessment|summative]] or high-stakes use ([[auto-marking-short-answer-science-2026]]). ## AI-generated feedback A large body of knowledge base research examines AI-generated formative feedback: - **The enactment problem:** [[ai-feedback-enactment-workflow-2026|Making AI-Generated Feedback Matter]] (13,037 students; 51,296 resources) shows feedback value depends on whether students *enact* it — the **Enacted Feedback** condition, where students select, evaluate, and apply AI feedback suggestions, outperformed simple directed feedback. - **Feedback is not information transfer:** [[care-full-feedback-genai|The care-full craft of feedback]] argues feedback is an ethical, relational practice, not information transmission — feedback only constitutes feedback when students make sense of and act on it, and contrasts mass-produced "AI slop" with human comment-bank shortcuts. - **Sequenced feedback can backfire:** [[sequenced-ai-feedback-learning|Sequenced AI feedback]] (encouragement → hints → correct answer, designed to promote autonomy) actually **harmed learning** in a randomized experiment (N=199) despite boosting [[student-engagement|engagement]] and positive perceptions — a cautionary finding about feedback design. - **Learner-centered tools:** [[learner-centered-feedback-ai|PolyFeed]] combines ML suggestion models with [[teacher-role|teacher]] practice, showing how teachers adopt and adapt AI feedback suggestions; [[ai-internal-feedback-evaluative-judgments|AI-supported internal feedback]] helps undergraduates develop [[evaluative-judgment|evaluative judgment]]. - **Rubric-guided prompting and role-aware feedback:** [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] showed that iterative rubric co-refinement — clarifying performance descriptors and explicitly accepting implicit indicators of learning — drove LLM–human agreement on student design work from 54.75% to 81.25% (Cronbach's Alpha rising from 0.393 to 0.798), with the largest gains in the cognitively demanding Iteration & Reflection category. Prompting the same model under instructor, peer-reviewer, and grant-reviewer roles produced distinct evaluative feedback, and post-revision LLMs were more consistent than some human raters in applying performance thresholds — positioning rubric-guided LLMs as calibration and co-design partners in formative feedback environments, with human-in-the-loop oversight remaining essential. - **AI feedback sustains participation and drives gains at scale:** [[gpt4-feedback-student-activation-2026|Geschwind et al. (2026)]]'s semester-long field experiment across undergraduate tutorials found that individual GPT-4 formative feedback (spanning all three Hattie & Timperley dimensions — Feed-Back, Feed-Up, Feed-Forward) sustained the highest participation across eight open-ended tasks, lengthened student answers, and produced the strongest content learning gains — an effect driven by reliable, consistent AI provision, since when high-quality textual peer feedback was actually received, peer outcomes matched AI's. - **Feedback futures:** [[feedback-futures-genai|Feedback Futures]] synthesizes a special issue and argues the question is not *whether* [[generative-ai|GenAI]] can produce feedback but how to design feedback that supports learning, distilling recurring tensions across the field. - **Diagnosis-first feedback for open-ended quantitative problems:** [[yin-arthur-ai-teaching-assistant-engineering-econ-2026|Arthur (Yin et al. 2026)]] delivers real-time, personalized formative feedback on Engineering Economics Calculated Formula Questions, a domain where handwritten, unstructured solutions had previously blocked AI support. A per-question [[machine-learning|XGBoost]] backbone diagnoses likely rubric-labeled mistakes from students' submitted numerical answers (average precision 0.81, recall 0.79), and a dialogue-based scheme requests intermediate answers only when prediction confidence is low — balancing feedback accuracy against collection efficiency within a question-bank web interface. - **Scale and limits of LLM formative feedback (systematic evidence):** a PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 empirical studies (2023–2025) finds LLMs can reduce teacher workload and deliver rapid, personalized feedback at scale — especially in large or [[higher-ed|higher-education]] cohorts — but that feedback is sometimes too generic or misaligned with the assigned grade and reliability slips on longer, [[multilingual-learning|multilingual]], or nuanced tasks, reinforcing that formative AI feedback is best deployed under educator oversight ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). - **Adaptivity is a separable ingredient, not decoration:** [[ai-feedback-adaptivity-children-plans-2026|Sukjaitham, Schaaf, Brod & Breitwieser (2026)]] supply the direct causal test that most LLM-feedback studies assume away, pitting GPT-4 response-contingent feedback against expert-written generic guidance matched on structure, tone, length, and [[motivation|motivational]] phrasing (verified with a five-dimension quality rubric, κ = .76–1.00). In a preregistered within-subjects experiment, 155 German fifth- and sixth-graders (M = 12.08 years) revised six if-then plans: plan quality rose from a median of **2 → 5** under adaptive feedback versus **2 → 3** under generic guidance (within-person V = 10,440, p < .001, r = .86; condition × time interaction estimate = 1.68, SE = 0.14, p < .001, with no pre-support difference). Children rated adaptive feedback both more helpful (r = .67) and more motivating (r = .74), and trial-level perceptions predicted the size of revision gains — making [[technology-acceptance-model|perceived usefulness]] part of the pathway rather than an affective byproduct. Because the control was itself well designed, the study shows added value *beyond* good non-contingent guidance rather than the difference between feedback and nothing: generic guidance is a genuine but limited substitute that plateaus at its own median. Planning served as the test case as a core [[self-regulated-learning|self-regulated learning]] strategy with explicit quality criteria, which makes contingency — the property that distinguishes [[scaffolding]] from static support — directly measurable on a one-sentence response. The authors define adaptivity narrowly as response-contingent adaptation of feedback content to the learner's concrete response, distinguishing it from conversational interactivity, tone, and stable-trait [[adaptive-learning|adaptive learning]]. - **Perceived usefulness tracks actionability:** [[mendonca-llm-feedback-perceived-usefulness-programming-2026|Mendonça et al. (2026)]] had 144 programming students rate 893 LLM-generated feedback instances on five dimensions and 237 consolidated reports on six, holding domain, task, and instrument constant while educational level varied. Ratings were favorable throughout, with student-level means of 4.24 to 4.43 for individual answers and 4.11 to 4.38 for reports, yet actionability and usefulness drew the lowest ratings even as actionability and perceived accuracy carried the largest relative weights (32.6% and 29.0%) in a model explaining 76% of the variance in perceived usefulness, and motivation and personalization led a report-level model of intention to use explaining 53%. Since actionability is the dimension [[feedback]] research treats as hardest to provide, the pattern reads as [[technology-acceptance-model|technology acceptance]] applied to feedback: usefulness tracks whether a learner can act, not how polished the feedback sounds. ## Curriculum-grounded and educator-in-the-loop design [[ai-learning-tools-engineering-education-needs|LearnLens]] addresses three persistent problems in AI formative assessment: **error-aware assessment** (capturing nuanced reasoning errors rather than surface mistakes), **topic-linked memory chains** (replacing noisy similarity-based [[rag]] with structured [[curriculum-design|curriculum]]-grounded retrieval), and **educator-in-the-loop** design (teacher customization and oversight, not full automation). This connects to the broader tension in [[human-in-the-loop-ai]]: scalable automation with expert validation. [[hoppe-teachers-diagnostic-skills-ai-formative-assessment-2026|Hoppe, Loibl & Leuders (2026)]] sharpen what educator-in-the-loop means once a tool produces its own claims rather than raw observations. Their conceptual analysis argues that AI-generated diagnostic inferences are qualitatively different evidence, because they are already the product of algorithmic interpretation, so teachers need a further layer the authors call **meta-diagnosis**: deliberately accepting, rejecting, or modifying an inference and integrating it with their own contextual knowledge. Specifying the DiaCoM framework for this treats AI-generated inferences as a situation characteristic and the accept, reject, or modify decision as diagnostic behavior, while expanding the person characteristics teachers need to include knowledge of how AI systems actually work. Because current systems rest mostly on performance data such as task correctness and completion time, motivational states and classroom dynamics stay largely absent, so the paper keeps teachers, not the dashboard, as the responsible reflective agents and frames judging algorithmic claims as a target for professional development. ## Design trade-offs | Dimension | AI Suitability | Human Requirement | |-----------|----------------|-------------------| | Factual correctness | High | Low | | Concept alignment | High | Medium | | Distractor quality | Low | High | | Feedback depth | Low | High | | Rubric consistency | Medium | Medium | ## Assessment, feedback, and learning Formative assessment in AI education connects to the learning process itself: - **Feedback loops:** [[feedback|feedback loops]] are the mechanism by which formative assessment informs learning; AI tutors and adaptive systems close these loops at scale. - **Self-regulated learning:** formative feedback supports [[self-regulated-learning|self-regulated learning]] when students monitor progress and adjust; AI feedback should cultivate [[ai-internal-feedback-evaluative-judgments|evaluative judgment]], not displace it. - **Automated scoring reaches some phases of the cycle, not all:** [[chen-automated-scoring-interpreting-self-regulated-learning-2026|Chen & Liu (2026)]] ran a 14-week quasi-experiment in which 46 interpreting students submitted weekly renditions to an automated scoring system that returned an immediate score, transcript, marked errors, and a reference rendition, while a control group received whole-class teacher feedback only. The automated group improved more overall (d = 1.03), but the gain stayed where the deficit was decomposable, the signal reliable, and the scale sensitive: linguistic accuracy and logical coherence rose while information fidelity (agreement with human raters r = 0.12) and delivery fluency did not move. [[self-regulated-learning|Self-regulated learning]] was uneven in a matching way, with execution and monitoring correlating with score gains (r = 0.42) while planning and emotional motivation sat near the scale midpoint, and several students deferred engagement after low scores instead of analyzing causes. The system reached the performance phase of the cycle, not the planning that precedes it. - **From automated diagnosis to generated practice:** [[zhu-adaptive-teaching-assistance-genai-big-data-2026|Zhu, Luo & Li (2026)]] wire audio-score alignment, error quantification, and a Proximal Policy Optimization layer into one closed loop for music education, turning rhythm error signals (91.2% recall, 98.4% specificity) into rewards that steer generated practice tracks, and report a significant Group x Time interaction favoring the system group (beta = 0.52) across a 12 week quasi-experiment with 120 undergraduate music majors. The authors frame this as technical feasibility, and the practical reading is triage rather than assessment: precision of 89.7% means roughly one flagged rhythm error in ten is a false alarm, and the study's expert reviewers rated support for musical expression as the weakest area, so automated loops suit technical drills while expressive judgment stays with teachers. - **Scaffolding:** [[scaffolding]] and formative assessment work together — AI can provide just-in-time hints and prompts, though sequenced feedback research cautions against over-structuring. - **Validity and quality:** the [[ai-feedback-quality|quality]] and [[assessment-validity|validity]] of AI-generated formative items and feedback must be evaluated; [[ai-ed-evaluation]] provides the methods. ## Risk: Assessment as surveillance Formative assessment systems can shift from learning-support tools to behavior-monitoring infrastructure. The same data streams that enable adaptive tutoring can enable punitive tracking if [[governance]] is weak. This connects to [[privacy]] and [[well-being|student well-being]], and argues for formative systems that support learning rather than surveil it. ## Implications for AI in education - **Match item type to AI reliability:** use AI for verifiable dimensions (concept alignment, correctness) and retain human judgment for instructional dimensions (distractor quality, feedback depth). - **Design for enactment, not just provision:** AI feedback only helps when students select, evaluate, and apply it — structure workflows that support enactment. - **Feedback design matters more than volume:** sequenced or over-structured feedback can backfire; prioritize feedback that supports student sense-making and autonomy. - **Keep educators in the loop:** curriculum-grounded, educator-in-the-loop systems improve relevance and reduce noise. - **Evaluate quality and validity:** assess AI-generated items and feedback for quality, validity, and [[equity-in-ai-education|equity]], not just generation speed. - **Treat formative assessment as a philosophy, not a toolkit.** [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] synthesize the wider higher-education literature into an "assessment for learning" paradigm that frames formative, ongoing, and individualized [[feedback]] as an overarching philosophy rather than a mere toolkit — balancing formative with [[summative-assessment|summative]] purposes and foregrounding [[agency|student agency]], self-regulation, and [[metacognition|metacognitive]] skill. They identify five mutually reinforcing practices ([[authentic-assessment|authentic assessment]], self- and [[peer-assessment]], reassessment, [[mastery-learning|standards-based grading]], ungrading) through which this philosophy can be enacted, while noting their uptake remains highly uneven across higher-education fields. ## Connected Concepts - [[assessment]] - [[educational-measurement]] - [[automated-assessment]] - [[automated-question-generation]] - [[assessment-validity]] - [[feedback]] - [[ai-feedback-quality]] - [[feedback-literacy]] - [[self-assessment]] - [[learning-analytics]] - [[personalized-learning]] - [[adaptive-learning]] - [[scaffolding]] - [[self-regulated-learning]] - [[human-in-the-loop-ai]] - [[intelligent-tutoring]] - [[ai-ed-evaluation]] - [[summative-assessment]] — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams) ## Connected Articles - [[ssail-safe-sound-ai-learning-2026]] — SSAIL: A Design Framework for Safe and Sound AI for Learning - [[causal-modeling-competency-assessment-2026]] — Causal Modeling of Support Interventions for Student Competency Assessment - [[nicola-richmond-programwide-assessment-genai-2025]] — Program-wide approaches to redesigning assessment in the GenAI era - [[ni-lam-multiliteracies-ai-portfolio-2026]] — Students' perceptions of multiliteracies development with AI-assisted portfolio assessment - [[ai-feedback-enactment-workflow-2026]] — Making AI-generated feedback matter: from provision to enactment - [[care-full-feedback-genai]] — The care-full craft of feedback in an age of GenAI - [[feedback-futures-genai]] — Feedback futures: beyond the limits of human and GenAI capacities - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans (Mathew et al. 2026) - [[sequenced-ai-feedback-learning]] — Impact and pathways of sequenced AI feedback - [[learner-centered-feedback-ai]] — Enhancing learner-centered feedback with AI - [[ai-internal-feedback-evaluative-judgments]] — Developing evaluative judgments through AI-supported internal feedback - [[cotal-formative-assessment-scoring-2026]] — CoTAL: formative assessment scoring with human-in-the-loop prompting - [[automated-formative-assessments-a-level-sciences]] — High-frequency automated formative assessment - [[ai-generated-feedback-higher-ed]] — AI-generated feedback in higher education - [[ai-learning-tools-engineering-education-needs]] — LearnLens: curriculum-grounded AI feedback - [[genai-teacher-feedback-comparison]] — GenAI vs. teacher feedback comparison - [[chatgpt-feedback-engagement-genai]] — ChatGPT feedback and engagement - [[becerra-aicofe-feedback-2026]] — AI-coffee feedback framework - [[code-gen]] — CODE-GEN: validated MCQ generation - [[responsible-assessment-ai-era-stanford-2026]] — Responsible assessment in the AI era - [[zhan-boud-du-authentic-assessment-scoping-review-2025]] — Designing for authentic assessment - [[automated-grading-linux-bash-examinations-large-language-models]] — Automated grading of Linux/bash exams - [[instructor-ai-roles-chatgpt-formative-assessment-2026]] — Instructor and AI roles in ChatGPT-enhanced formative assessment - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant assessment - [[roe-assessment-twins-2026]] — Assessment twins for strengthening assessment validity in the age of GenAI (Roe, Perkins & Giray 2026) - [[harmogen-ai-assessment-rubric-generation]] — HARMOGEN-R: AI assessment rubric generation - [[ai-assisted-instructor-supervised-grading-feedback]] — AI-assisted instructor-supervised grading and feedback - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[assessing-student-drive-framework-2025]] — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise) - [[chatgpt-qiskit-homework-autogradable-2026]] — ChatGPT solves Qiskit homework; autogradable design - [[llm-adaptive-programming-error-explanations-2026]] — LLM adaptive explanations of programming errors - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design - [[auto-marking-short-answer-science-2026]] - [[gpt4-feedback-student-activation-2026]] - [[yin-arthur-ai-teaching-assistant-engineering-econ-2026]] - [[mesny-innovative-assessment-grading-management-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] - [[ai-feedback-adaptivity-children-plans-2026]] — Adaptivity makes feedback effective: evidence from AI-generated feedback on children's plans - [[peer-group-vs-ai-feedback-2026]] — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education - [[chen-automated-scoring-interpreting-self-regulated-learning-2026]] — Automated scoring, interpreting performance, and self-regulated learning (Chen & Liu 2026) - [[hoppe-teachers-diagnostic-skills-ai-formative-assessment-2026]] — Teachers' diagnostic skills in AI-supported formative assessment: from diagnosis to meta-diagnosis - [[mendonca-llm-feedback-perceived-usefulness-programming-2026]] — Perceived usefulness and intention to use LLM-generated feedback in programming across three educational levels - [[zhu-adaptive-teaching-assistance-genai-big-data-2026]] — Adaptive teaching assistance combining generative AI and big data analytics in music education --- ## [Summative Assessment](https://edtechdev.github.io/aied/concepts/summative-assessment/) > **Summative assessment** — assessment used to evaluate and certify what a learner has learned at the end of a unit, course, or program, in contrast to [[formative-assessment|formative assessment]] which supports learning during instruction. Summative assessment typically takes the form of high-stakes examinations — written, oral, proctored, or closed-book — that assign grades, gate progression, and certify competence. In the AI era, summative assessment has become a central battleground over [[academic-integrity]] and validity: [[generative-ai|generative AI]] can inflate performance on unproctored or take-home tasks, making the choice of summative format — and how it resists AI substitution — a pivotal design decision. ## Questions to Consider - Summative assessment certifies what a student has learned at the end of a course, while formative assessment supports learning during it. Where have you seen the line between these two blur, and why might it matter that they serve different functions? - The page frames generative AI as reshaping summative assessment in two directions at once: AI scores exams, and students use AI to evade exam-based measurement. Which of these two pressures do you think is the bigger threat to validity, and why? - If unproctored or take-home tasks lose validity because AI can produce the answers, what does that imply for how assessments should be designed — and what might be sacrificed in the process? - [[research-methods-aied|Research]] cited on the page finds LLMs do not grade essays the same way humans do. If automated scoring is fast and consistent but grades differently, is that a [[bias-mitigation|fairness]] problem, an opportunity, or both? - What does a high-stakes result (a grade, a credential, admission) mean if the work behind it could have been produced by AI? How would you design an assessment you could actually trust? ## Introduction Summative assessment serves a fundamentally different function from formative assessment: it measures and certifies achievement rather than guiding next steps. It includes end-of-unit tests, final examinations, standardized and high-stakes tests (e.g., entrance exams), oral defenses, and cumulative performance assessments. Because summative results carry real consequences (grades, progression, credentials, university admission), they face particular pressures in the AI era — both as *targets* of automated scoring and as *vulnerable* measures that students may seek to game using generative AI. ## The AI-era stakes: validity and integrity The knowledge base's research documents how generative AI has fundamentally reshaped the summative-assessment landscape in two directions: AI is used to **score** exams at scale, and AI can be used by students to **evade** exam-based measurement of their own learning. - **AI as scorer.** Summative assessment increasingly relies on [[automated-assessment|automated scoring]] of exams, essays, and short answers. [[llms-do-not-grade-essays-like-humans-2026|Research on LLM essay grading]] finds [[llm|LLM]] do not grade essays the same way humans do, raising validity and fairness questions for high-stakes automated scoring. [[llm-automated-assessment-student-self-explanations|LLMs assessing student self-explanations]] and [[cong-confidence-asag-2026|automatic short-answer grading]] explore the reliability of LLM scoring in summative contexts, while [[psyscore-essay-scoring-zpd-feedback|psychometrically-aware frameworks]] seek to keep automated scoring trustworthy and adaptive. A 296-student handwritten general-chemistry exam illustrates why selective oversight is required: a multimodal LLM's total-score agreement with TA grading was high (R² = 0.91) yet item-level reliability varied sharply by format, and false positives (AI crediting genuinely wrong answers) tend to go undetected because students rarely contest them — so a uniform "grade everything" AI scorer is not defensible for high-stakes use without confidence-based deferral to humans ([[cvengros-grading-handwritten-chemistry-ai-2026]]). - **AI as evasion.** Because generative AI can produce answers to written questions, unproctored and take-home summative tasks lose validity: [[generative-ai-reduced-study-time-math|proctored, unassisted measures are essential]] because non-proctored performance is inflated by AI, and [[generative-ai-guardrails-harm-learning|guardrailed (hint-not-answer) tools]] can eliminate the exam penalty that unguarded AI causes. [[chirikov-ai-grade-inflation-2026|Chirikov's (2026)]] quasi-experiment on 500,000+ grades makes the mechanism concrete: after ChatGPT's release, courses with more AI-exposed tasks saw the share of A grades rise by 13 percentage points, and the effect concentrated in **homework-heavy courses** (an additional 16 pp in the triple-differences estimate) — direct evidence that unproctored homework, not genuine [[learning-gains|learning gains]], is where AI inflates summative outcomes. - **AI-generated exams.** [[assessing-quality-ai-generated-exams-field-2025|A large-scale field study]] and [[ai-vs-human-assessment-efl-tpck-2026|EFL assessment research]] examine whether AI can *generate* high-quality exams and assessment tasks — an emerging summative-design use of AI. ## AI-resistant summative formats A key theme in the knowledge base is that **summative format determines AI-resistance** — the more a task requires live, in-person, individually-probed performance, the harder it is for students to substitute AI for their own learning. - **Oral exams and assessments.** [[fenton-oral-exams-ai-authentic-assessment-2025|Fenton (2025)]] argues the oral exam is a low-tech, inherently AI-resistant summative format: its real-time, interactive dialogue tests comprehension, [[critical-thinking|critical thinking]], and reasoning rather than memorization, prevents students from using AI to generate and memorize answers, and mirrors professional practice. [[socratic-tests-conversational-assessment|Socratic tests]] and [[code-review-genai-cs1|code-review interviews]] extend this to dynamic, conversational, and interview-based summative assessment. - **Closed-book, proctored, unassisted measures.** [[generative-ai-reduced-study-time-math|Evidence]] and [[stromberg-generative-ai-learning-penalty-secondary-2026|large-scale field data]] show that proctored closed-book exams — not inflated homework or take-home work — are the reliable signal of actual learning when students use AI. [[responsible-assessment-ai-era-stanford-2026|Responsible assessment]] frameworks embed these unassisted measures within a validity-driven redesign. ## High-stakes and standardized summative assessment High-stakes summative assessment — entrance exams, standardized tests, and certification — carries outsized consequences and is a focus of AI-era concern. [[stromberg-generative-ai-learning-penalty-secondary-2026|The generative AI learning penalty study]] measured outcomes on high-school (Zhongkao) and college (Gaokao) entrance exams, finding entrance-exam scores fell 18–24% after prolonged AI use. [[brcic-effortless-trap-productive-struggle-2026|The Effortless Trap]] and [[genai-performance-vs-learning|performance-vs-learning research]] warn that gains on AI-assisted tasks do not transfer to unassisted high-stakes measures. ## Summative vs. formative in the AI era The knowledge base's assessment literature consistently emphasizes that [[assessment]] is most effective when it combines [[formative-assessment|formative]] and summative functions — but the AI era sharpens the distinction. Because AI inflates performance on low-stakes, unproctored, and process-hidden tasks, **summative (especially proctored/closed-book/in-person) measures become the crucial check** on whether learning actually occurred. This motivates assessment redesign that keeps authentic, AI-resistant summative tasks (oral exams, code-review interviews, proctored examinations, process-based [[eportfolio|portfolios]]) as the anchor of [[academic-integrity|integrity]] while using formative assessment to support learning along the way. See [[authentic-assessment]] for the constructive design response. ## Implications for AI in education - **Summative format is a validity and integrity lever:** AI-resistant summative formats (oral, proctored, closed-book, in-person) preserve the connection between assessed performance and actual learning. - **Proctored/unassisted measures are the reliable signal:** when students use AI, unassisted summative exams — not homework — reveal genuine learning. - **Automated scoring needs psychometric scrutiny:** using LLMs to grade high-stakes exams requires evaluation of reliability, fairness, and validity, not just accuracy. - **AI can also generate exams:** AI-assisted exam and task generation is an emerging summative-design application that itself needs quality evaluation. - **Reconsider grading purpose, not just format.** [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] critique traditional summative, norm-referenced grading for encouraging superficial, fragmented learning, giving students little control or transparency, harming intrinsic [[motivation]], fueling stress and [[well-being|anxiety]], and perpetuating inequities while largely assessing recall rather than real-world application. They position reassessment, [[mastery-learning|standards-based grading]], and ungrading as grading-focused innovations that can soften summative-heavy practice, while acknowledging these remain marginal in management education because of normative barriers — grading on a curve, external signaling (rankings, internships, accreditation), and students' instrumental mindset — and recommend incremental experimentation with institutional support. ## Connected Concepts - [[remote-proctoring]] - [[assessment]] - [[formative-assessment]] - [[authentic-assessment]] - [[automated-assessment]] - [[assessment-validity]] - [[academic-integrity]] - [[ai-ed-evaluation]] - [[higher-ed]] - [[k-12]] ## Connected Articles - [[academic-dishonesty-automated-proctoring-ai-2026]] - [[automated-online-exam-proctoring-decade-review-2026]] - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant summative assessment - [[ivory-psychology-assessment-integrity-2026]] — 90% of psychology assessments passable at minimum effort (Ivory et al. 2026) - [[stromberg-generative-ai-learning-penalty-secondary-2026]] — The generative AI learning penalty: proctored/closed-book exam evidence - [[chirikov-ai-grade-inflation-2026]] — AI task displacement as a mechanism of grade inflation; homework-heavy courses (Chirikov 2026) - [[generative-ai-reduced-study-time-math]] — Faster completion, less learning: proctored measures essential - [[generative-ai-guardrails-harm-learning]] — Generative AI without guardrails harms learning - [[assessing-quality-ai-generated-exams-field-2025]] — Assessing the quality of AI-generated exams - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans - [[llm-automated-assessment-student-self-explanations]] — LLMs for automated assessment of student self-explanations - [[cong-confidence-asag-2026]] — Automatic short-answer grading - [[psyscore-essay-scoring-zpd-feedback]] — Psychometrically-aware trait-adaptive essay scoring - [[socratic-tests-conversational-assessment]] — Socratic tests: dynamic, conversational, multimodal assessment - [[code-review-genai-cs1]] — Code review interviews in CS1 - [[responsible-assessment-ai-era-stanford-2026]] — Responsible assessment in the AI era - [[test-driven-ai-assisted-learning]] — Test-driven AI-assisted learning - [[genai-oop-programming-assessments-2026]] — GenAI performance on object-oriented programming assessments - [[brcic-effortless-trap-productive-struggle-2026]] — The Effortless Trap: productive struggle and the illusion of learning - [[ai-vs-human-assessment-efl-tpck-2026]] — AI-generated versus human-developed assessment tasks in EFL - [[roe-assessment-twins-2026]] — Assessment twins for strengthening assessment validity in the age of GenAI (Roe, Perkins & Giray 2026) - [[ai-grading-handwritten-physics-2026]] — AI grading of handwritten physics assessments (Olympiad) - [[mesny-innovative-assessment-grading-management-2026]] - [[cvengros-grading-handwritten-chemistry-ai-2026]] --- ## [Authentic Assessment](https://edtechdev.github.io/aied/concepts/authentic-assessment/) > **Authentic assessment** — the design of assessments that examine student performance on worthy, realistic intellectual tasks, rather than isolated, standardized test items. Originating with Wiggins (1990) as a counterbalance to standardized tests, authentic assessment has evolved from replicating workplace tasks toward a multi-dimensional framework encompassing professional, digital, personal, and social authenticity. [[generative-ai|Generative AI]] has made authentic assessment newly essential: any task a [[llm|language model]] can credibly simulate in a take-home setting loses its validity as evidence of original student competence, so authentic forms must be redesigned around what AI cannot credibly counterfeit. ## Questions to Consider - Authentic assessment examines performance on worthy, realistic intellectual tasks rather than isolated test items. Before reading, did you assume an 'authentic' task simply replicates a real-world or workplace scenario? This page argues authenticity has evolved well beyond mere replication — what other forms can it take? - A central claim is that any task a language model can credibly simulate in a take-home setting loses its validity as evidence of original student competence. Which assessment in your own teaching or study would you now consider 'counterfeitable' — and what makes you say that? - The page highlights a striking gap: only 3 of 37 studies addressed social authenticity — whether assessment helps students contribute to societal transformation. Why do you think the social dimension of authenticity is so neglected, and what might assessing it actually look like? - The knowledge base argues authenticity must be 'redesigned, not policed.' When students co-author rubrics, choose what and how to submit, and engage in real-time interaction, AI use becomes expected and declared rather than concealed. How would that shift the relationship between student and assessor? - Oral exams and process-based portfolios are offered as AI-resistant authentic forms. But graded reflection and self-assessment risk becoming performative — students performing the reflection they think is wanted. How could you design for genuine reflection rather than performance? - When AI supplies the rubric, the feedback, and the monitoring, the page warns that a student's metacognitive practice may be displaced. If the tool does the evaluating, what is left for the learner to actually practice or learn? ## Introduction Authentic assessment sits at the heart of how [[assessment]] is being rethought in the AI era. It connects to [[assessment-validity]] (does the assessment measure what it claims?), [[formative-assessment]] (authentic tasks that inform learning), and [[academic-integrity]] (moving from detection to designing tasks where AI use is expected and declared). It is a central response in the knowledge base's assessment-redesign literature. ## The evolution of authenticity - **1990s origins — worthy intellectual tasks:** Wiggins (1990) proposed authentic assessment as direct examination of "student performance on worthy intellectual tasks," a counterbalance to standardized tests. - **Late-1990s uptake — workplace replication:** Joughin (1998) framed authenticity as the extent to which assessment replicates professional practice or real life — a view that dominated for two decades. - **2020s critique — beyond replication:** McArthur (2023) argued authentic assessment must enable students to "influence the future and transform society" rather than merely replicate existing tasks; Ajjawi et al. (2024) broadened authenticity to contextual, task, and personal forms that reflect [[student-experience|student experience]]. - **Generative AI as existential challenge:** [[zhan-boud-du-authentic-assessment-scoping-review-2025|Zhan, Boud & Du (2025)]] note that generative AI makes workplace-replication authenticity newly vulnerable, driving a pivot toward [[ai-literacy|digital literacy]], [[collaborative-learning|real-time collaboration]], social contribution, and individual meaning-making that AI cannot credibly counterfeit. ## A six-dimensional design model [[zhan-boud-du-authentic-assessment-scoping-review-2025|Zhan, Boud & Du's (2025) scoping review]] of 37 empirical studies (2000–2024) proposes six design dimensions: (1) authenticity in assessment (assessment, professional, digital, self, and social authenticity — with only 3/37 studies addressing social authenticity, a critical gap), (2) cognitive challenges, (3) assessment criteria (with students often passive recipients rather than co-authors of rubrics), (4) feedback (formative-dominated, but sustainable feedback rare), (5) [[agency|student agency]] (choice in what/how/when/where to submit was rare), and (6) social collaboration. It also proposes a cyclical co-design model — negotiate goals, create context, co-design criteria, plan feedback — that AI tools could operationalize. ## Authentic assessment in the AI era The knowledge base's assessment-redesign literature argues that authenticity must be **redesigned, not policed**: - [[beyond-detection-authentic-assessment-ai-2025|Beyond Detection]] contends that authenticity cannot be policed into existence; it must be designed, positioning AI as a declared collaborator rather than a cheating application, and prioritizing authentic, process-based assessment over surveillance. - [[responsible-assessment-ai-era-stanford-2026|Responsible Assessment]] reframes assessment around validity evidence and authentic tasks that mirror students' future work. - [[authentic-products-authenticated-processes-2026|Authentic products, authenticated processes]] examines how AI-rich [[higher-ed|higher education]] can assess both genuine outputs and the processes that produced them. - [[tool-invariant-framework-agentic-ai|The tool-invariant framework]] argues for assessing computational methods and process rather than tool-specific outputs, using oral defense and verification. - [[fenton-oral-exams-ai-authentic-assessment-2025|Reconsidering oral exams]] positions the oral exam/assessment as a low-tech authentic alternative that is inherently AI-resistant — its real-time, interactive dialogue tests comprehension, [[critical-thinking|critical thinking]], and reasoning (not memorization), mirrors professional practice, and prevents students from using AI to generate and memorize answers. It offers a concrete set of practical recommendations (clear rubrics, standardized content, assessor training, [[prompt-engineering|prompting]] guidelines, [[bias-mitigation|bias mitigation]]) for reintroducing [[oral-assessment]] across high school and higher education. - [[eportfolio|E-portfolio assessment]] is another authentic, process-based form that resists AI fabrication: [[zhan-boud-du-authentic-assessment-scoping-review-2025|Zhan, Boud & Du (2025)]] identify social contribution portfolios among the authentic forms most robust to generative AI, and [[beyond-detection-authentic-assessment-ai-2025|Beyond Detection]] recommends annotated portfolios and recorded walkthroughs that probe reasoning in real time. [[ni-lam-multiliteracies-ai-portfolio-2026|Ni & Lam (2026)]] and [[sutama-chatgpt-eportfolio-speaking-2026|Laksana et al. (2026)]] show generative AI can assist the portfolio process — feedback, drafting, reflection — while the portfolio's reasoning traces and drafts preserve authenticity. - **Authenticity criteria can themselves restrain AI use.** [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] found that a rubric requiring students to ground a group presentation in their own first-hand teaching experience, and to reflect on shared classroom observations, led seven of fifteen student groups to *deliberately reduce* their GenAI use. Students argued the tool could not meet the epistemic demand — "AI only knows that moment when you type" — because it lacked the longitudinal, situated knowledge their classmates and [[teacher-role|teacher]] had. Both layers mattered: the authenticity of the task and the relational authenticity of contributing one's own thinking to a group, which reframed heavy AI use as free-riding on peers. The design implication is that authentic, experience-grounded criteria do evaluative work even without enforcement — they supply a reason for restraint that policy statements cannot. [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] argues that integrity-oriented design and authentic design are not the same thing, and that the difference is epistemic rather than stylistic. Authenticity asks whether a task mirrors worthwhile real-world practice; integrity-oriented design asks whether learners can justify their decisions and assume responsibility in relation to disciplinary standards, which makes integrity an explicit, assessable criterion embedded in the task architecture rather than an incidental by-product of realism. The practices he offers — annotated decision trails, verification of GenAI-contributed claims, oral defense and dialogic accountability, draft differences with version history — are recognisably authentic-assessment forms, but they are selected for the [[evaluative-judgment|judgment]] they make visible rather than for their realism, which is a useful corrective for tasks that look authentic and remain counterfeitable in substance. His caution belongs on this page too: requiring documented reasoning privileges learners more fluent in reflective discourse, and judgment as evidence "remains relational and situated rather than mechanically verifiable", so the design carries its own interpretive-reliability load. ## When practitioners retreat to the policed formats The knowledge base argues that authenticity must be designed rather than policed, and [[teacher-educators-ai-integration-preservice-2026|Goldstein, Marae-Haj and Zidan (2026)]] supply the counter-evidence from practice: after a trust crisis in which pre-service teachers submitted raw AI-generated work as their own, teacher educators in seven Israeli colleges redesigned assessment around process evidence (prompts, document version history, monitored group contributions) and in-class performance — but several also described falling back, as a last resort, on supervised examinations and anticipated oral defenses on theses, one participant calling the return to exams personally painful yet unavoidable. Read against the design-first position, the finding marks a failure mode rather than a solution: where authenticity is not redesigned in advance, the practical response to unassessable submitted work is reinstated control — surveillance by another name — which reintroduces the very formats the authentic-assessment literature tries to move beyond. It also shows the felt cost of the alternative: participants reported their take-home written assessments could no longer evidence learning at all, and that process documentation took real work to establish. The retreat carries an evidentiary cost that the case-file evidence makes visible. [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] coded 1,162 generative-AI misconduct allegations and found that the strongest evidence types are generated by the investigation or by supervision rather than by the allegation, and that process evidence such as drafts, supervision meetings and presentations appears only where those practices already exist — so in an unredesigned assessment the process evidence this literature recommends is simply not available, and cases rest on weaker categories. What stands in for it is the weakest material in their corpus: [[ai-detection|detector]] output drew the lowest probative ratings of any category, and [[hadra-ai-detector-accuracy-efl-2026|Hadra et al. (2026)]] show why, with macro accuracy of 0.69 (Originality) and 0.61 (Turnitin) across 192 texts, near-total failure on hybrid human–AI writing, and a borderline tendency to misread EFL student work as AI. [[wright-transcription-not-generation-2026|Wright (2026)]] adds the rule-design dimension: prohibitions written by platform identity rather than function are over-inclusive by definitional accident, so a policed format can sanction a student who did not do the thing the rule was designed to prevent, falling hardest on the students who relied on transcription tools for [[accessibility]]. Read together, the retreat substitutes the least probative evidence available for the process evidence that authentic design would have generated in the first place. - **Self-regulated learning:** student agency in authentic assessment (choice, self-reflection, co-design) mirrors the [[self-regulated-learning|forethought → performance → self-reflection]] cycle, though graded reflection risks becoming performative. - **Metacognition:** [[metacognition]] is required for students to evaluate their work against co-designed rubrics; when AI supplies the rubric, feedback, and monitoring, the student's metacognitive practice may be displaced. - **Formative and sustainable feedback:** authentic assessment emphasizes [[formative-assessment|formative]], future-oriented feedback that transfers to later contexts. ## Implications for AI in education - **AI-proof assessment types:** in-vivo demonstrations, social-contribution portfolios, co-created artifacts with auditable provenance, and real-time [[embodied-learning|embodied]] interaction are more resilient to generative AI than take-home essays or MCQs. - **Co-design at scale:** AI tools could enable rubric co-design and student co-creation of assessment parameters at classroom or [[online-teaching-and-learning|MOOC]] scale — though machine-mediated agency must be designed carefully. - **Address the social-authenticity gap:** only 3/37 studies addressed social issues; AI assessment tools should help students contribute to societal transformation, not merely simulate it. - **Sustainable feedback:** [[ai-feedback-quality|AI feedback]] should be designed to transfer to future contexts, not just provide reactive, momentary corrections. - **Authentic assessment suits practice-oriented fields.** [[mesny-innovative-assessment-grading-management-2026|Mesny, Roberge-Maltais & Galy (2026)]] find authentic assessment especially well-suited to [[business-education|management education]]: tasks mirroring real professional problems (live consulting projects, [[visualization|dashboards]] with executive briefings) align with the field's practice-oriented, [[career-development-and-readiness|employability]] focus and can support inclusive, integrity-preserving alternatives to exam-centered assessment in the generative AI era. In their review of 58 articles from four management-education journals, however, authentic assessment appeared mainly via technology-mediated [[simulation|simulations]] and was often conflated with [[experiential-learning|experiential learning]] — a terminology gap that can obscure its broader value and uptake. ## Connected Concepts - [[learner-identity]] — evolving disciplinary, professional, creative, and academic learner identities - [[eportfolio]] - [[problem-based-learning]] - [[assessment]] - [[assessment-validity]] - [[formative-assessment]] - [[automated-assessment]] - [[academic-integrity]] - [[ai-detection]] - [[self-regulated-learning]] - [[metacognition]] - [[ai-ed-evaluation]] - [[higher-ed]] - [[generative-ai]] - [[ai-education]] - [[feedback]] - [[summative-assessment]] — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams) - [[arts-design-and-media-education]] ## Connected Articles - [[evaluation-age-ai-output-evidence-2026]] — Evaluation in the Age of AI - [[ivory-psychology-assessment-integrity-2026]] — What AI could not pass: presence, visual artifacts, and the student's own data (Ivory et al. 2026) - [[paternalistic-filter-llm-history-education]] — Paternalistic AI use and student identity in history education - [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene]] - [[benali-genai-academic-writing-2026]] - [[ying-genai-journalism-assessment-2026]] - [[pedlow-genai-selfassessment-2026]] - [[dollinger-equitable-assessment-ai-2026]] - [[sutama-chatgpt-eportfolio-speaking-2026]] - [[nicola-richmond-programwide-assessment-genai-2025]] - [[ni-lam-multiliteracies-ai-portfolio-2026]] - [[espino-ai-business-education-review-2026]] - [[genai-simulate-patient-history-pbl-2026]] - [[pbl-structural-conditions-ai-2026]] - [[best-response-student-ai-dialog-2026]] - [[zhan-boud-du-authentic-assessment-scoping-review-2025]] — Designing for Authentic Assessment: A Scoping Review - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond Detection: Authentic Assessment - [[responsible-assessment-ai-era-stanford-2026]] — Responsible Assessment in the AI Era - [[authentic-products-authenticated-processes-2026]] — Authentic Products, Authenticated Processes - [[tool-invariant-framework-agentic-ai]] — Tool-Invariant Framework for Teaching and Assessing Computational Methods - [[cong-confidence-asag-2026]] — Confidence-aware automated short-answer grading - [[llm-fallacy-misattribution]] — The LLM Fallacy and Misattribution of Competence - [[universities-ai-era-rethinking]] — Rethinking Universities in the AI Era - [[ssaho-ai-academic-integrity-review-2025]] — Multiple assessment methods to counter AI misconduct - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering oral exams as authentic, AI-resistant assessment - [[roe-assessment-twins-2026]] — Assessment twins for strengthening assessment validity in the age of GenAI (Roe, Perkins & Giray 2026) - [[lodge-adaptive-capabilities-genai-future-2026]] — Adaptive capabilities for assuring quality learning in a gen AI-integrated future (Lodge et al. 2026) - [[bassett-ai-detectors-education-2026]] — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026) - [[asynchronous-oral-assessment-2026]] — Asynchronous Oral Assessments in the AI Era (Pentland 2026) - [[assessing-student-drive-framework-2025]] — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise) - [[mesny-innovative-assessment-grading-management-2026]] - [[chen-zou-genai-group-assessment-agency-2026]] — Authenticity criteria that led student groups to reduce GenAI use - [[teacher-educators-ai-integration-preservice-2026]] — The AI trust crisis that pushed teacher educators back toward supervised exams and oral defenses (Goldstein et al. 2026) - [[sharma-judgment-visible-genai-assessment-2026]] — Integrity-oriented design distinguished from authentic assessment: making judgment visible (Sharma 2026) - [[munoz-misconduct-allegation-evidence-2026]] — What misconduct allegation files actually contain as evidence, and the process evidence they lack (Munoz et al. 2026) - [[hadra-ai-detector-accuracy-efl-2026]] — Detector accuracy and hybrid-writing failure: why detection is weak fallback evidence (Hadra et al. 2026) - [[wright-transcription-not-generation-2026]] — Over-inclusive AI prohibitions: transcription is not generation (Wright 2026) - [[authentic-assessments-generative-ai-pilot-2026]] — Designing Authentic Assessments with Generative AI: A Pilot Study of Assessment Authentifire in Higher Education --- ## [Group Work](https://edtechdev.github.io/aied/concepts/group-work/) > **Group work** — learning and graded work produced by a team rather than an individual, used to develop collaboration, communication, and shared responsibility while also generating an outcome and often a grade for each member. It sits where [[assessment|assessment design]] and [[collaborative-learning|collaborative learning]] meet, and it inherits the tensions of both: free-riding and social loafing, unequal contributions, conflict avoidance, and the difficulty of attributing a collective product to individual learning. [[generative-ai|Generative AI]] intensifies these because teams must now negotiate whose and what kind of AI [[student-engagement|engagement]] counts as acceptable — a negotiation that [[agency]] [[research-methods-aied|research]] shows can resolve in opposite directions within a single cohort, and that [[agentic-ai|AI agents]] joining the group can themselves reshape. ## Questions to Consider - What is a group task actually producing — the group's output, each member's learning, or their ability to collaborate? If those three can diverge, which one does your grading claim to warrant? - Think of a group project you have been part of. Who did the work, who decided what counted as good work, and who stayed quiet? How much of that was the task design rather than the people? - Research on group work finds students often decline to hold free-riders accountable to protect relationships. When a group member uses [[generative-ai|GenAI]] to produce their section, is that free-riding — or is the judgment more complicated than the word suggests? - A group can access AI through a single shared interface or through each member's private prompts. Which produces more honest collaboration, and what does each conceal? - If an AI agent joined your group as a teammate — praising, questioning, or provoking — would it deepen the collaboration or reshape who drives it without anyone noticing? How would you tell the difference? - Collaborative AI that maximizes task efficiency tends to minimize learners' self-regulatory engagement. If you had to choose between the outcome and the struggle, which would you protect? - A mediator AI is trusted only while it stays neutral. When it moves from summarizing to advising or challenging, the trust erodes. How neutral should a group's AI really be? ## Introduction Group work occupies a distinctive place in [[higher-ed|higher education]] and [[professional-training|professional training]]: it is valued for building collaboration, communication, and shared responsibility, and it is pedagogically central in professionally oriented programs such as [[teacher-education|teacher education]]. Its promise is not automatically realized, however. Both teachers and students report obstacles to effective collaboration, and a graded group product carries two different claims at once — about what the group accomplished and about what each individual learned. AI, and [[generative-ai|generative AI]] in particular, presses on both: it changes how work can be divided, how easily contributions can be fused, and how each member's thinking can be distinguished from a tool's output. A large part of the knowledge base's collaborative-learning research bears on group work, and this page draws on several strands of it rather than any single study. ## Why group work is difficult by design - **Free-riding and social loafing.** Some members reduce their contribution and the workload is redistributed onto more engaged peers. Task scope, group size, and [[writing-education|composition]] raise the likelihood; accountability mechanisms and group efficacy reduce it. - **Accountability is relational as well as structural.** In Hong Kong — where one of the knowledge base's group-assessment studies was conducted — undergraduates often chose *not* to report free-riding to protect interpersonal relationships, completing the work themselves instead of confronting the problem. Students' [[agency]] can therefore take the form of pursuing the grade while avoiding the conflict. - **The collaboration has to be designed, not assumed.** Where a task can be partitioned into independent subtasks, students take that route and the group becomes a submission format rather than a joint intellectual endeavor. [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] found three of their fifteen groups doing exactly this — dividing the work, working on "individual platforms," and never pooling their individual GenAI capabilities even when coherence was an explicit assessment criterion. - **Process versus product.** Group work is strongest when the process is assessed alongside the output, which requires [[peer-assessment|peer assessment]] or other visible intermediates rather than a single final artifact. - **[[equity-in-ai-education|Equity]] and [[inclusive-learning|inclusion]].** Contribution norms, language, and confidence distribute unequally inside groups. [[neurodivergent-computing-students|Neurodivergent students]] report needing structured assignments, small consistent teams, and explicitly defined roles — requirements that AI collaboration tools, often built for the "average" learner, routinely fail to accommodate. ## How AI changes group work Research on AI in group work addresses two distinct things: what AI adds to a group's *task*, and what AI does to the group's *process*. Both matter, and the knowledge base's findings span them. ### AI as coordination infrastructure The most direct finding on GenAI and group assessment is [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou's (2026)]] three-pattern study of fifteen pre-service teacher groups. Rather than a single enthusiasm-to-avoidance spectrum, agency ran in three directions at once. | Pattern | Groups | Practice | Rationale | |---|---|---|---| | Cooperation-oriented agency | 5 | Intensified GenAI use | Coherence, performance, shared norms lowering perceived risk | | Normative agency | 7 | Deliberately restrained use | Authenticity, [[bias-mitigation|fairness]], originality, diversity of perspectives | | Non-enacted agency | 3 | Unchanged from individual work | Task partitioned; collective interaction viewed as costly | - **GenAI as coordination infrastructure.** Students intensified use to solve a familiar [[collaborative-learning|collaboration]] problem — not knowing what peers' sections contained. Feeding those sections into a [[conversational-ai|chatbot]] to decode and align them made distributed knowledge mutually intelligible, and one group rebuilt its workflow as "discussion → externalisation to GenAI → collective review → re-discussion." The authors read this as going beyond [[cognitive-offloading|cognitive offloading]], since judgment stayed with students while the tool absorbed coordination work, while warning that the smoother workflow may bypass the disagreement through which cohesion is built. - **Group norms can invert accountability.** A permissive collective climate lowered the perceived [[ai-misuse-learning-harm|risk of misuse]] ("it is more comfortable because besides me, everyone in my group is using GenAI"), turning shared norms into shared risk management rather than commitment to learning — the opposite of what group accountability is meant to achieve. - **Restraint is also agentic.** Seven groups cut their GenAI use to protect the task's [[situated-learning|situated]] knowledge ("AI only knows that moment when you type"), to avoid effectively free-riding on groupmates, to preserve the originality that distinguishes groups from one another, and to keep the diversity of perspectives the group already had — normative [[self-regulated-learning|self-regulation]] rather than mere compliance, and hard to distinguish from disengagement from the outside. ### How the group accesses AI shapes the collaboration [[xu-genai-collaborative-space-2026|Xu et al. (2026)]] show that *how* a team accesses GenAI is itself a design decision. With a single shared interface in synchronous work, teams co-construct "collective prompts," run a surface–evaluate–embed cycle, and treat the chat as shared memory; in asynchronous work, private [[prompt-engineering|prompting]] and output "de-labeling" fragment [[explainable-ai|transparency]] and raise the cost of sustaining a shared cognitive model. Access configuration therefore connects directly to the quality of interactive engagement — whether the group is genuinely co-constructing or merely co-approving. ### AI as a group member reshapes the group A second line of research treats AI as a participant rather than a tool, and its findings complicate the assumption that "AI teammate" means "better teamwork." - **Agents as social stabilizer — and a magnet for questions.** In fifteen groups of three [[stem-education|STEM]] students deliberating [[ethics]] with three [[llm]] participants ([[ethics-training-agents-group-ethics-discussion-2026|Seo et al., 2026]]), the agents kept groups on topic when human members drifted ("even if the other two participants strayed, I didn't have to handle it") and prompted clarifications humans avoided for relational reasons. But the group dynamic shifted interaction away from humans — 79.7% of questions were directed at agents (*p* = .017) — and participants noted a breadth–depth trade-off in the turn-stacking format, where each speaker took the floor in turn and "other participants couldn't intervene," making it easy to cover many ideas but hard to explore any one in depth. - **AI personas reconfigure emergent [[agency]].** [[jin-emergent-learner-agency-implicit-hai-2026|Jin et al. (2026)]], in an experiment where AI operated as an undisclosed teammate, found supportive and contrarian personas still reshaped the group: contrarian AI pulled discourse toward challenge and reflection ([[desirable-difficulties|productive friction]]), while supportive AI stabilized agreement and renewed ideation. Critically, the friction had a [[affective-computing|affective]] cost — contrarian AI reduced teamwork satisfaction and psychological safety *without* yielding creative gains — so "productive" discourse structures must be weighed against their emotional consequences. Because the personas worked even without disclosure, they position persona design as a form of invisible [[governance]] over group collaboration. - **Controversy can counter groupthink.** [[genai-counter-learner-groupthink-2025|Wiss et al. (2025)]] took the opposite tack deliberately: an AI agent (CALIE) prompted to inject controversial viewpoints into twelve interprofessional [[problem-based-learning|PBL]] teams stimulated [[critical-thinking|critical thinking]] and positive group dynamics, countering the conformity pressure that group work can otherwise produce. The contrast with Jin et al. is instructive — challenge helps when it is the intended design, and hurts when it is an unacknowledged side effect. - **The neutral-mediator constraint.** [[spritz-ai-disciplinary-mediation-student-teams-2026|Spritz]], an AI that mediates disciplinary boundaries in interdisciplinary teams by surfacing implicit assumptions, was valued as both cognitive support and a relational buffer — but students' trust was load-bearing and eroded the moment the AI moved from neutral mediator to advisor or challenger. Role switches therefore need to be explicit and configurable, not silent. - **Collaboration can itself be the object of instruction.** [[golrang-propact-pair-programming-2026|ProPACT]] treats the *dyad* — not the individual — as the unit of analysis, modeling joint attention and effort to predict collaborative breakdowns up to 30 seconds in advance. Dyads receiving proactive feedback achieved higher debugging success and showed sustained gains in collaborative [[regulation]] afterward: AI can teach collaboration itself, not just support a task. ### The efficiency–regulation trade-off [[hao-human-ai-collaborative-problem-solving-cognition|Hao et al.]] identified three collaborative modes with AI on complex problems — *Delegated Reasoning*, *Concerted Interpretation*, and *Delegated Elaboration*. The most efficient mode (delegated reasoning) yields the best task performance but the lowest self-regulatory engagement; the mode with greatest self-regulation (concerted interpretation) underperforms on the task. This is the central design tension for AI-mediated group work: balance the efficiency of the distributed human–AI system against the depth of learners' regulatory engagement. ### The epistemic risk of polished output [[polished-artifacts-fragile-engagement-2026|Kimmerle]] conceptualizes the risk of reduced epistemic effort when learners use AI to produce polished knowledge artifacts — automation bias on the social-cognitive side and epistemic closure induced by the finished artifact on the other. The remedy is to structure AI as an argumentative partner or challenger that preserves conflict. The counterpoint comes from [[oppenheimer-llms-collaborative-learning-partners-2026|Oppenheimer et al.]], who found that when LLMs act as critique partners and students actively rebut their claims, learners behave as critical consumers who preserve rather than surrender the cognitive conflict of critique. The difference between an answer-giving teammate and a challenging partner is what determines whether group work with AI builds or bypasses thinking. ### Classroom-level, non-evaluative support [[breideband-community-builder-cobi-2026|CoBi]] detects "uplifting" small-group discourse (respect, equity, community, moving thinking forward) and returns non-evaluative, classroom-level visualizations — deliberately withholding student- or group-level feedback to protect [[privacy]] and [[trust]]. Students preferred [[qualitative-research|qualitative]] visualizations (an organic tree) over [[quantitative-research|quantitative]] ones (a radar chart). For the relational dimension of group work, class-level aggregated feedback supports community building where individual scoring would feel surveilled. ## Design implications - **Make the AI-use negotiation an assessable outcome.** [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] argue teams should justify and document how GenAI will and will not be used, turning an implicit peer norm into explicit, gradeable reasoning rather than leaving it to perceived risk. - **Assess process, not only product.** [[peer-assessment|Peer assessment]], intermediate deliverables, and reflective contributions create the interactions through which norms and shared understanding are actually constructed. - **Do not rely on a coherence criterion alone.** A rubric line demanding integration does not produce collaboration if the task structure still permits divide-and-conquer; the workflow has to require joint work. - **Design access deliberately.** Whether the group shares one AI interface or each member prompts privately changes the transparency of the collaboration and the cost of sustaining a shared cognitive model. - **Decide, visibly, what the AI's role is.** A neutral mediator is trusted; an advisor or challenger is not, unless that role is explicit and intended. A contrarian persona can counter groupthink, but only as a designed feature with its psychological-safety costs acknowledged. - **Balance efficiency against self-regulation.** Collaborative AI that maximizes task efficiency can undercut learners' regulatory engagement; design should deliberately protect space for concerted interpretation. - **Prefer an argumentative partner over an answer-giver.** Structure AI to surface disagreement rather than smooth it over, so group work preserves the cognitive conflict that builds understanding. - **Accommodate [[neurodiversity|neurodivergent]] learners.** Structured assignments, small consistent teams, and explicit role definitions are requirements AI collaboration tools must support. - **Read restraint carefully.** Students who avoid AI in groups may be exercising normative self-regulation — or guarding against risk and unfamiliarity. The two call for different [[teacher-role|instructor]] responses. ## Connected Concepts - [[ai-education]] — AI in education (umbrella) - [[collaborative-learning]] - [[assessment]] - [[peer-assessment]] - [[academic-integrity]] - [[agency]] - [[authentic-assessment]] - [[ai-use-disclosure]] - [[higher-ed]] - [[teacher-education]] - [[formative-assessment]] - [[student-engagement]] - [[learning-design]] - [[problem-based-learning]] - [[sociocultural-learning]] ## Connected Articles - [[chen-zou-genai-group-assessment-agency-2026]] — Three patterns of student agency in GenAI-mediated group assessment - [[jin-emergent-learner-agency-implicit-hai-2026]] — Emergent learner agency when AI joins a group: supportive vs. contrarian personas - [[genai-counter-learner-groupthink-2025]] — An AI agent that injects controversy to counter groupthink in PBL teams - [[xu-genai-collaborative-space-2026]] — Shared versus private GenAI access across a team - [[hao-human-ai-collaborative-problem-solving-cognition]] — Collaboration modes and the task-performance versus self-regulation trade-off - [[golrang-propact-pair-programming-2026]] — ProPACT: treating the dyad as the unit of analysis in pair programming - [[spritz-ai-disciplinary-mediation-student-teams-2026]] — AI mediation of assumption-surfacing in interdisciplinary student teams - [[oppenheimer-llms-collaborative-learning-partners-2026]] — LLMs as critique partners: active rebuttal preserves cognitive conflict - [[polished-artifacts-fragile-engagement-2026]] — Polished artifacts, fragile epistemic engagement - [[breideband-community-builder-cobi-2026]] — Classroom-level discourse analytics for group collaboration - [[niari-ai-pedagogical-mediator-collaborative-learning]] — AI as pedagogical mediator of collaboration - [[neurodivergent-computing-students]] — Neurodivergent students' requirements for collaborative learning - [[academic-league-of-ai-2026]] — Democratic student governance and project teams for AI education - [[teacher-student-agency-orchestration]] — Co-orchestration of teacher and student agency in real time - [[dollinger-equitable-assessment-ai-2026]] — Equitable assessment in the AI era - [[ethics-training-agents-group-ethics-discussion-2026]] — Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration - [[durable-skills-measurement-ai-teammates-2026]] — Toward Scalable Measurement of Durable Skills - [[peer-group-vs-ai-feedback-2026]] — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education --- ## [E-Portfolio](https://edtechdev.github.io/aied/concepts/eportfolio/) > **E-Portfolio (e-portfolio)** — a digital collection of a learner's work, reflections, and evidence of achievement, used for [[assessment]], learning, and evaluation. In the AI era, e-portfolios have emerged as a relatively **AI-robust and AI-assisted** assessment form: they capture the *process* of learning (reasoning, drafting, reflection) rather than just the final artifact, making them valuable against the collapse of the artifact-as-proxy, and [[generative-ai|generative AI]] can support their creation, [[feedback]], and evaluation. ## Questions to Consider - An e-portfolio captures the process of learning — drafts, revisions, reflections — not just the final artifact. Why might the process be exactly what AI cannot easily fabricate, and what does that make portfolios good for in an AI-heavy classroom? - With AI able to produce polished final products on demand, the 'artifact-as-proxy' for learning has collapsed. What do you think an assessor can learn from a portfolio that they can no longer trust from a single submitted essay or exam? - Some research recommends annotated portfolios, oral defenses, and recorded walkthroughs to probe reasoning in real time, mirroring professional practice. How different would that make assessment from the essays and exams you've taken — and fairer or harder to scale? - AI can support the portfolio process itself — giving feedback, aiding drafting, and prompting reflection. Where's the line between AI genuinely helping a student learn and AI quietly doing the thinking the portfolio is meant to reveal? ## Introduction E-portfolios are a form of [[authentic-assessment|authentic assessment]] and [[formative-assessment|formative assessment]]: they assemble student work over time, often with reflective commentary, into a body of evidence that can be assessed holistically. Their strength is that they surface the learning *process* — drafts, revisions, feedback, and self-reflection — which is exactly what AI cannot easily fabricate and what assessors need to evaluate genuine learning. This makes portfolios a leading candidate for assessment redesign in the generative AI era. ## E-portfolios in the AI era The knowledge base's research shows e-portfolios are increasingly central to productive AI integration and assessment redesign. - **Portfolio assessment is relatively robust to generative AI.** [[zhan-boud-du-authentic-assessment-scoping-review-2025|Zhan, Boud & Du (2025)]] identify **social contribution portfolios** and co-created artifacts with auditable provenance chains among the authentic-assessment forms most robust to generative AI. [[beyond-detection-authentic-assessment-ai-2025|Beyond detection (2025)]] likewise recommends **annotated portfolios / oral defenses / recorded walkthroughs** to probe reasoning in real time, mirroring professional practice. - **AI-assisted portfolio assessment.** [[ni-lam-multiliteracies-ai-portfolio-2026|Ni & Lam (2026)]] study students' perceptions of **AI-assisted portfolio assessment** for multiliteracies development, showing AI can support the portfolio *process* (feedback, drafting, reflection) and enhance engagement while students evaluate AI feedback critically. - **ChatGPT + e-portfolio for language learning.** [[sutama-chatgpt-eportfolio-speaking-2026|Laksana et al. (2026)]] combine ChatGPT with e-portfolio assessment (CEA model) for EFL speaking: MANOVA showed significant gains in speaking performance (F(1)=48.554, p<.001) and [[feedback-literacy|feedback literacy]] (F(1)=16.135, p<.001), with e-portfolios supporting both speaking ability and the metacognitive abilities to use feedback effectively. - **Portfolios as AI-era assessment strategy.** [[responsible-assessment-ai-era-stanford-2026|Responsible assessment (2026)]] argues for shifting from one-shot testing to **continuous embedded assessment using portfolios** and competency-based approaches. [[pbl-structural-conditions-ai-2026|Rowe (2026)]] names portfolios (with orals and observed practice) as partial answers to the open problem of assessment at scale when the artifact-as-proxy has broken. - **Portfolios as assessable learning traces.** [[learn-framework-responsible-genai-pbl-2026|Uden & Hwang (2026)]] use **personal learning portfolios** as a LEARN-framework mechanism — assessable learning traces (portfolios, AI-use disclosure) that make AI reliance transparent and auditable. - **Symbiotic portfolios for hybrid learners.** [[elsayed-pedagogical-symbiosis-posthuman-learner|Elsayed (2026)]] proposes a **Symbiotic Portfolio** for the post-human/hybrid learner, assessed by a five-criterion rubric (Generative Dialogue, Epistemic Auditing, Critical Reflection, etc.) — an explicit model for evaluating human+AI collaborative work. ## Why e-portfolios matter for AI integration Because e-portfolios foreground **process, reflection, and demonstrated understanding** over a single final artifact, they directly address the central AI problem: when AI can produce a polished product, the product no longer evidences the engagement behind it. E-portfolios shift assessment toward the evidence that survives — drafts, revisions, reasoning traces, and reflective self-assessment — while AI itself can assist in generating feedback, scaffolding reflection, and (with appropriate rubric design) supporting evaluation. This makes e-portfolios a cornerstone of [[authentic-assessment|authentic]], [[formative-assessment|formative]], and process-based assessment in the [[generative-ai|generative AI]] era, and a natural home for [[feedback-literacy|feedback literacy]] and [[self-regulated-learning|self-regulated learning]]. ## Connected Concepts - [[assessment]] - [[authentic-assessment]] - [[formative-assessment]] - [[feedback]] - [[feedback-literacy]] - [[generative-ai]] - [[self-regulated-learning]] - [[metacognition]] - [[student-engagement]] - [[language-learning]] - [[higher-ed]] ## Connected Articles - [[ni-lam-multiliteracies-ai-portfolio-2026]] — Students' perceptions of multiliteracies development using AI-assisted portfolio assessment (Ni & Lam 2026) - [[sutama-chatgpt-eportfolio-speaking-2026]] — ChatGPT aligned with e-portfolio assessment for EFL speaking and feedback literacy (Laksana et al. 2026) - [[responsible-assessment-ai-era-stanford-2026]] — Portfolio- and competency-based assessment in the AI era - [[zhan-boud-du-authentic-assessment-scoping-review-2025]] — Authentic assessment robust to generative AI, incl. social contribution portfolios - [[beyond-detection-authentic-assessment-ai-2025]] — Annotated portfolios / oral defenses as AI-resistant assessment - [[elsayed-pedagogical-symbiosis-posthuman-learner]] — The Symbiotic Portfolio for hybrid learners - [[pbl-structural-conditions-ai-2026]] — Portfolios as partial answers to assessment at scale - [[learn-framework-responsible-genai-pbl-2026]] — Personal learning portfolios as assessable learning traces --- ## [Peer Assessment](https://edtechdev.github.io/aied/concepts/peer-assessment/) > **Peer assessment** — the practice in which students evaluate, grade, or give [[feedback]] on one another's work, through written peer review, peer grading, peer code review, or group and team assessment. In [[writing-education|writing]] [[pedagogy]] it is a long-standing best practice: students learn both from receiving criteria-based feedback and from providing it, and peer talk about shared work correlates with deeper learning, audience awareness, and social development. In the AI era it is being re-examined as an [[assessment|assessment design]] choice rather than a single activity — a human complement to [[ai-feedback-quality|AI-generated feedback]], a training ground for [[feedback-literacy]], and a site where [[academic-integrity]] and [[agency|student agency]] are renegotiated. ## Questions to Consider - Think back to a time your work was assessed by a peer—or you assessed theirs. What did you learn more from: giving feedback or receiving it, and why? - Why might a student's feedback be more context-aware and emotionally supportive than AI feedback, even while AI feedback is more consistent and rubric-driven? How could the two complement each other? - Peers tend to under-grade strong work and AI tends to inflate weak work. If neither is reliably accurate across the whole quality range, what should a grade from either source be used for? - Peer assessment depends heavily on [[scaffolding]] and clear criteria, and friendship bias and social anxiety can make untrained peer review worse than none. What does adequate training actually look like in your context? - If AI can draft comments on organization and structure, does that make peer assessment redundant—or does it free peers to give the specific, audience-aware feedback only they can give? - GenAI-supported peer feedback outperformed plain peer feedback in one study, but only with prompt scaffolding added. What is being scaffolded there—the AI, the student, or the assessment? - Critically assessing AI-generated feedback is described as building AI literacy and writerly agency. What does it mean to have agency over your own work when machines increasingly comment on it? ## Introduction Peer assessment is an assessment design choice in which students take on part of the evaluative work usually reserved for instructors. Its forms differ in what students produce and what is at stake: peer feedback (comments on a draft, usually [[formative-assessment|formative]] and low-stakes), peer grading (a mark contributing to a grade), peer code review, and group or team assessment (where peers assess the collective product or one another's contributions). The form determines what students practice — giving a criteria-based comment builds [[evaluative-judgment|evaluative judgment]] differently from assigning a number, and defending one's own code orally builds something different again. [[self-assessment]], the sibling practice, turns that same criteria-based scrutiny on the learner's own work, whereas peer assessment directs it at a peer's submission and the audience awareness that comes with it. The research base agrees that peer assessment produces learning and disagrees about how much depends on design. It gives students an authentic audience, develops evaluative judgment through criteria-based responding, and builds the social context that supports [[student-engagement|engagement]] and [[motivation]]. Its quality depends heavily on [[scaffolding]] — how well it is structured, and whether students get clear criteria and training. That dependency is where AI enters, as a consistent, rubric-driven complement to the specific, context-aware feedback peers give, and increasingly as a scaffold built into the peer-assessment process itself. ## What students learn from assessing peers The strongest argument for peer assessment is that the assessor learns. [[code-review-genai-cs1|Fowles et al. (2026)]] made every CS1 submission subject to a 15-minute oral code review interview with a trained teaching assistant, weighted at 70% of the assignment grade. Pasted-to-total characters rose from 61.0% to 68.1% (p < 0.0001), yet exam scores did not decline and time-on-task held steady; 90% of students said the reviews motivated them to understand their code better and 65% that they helped avoid [[cognitive-offloading|over-reliance]] on AI. Requiring students to explain their work to a trained peer converted potential offloading into a [[self-regulated-learning|self-regulated]] practice opportunity. The mechanism recurs across the literature. In a PRISMA 2020 [[meta-analysis-systematic-review|systematic review]] of 22 articles from 203 screened, [[llm-critical-thinking-teamwork-review|Martínez-Peláez et al. (2025)]] found LLMs support [[collaborative-learning|collaboration]] partly by aiding peer feedback and simulating rubric-based evaluations, and that students who verify and correct model output develop critical thinking. [[learning-by-teaching]] shows up even among machines: peer-learning-like discourse among over 2.4 million [[agentic-ai|AI agents]] was mostly assertion rather than inquiry (statement-to-question ratio 11.4:1, [[metacognition|metacognitive]] reflection only 7% of a coded taxonomy across 28,683 posts), which [[ai-agents-peer-learning-discourse|Chen et al. (2026)]] treat as a caution that surface discourse does not establish that anything is learned. ## Training, calibration and feedback literacy Peer assessment fails predictably when students are untrained. [[scaffolding-srl-feedback-genai-human-peers|Gu, Chen, and Yan (2026)]] ran a [[mixed-methods-research|mixed-methods]] quasi-experiment with 118 first-year [[higher-ed|undergraduates]] in China, comparing [[generative-ai|ChatGPT]]-4o under pre-trained rubrics against structured peer-review worksheets with a high-quality exemplar. Feedback literacy rose slightly more in the GenAI group (ANCOVA p = 0.049, η²p = 0.03), but the [[qualitative-research|qualitative]] findings matter more: peer-group students chose feedback sources by social convenience — nearby peers, same-major peers, roommates — rarely held specific feedback goals, and faced social anxiety about seeking feedback. Peer evaluation was more often distorted by friendship bias and perceived peer proficiency, and peer reflection was often delayed until later exams. The authors recommend multi-stage designs with anonymous peer feedback, since peer review still uniquely builds audience awareness and evaluative judgment through giving feedback. Design frameworks target those bottlenecks. [[irwin-muller-efl-peer-feedback-literacy|Irwin and Muller (2026)]] propose two GenAI roles in EFL/ESL peer feedback on speaking, a case they argue is harder than writing because of time pressure, fleeting oral performance, and heightened affect: a Trainer supporting feedback givers through exemplar-based calibration and feedback-on-feedback, and a Synthesizer aggregating peer comments into a criteria-linked uptake report that normalizes formats, preserves minority views, and flags contradictions. Their principles are careful timing and sequencing, short repeated Trainer units, preserving the givers' voice, and [[guardrails]] that keep the [[teacher-role|teacher]] in the loop with no [[automated-assessment|automated grading]]. The paper is conceptual, with no new empirical data. ## Validity, fairness and grading accuracy: peer vs AI vs instructor Where peer assessment produces a grade, the question is whether the grade is defensible. [[usher-faraon-who-grades-best-2026|Usher and Faraon (2026)]] found peer–instructor alignment was strongest for lower-quality work (r = 0.51 in the low-quality tier) and weakened for high-quality projects, which peers tended to under-grade — plausibly reflecting reluctance to criticize strong work. ChatGPT showed the opposite pattern: better alignment on high-quality work, but inflated grades for weak submissions. Students perceived peers' grades as more consistent with the peers' own written feedback, and valued peers' relationship-aware judgment alongside ChatGPT's neutral consistency. Peer accuracy is quality-dependent and relationship-sensitive, and no single source is reliable across the whole range. That constrains what a peer grade should be used for. A mark calibrated in the middle of the distribution need not be calibrated at the top or bottom, so peer grades are more defensible as input to a moderated decision or as evidence of the assessor's judgment than as a final mark on strong work. [[multimodal-affective-its-presentation|Suen and Hung (2026)]] underline why institutions keep seeking scalable alternatives: peer feedback and expert coaching are time-consuming, costly, and hard to scale consistently. Their closed-loop system scored 204 [[adult-learning|adult learners]] on presentation skills with an XGBoost backbone approaching expert-rater reliability (ρ = 0.69–0.78) and pre–post gains of Cohen's d = 0.39–0.90 — one pre–post field study, not a comparison against peer assessment, substituting machine judgment rather than augmenting peer judgment. ## AI-augmented peer assessment: PAIRR and the Trainer/Synthesizer split The dominant design pattern keeps the peer central and adds AI around them. The Peer and AI Review + Reflection model is the best-documented instance. [[pairr-ai-peer-review-2025|Sperber et al. (2025)]] implemented PAIRR across 10 writing courses and three writing-intensive [[stem-education|STEM]] courses, with 654 students (37% first-generation, 13% international, 68% [[multilingual-learning|multilingual]]). Students drafted, exchanged peer review, prompted ChatGPT for rubric-driven feedback on the same draft, critically assessed both, planned revisions, and reflected. 58% preferred combined feedback, 36% peer feedback alone, and only 6% AI feedback alone. The two sources were often similar (75% reported similarities, described as confirming and strengthening each other) and, when they differed, complementary: AI feedback was often called overly general (31%) but gave actionable revision strategies on organization and structure, while peer feedback was more specific and detailed (28%) and drew on contextual knowledge of the assignment. Critically [[ai-ed-evaluation|evaluating AI]] output built [[ai-literacy]], and only 5.3% of students showed overconfidence in it. The pattern generalizes to professional writing. [[gift-ai-pairr-business-writing-2025|MacArthur et al. (2025)]] applied PAIRR to an upper-division business writing course where five major assignments each required a draft, audience analysis, peer review by 2–3 peers, and revision (34 of 46 enrolled students participated; 69% multilingual). Students valued AI feedback but sometimes found it too general, and valued peers' contextual knowledge — "my peers looked at it from the employer's perspective… ChatGPT did not do that as much." One quarter of coded reflections in the larger PAIRR study expressed skepticism about AI feedback or noted inaccuracies, which the authors read as developing AI literacy. The study is descriptive and course-specific. ## GenAI as a scaffold inside peer feedback and group assessment The most rigorous test of how much design matters is a multisite cluster-randomized experiment. [[genai-feedback-design-multisite-experiment|Ateş (2026)]] randomized 48 sections across 4 universities — 1,176 first-year undergraduates in [[biology-education|biology]], [[chemistry-education|chemistry]], and [[physics-education|physics]] — to four conditions for scientific argumentation: peer feedback only, direct GenAI feedback, reflective GenAI feedback (self-evaluation then AI critique), and a hybrid of self-evaluation → peer feedback → GenAI critique. Direct GenAI beat peer feedback on immediate argument quality but showed weaker [[transfer-of-learning|transfer]]; reflective and hybrid designs produced stronger feedback uptake and self-regulated learning; the hybrid showed the clearest advantage on conceptual learning; both outperformed direct GenAI on delayed AI-free transfer. GenAI's value, the authors conclude, depends less on access than on whether the environment preserves student agency and ownership during revision. Adding GenAI can also raise the quality of the peer feedback itself, but apparently only with prompt support. [[chang-genai-peer-feedback-collaborative-argumentation-2026|Chang et al. (2026)]] compared three conditions among 45 student teachers in 12 groups over four rounds of collaborative argumentation: plain peer feedback, peer feedback with GenAI, and peer feedback with GenAI under prompt scaffolding. The GenAI-supported groups outperformed plain peer feedback on argumentation performance, and the prompt-scaffolded group performed best on advanced elements such as "rebuttal data and warrant" and "addressing the opposing view". GenAI-supported groups produced more explanations, suggestions, and neutral or negative feedback, and the scaffolded group paired negative emotions with higher-order feedback content — [[critical-thinking|critical evaluation]] rather than passive acceptance. It is a small single experiment, but it isolates prompt scaffolding as the active ingredient. [[group-work|Group assessment]] changes the problem, because students must negotiate whose AI use is acceptable. [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] interviewed 15 focus groups of 52 pre-service teachers in a course where a group presentation worth 30% of the final grade required integration and coherence. Three patterns emerged: cooperation-oriented agency in five groups, who intensified GenAI use to hold the work together and protect a shared grade; normative agency in seven groups, who restrained use to protect authenticity and fairness — one student reasoned that generating a part in AI would be "not fair to other groupmates"; and non-enacted agency in three groups whose practice never changed from individual work. The authors argue that individual capability does not become collective agency on its own and that the negotiation of acceptable AI use should itself become an assessable outcome, with peer review tasks feeding the final product. ## Feedback architecture, disclosure, and the social conditions of peer assessment Peer assessment also depends on what students can see of each other and what they will admit. [[hao-peer-exposure-bridging-social-capital-ai-summaries-2026|Hao and Cukurova (2026)]] tested an AI-generated summary [[learning-design|learning design]] across three iterations with 128 students over eight weeks. Students in AI-supported iterations showed significantly higher standardized out-degree centrality than the baseline (β = 0.547 and β = 0.438, both p < 0.001): they viewed or interacted with more peers. Interviews described the summaries as a navigation map that reduced the effort of finding contributions buried in ill-named threads, and students attended to active contributors chosen for contrasting opinions rather than friendship. The design did not prevent declining viewing activity under rising workload, so technological affordances alone do not sustain engagement. Peer norms also shape honesty, which matters wherever peers assess AI-influenced work. [[qu-wang-disclose-or-not-genai-2026|Qu and Wang (2026)]] surveyed 409 Singaporean undergraduates about why students conceal GenAI use despite disclosure mandates, and found relational variables dominated: perceived peer disclosure and comfort with instructors were the strongest predictors of disclosure, while moral disengagement was weaker. Non-disclosure was strategic adaptation to perceived peer norms and low [[trust|interpretive trust]], not moral negligence. Peer norms can therefore support honesty, as when group-assessment students restrained AI use out of obligation to groupmates, or suppress it, as when collective adoption lowered the perceived risk of misuse. Transparency, the authors argue, depends less on compliance than on trust and positive normative climates — a relational condition that peer-assessment designs either create or destroy. ## Design implications and open questions Several design moves follow. Give assessors training, exemplars, and criteria before they assess, because untrained peer assessment is vulnerable to friendship bias, convenience-based source selection, and social anxiety. Sequence the work so self-evaluation precedes peer feedback and AI critique, since the hybrid condition produced the strongest conceptual learning and best delayed transfer. Where GenAI enters, scaffold how students prompt it. Keep the teacher in the loop for anything that becomes a grade, treat peer grades as quality-dependent evidence rather than uniform marks, and make AI use itself something groups negotiate and document. The open questions are about the strength of the evidence, not only about design. Much of the peer-and-AI evidence is small and context-bound: 45 student teachers in one course, 34 students in one business writing course, 52 pre-service teachers in 15 focus groups. The PAIRR survey is the largest dataset here and measures student perceptions, not the quality of the AI outputs students judged. Only the multisite experiment's 1,176 undergraduates approaches causal-comparative scale, and it tests feedback design for scientific argumentation rather than peer assessment as such. Meanwhile [[oneill-presumed-effective-meta-analysis-2026|O'Neill (2026)]] audited 14 peer-reviewed meta-analyses claiming AI improves education and found none provided a valid basis for its claims — all but two defined the treatment as a tool rather than a pedagogical intervention, heterogeneity was high in every meta-analysis reporting I² (77.2% to 94.4%), and an audit of 59 primary studies found 61% had [[assessment-validity|validity concerns]], most often a mismatch between the outcome measured and the claim made. Claims about what AI does in peer assessment should be treated as claims about a designed activity, tested in that activity's terms. ## Connected Concepts - [[writing-education]] - [[formative-assessment]] - [[self-assessment]] - [[ai-feedback-quality]] - [[ai-literacy]] - [[self-regulated-learning]] - [[metacognition]] - [[student-experience]] - [[collaborative-learning]] - [[academic-integrity]] - [[feedback-literacy]] - [[feedback]] - [[group-work]] - [[assessment]] - [[scaffolding]] ## Connected Articles - [[usher-faraon-who-grades-best-2026]] — Comparing ChatGPT, peer, and instructor grading across project quality levels (Usher & Faraon 2026) - [[pairr-ai-peer-review-2025]] — Peer and AI Review + Reflection (PAIRR) - [[becerra-aicofe-feedback-2026]] — AI Peer Feedback Systems - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond Detection: Redesigning Authentic Assessment - [[ai-internal-feedback-evaluative-judgments]] — Unravelling Undergraduates' Development of Evaluative Judgments - [[learner-centered-feedback-ai]] — Enhancing Learner-Centered Feedback With AI - [[genai-linguistic-diversity-academic-writing]] — Generative AI and Linguistic Diversity in Academic Writing - [[gift-ai-pairr-business-writing-2025]] — PAIRR in a business writing course: peer review, chatbot feedback, reflection (MacArthur et al. 2025) - [[scaffolding-srl-feedback-genai-human-peers]] — GenAI vs. human peers for scaffolding self-regulated feedback and feedback literacy (Gu et al. 2026) - [[irwin-muller-efl-peer-feedback-literacy]] — GenAI as Trainer and Synthesizer in EFL peer feedback on speaking (Irwin & Muller 2026) - [[genai-feedback-design-multisite-experiment]] — Multisite experiment comparing peer-only, direct, reflective, and hybrid GenAI feedback (Ateş 2026) - [[chang-genai-peer-feedback-collaborative-argumentation-2026]] — Prompt-scaffolded GenAI peer feedback in collaborative argumentation (Chang et al. 2026) - [[chen-zou-genai-group-assessment-agency-2026]] — Agency in GenAI-mediated group assessment: cooperation, restraint, non-enactment (Chen & Zou 2026) - [[code-review-genai-cs1]] — Oral code review interviews as harm reduction for GenAI in CS1 (Fowles et al. 2026) - [[multimodal-affective-its-presentation]] — Automated multimodal presentation coaching as a scalable alternative to peer feedback (Suen & Hung 2026) - [[hao-peer-exposure-bridging-social-capital-ai-summaries-2026]] — AI-generated discussion summaries broaden peer exposure in online forums (Hao & Cukurova 2026) - [[qu-wang-disclose-or-not-genai-2026]] — Peer influence and relational trust in students' GenAI disclosure (Qu & Wang 2026) - [[llm-critical-thinking-teamwork-review]] — Systematic review of LLMs for critical thinking, teamwork, and problem solving (Martínez-Peláez et al. 2025) - [[ai-agents-peer-learning-discourse]] — Peer-learning-like discourse among 2.4 million AI agents (Chen et al. 2026) - [[oneill-presumed-effective-meta-analysis-2026]] — Audit of 14 AIED meta-analyses and 59 primary studies (O'Neill 2026) - [[peer-group-vs-ai-feedback-2026]] — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education --- ## [Self-Assessment](https://edtechdev.github.io/aied/concepts/self-assessment/) > **Self-assessment** — judging your own work, competence, or progress against criteria, and the practices and instruments built on that act. The term carries two faces that are easy to collapse into one. As an **instructional technique** it is self-grading, rubric-based self-review, self-checking, and guided reflection: activities meant to develop the [[metacognition|monitoring]] and [[evaluative-judgment|judgment]] that [[self-regulated-learning]] runs on. As an **educational measure** it is a [[self-report-measures|self-report instrument]], in which a learner's estimate of their own skill, confidence, or learning stands in for an observation nobody made. It sits beside [[peer-assessment]] as the other half of students-as-assessors, and inherits the same dependence on [[scaffolding]] and explicit criteria. Both faces turn on the same faculty and both fail in the same direction: the estimate is systematically generous, and the more a learner's submitted work can be produced by a [[generative-ai|generative AI]] tool, the less that work says about whether the faculty is there at all. ## Questions to Consider - When you finish a piece of your own work, what tells you whether it is good: the criteria you were handed, the criteria you have internalized, or how it feels? Could you tell which one you actually used? - In a study of 288 K-12 teachers, a self-reported measure of [[ai-literacy|AI literacy]] correlated with an objective measure of the same skill at only r = 0.07 to r = 0.24. Where in your own practice would a self-estimate most likely diverge from a test of the same skill, and why there rather than somewhere else? - Self-grading can save instructor time and build judgment, and it can also hand out a mark the student has not earned. What would have to be true about the task, the rubric, and the training for a self-grade to carry credit? - If a tool can produce the essay, what is left for a student to assess about their own learning? What would you ask them to estimate instead? - Students who assess a peer's work and students who assess their own are doing different things. Which one transfers better to the next task, and what evidence would settle it? ## Introduction Self-assessment is one of the older ideas in education and one of the more confused ones in practice, because two distinct things travel under the name. The first is a technique: students reviewing their own work against criteria, grading it, checking it, reflecting on it. The second is a measurement: the learner's own rating of their skill, confidence, or learning, used as data. The technique is judged by what it develops in the learner; the measurement is judged by whether the number can be trusted. Both matter in an AI-era [[assessment|assessment design]], and they matter for different reasons. The technique is where [[evaluative-judgment|evaluative judgment]] is trained, and evaluative judgment is exactly the capability that becomes scarce when a [[llm|language model]] can produce competent-looking work on demand. The measurement is where the field routinely overstates what it knows, because a self-report cannot measure learning: it can only record what someone says about it. ## Self-assessment as an instructional technique The practices that fall under this heading are familiar: self-grading against a rubric, self-review of a draft before submission, self-checking of problem solutions, reflective prompts asking what a learner would do differently, and prediction tasks where a student estimates their performance and then compares. Their purpose is developmental. A student who reviews their own draft against criteria is rehearsing the appraisal that peer review and instructor feedback also train, and is building the internal standard they will use when nobody is watching. Evidence that the technique works, and about which version works, comes mainly from writing and from self-regulated learning research. In a comparative study of feedback literacy, first-year undergraduates (N = 118 in China, 56 in a GenAI group and 62 in a peer group) worked through three self-assessment cycles across one semester. The GenAI group's advantage over the peer group was small, only 0.17 scale points (partial eta squared = 0.03), and the study's practical conclusion is the one that recurs: self-assessment needs scaffolding. Pre-trained rubrics, prompt guidelines, and worksheets are what separate a self-assessment cycle that builds [[feedback-literacy]] from one that produces [[cognitive-offloading|offloading]]. Self-assessment is also where the cycle of [[self-regulated-learning]] becomes visible as a skill rather than a disposition. In an undergraduate statistics study, students who received explicit training in reasoning-focused scaffolding (stepwise hints and verification prompts) performed better on a later task attempted with no model assistance and showed better self-assessment calibration than students who leaned on the model uncritically. The interesting part is the second measure: the trained group's estimates of their own understanding aligned better with what they could actually do. A technique aimed at reflection changed a measurement property of the learner. ### Calibration training The most direct technique is calibration training, where a student predicts their score or confidence before seeing a result and then compares the prediction with the outcome. In an AI-assisted writing study, a feedback literacy script improved writing quality, feedback uptake, and deep revision, while an Assessment-Performance Calibration Activity primarily improved the accuracy of students' self-assessment and reduced overconfidence. Combining the two produced the largest writing gains and the strongest retention after AI support was withdrawn, though the calibration activity alone remained the better route to accurate self-estimates. That division of labor is the practical point: feedback literacy improves what a student does with comments, calibration improves how well they know their own standing, and they are not the same intervention. An essay-level paper on [[ethics|ethical]] AI use at one Australian university treated guided self-assessment as the vehicle rather than the subject: across commencing nursing, health sciences, engineering, and science students from 2021 to 2025, pre- and post-semester self-assessments combined Likert confidence items with an open reflection prompt, and the authors argued that [[ethics|ethical]] practice is a developmental capability to be built through reflection rather than a compliance rule to be enforced. ### What AI adds to the technique AI changes three things about the technique. It can generate the items and rubrics, which makes repeated self-assessment cycles cheap. It can be the thing assessed, which is how [[evaluative-judgment|evaluative judgment]] gets trained: a student judging an AI draft is practicing appraisal with an unlimited supply of material. And it can make the learner's judgment visible as process evidence. In interactive learning dashboards, a conventional analytics dashboard was extended with a Judgment of Learning self-assessment feature and a conversational agent, on the argument that eliciting a learner's own judgment does more for engagement than telling them what the data shows. Institutional frameworks have started to treat the same process traces, including structured self-assessments and metacognitive prompts, as evidence of adaptive capabilities that a transcript cannot show. ## Self-assessment as a measure Used as a measure, self-assessment usually arrives dressed as something else: a confidence rating on a Likert scale, a predicted score, a Judgment of Learning item, or a taxonomy-based inventory of perceived skill. It is one family within the wider set of [[self-report-measures]], and it inherits that family's categorical limit. A self-report cannot measure learning or behavior; it can measure only what a person is willing to say about their learning or behavior, and those two things come apart often enough in this research base to be a finding in its own right. ### How accurate the estimates are Accuracy is the headline problem, and the direction is consistent. The cleanest evidence comes from a study that built parallel self-report and objective measures of teacher [[ai-literacy|AI literacy]] inside a single framework. Across 288 K-12 teachers, correlations between the objective and self-reported factors ranged from r = 0.07 to r = 0.24, and latent profile analysis found six profiles: 43 teachers rated themselves consistently high while scoring lower on the objective measure, 59 showed the reverse pattern, and the remaining profiles clustered near the mean or split by prior AI literacy experience. Two measures of one skill, built by the same team on the same construct, share almost nothing. The field-level picture matches that single study. [[assessing-teachers-ai-literacy-measurement-tools-2026|Zainal, Mohd Matore and Maat's 2026 review of teacher AI literacy measurement tools]] appraised 33 instruments published between 2019 and 2025 and found that 31 (93.9%) were self-report scales of perceived confidence, only two (6.1%) tested knowledge objectively, and none used performance-based tasks. The authors' reading is the one this page's measurement face turns on: a self-report score records reported confidence rather than capability, so it cannot stand in for competence, and they argue for performance tasks and item-response-theory analysis alongside self-assessment to separate validated capability from reported confidence. Accuracy also depends on what is being estimated. A psychometric analysis of a taxonomy-based GenAI literacy self-assessment with 158 university staff and students found an inverted profile, with respondents claiming mastery of creation before the conceptual foundations underneath it, and only a weak correlation (r = 0.188) between student and academic profiles. A survey of teacher-education students found the same shape from the other side: nearly all (97.8%) rated critical media analysis important while only 13.8% said their coursework addressed it, and fewer than half (44.9%) believed they had the skills to analyze media information critically. Self-assessed media competence lagged self-assessed importance. Two mechanisms explain the generosity. The first is motivational and well documented outside this literature: people rate themselves generously, and the least competent overestimate most. The second is specific to AI-native cohorts. A 2026 theoretical paper on the **absent cognitive baseline** argues that sustained substitutive AI use during the formative secondary and high-school years reduces the independent cognitive encounters that any academic self-assessment depends on. The claim is structural rather than individual: self-assessment works by comparing a current performance with a record of previous ones, and where the record was never built, the estimate has nothing to be calibrated against. That is presented as distinct from [[cognitive-offloading|offloading]] and [[cognitive-surrender|surrender]], which describe processes during AI use, because the gap persists when the tool is absent. Developmental stage sets one boundary on accuracy, and the youngest learners are where a validated self-assessment instrument is scarcest. [[ai-literacy-self-assessment-questionnaire-primary-2025|Thianwan and Srikoon's 2025 validation study of an AI literacy self-assessment questionnaire for upper primary students]] built a 15-item measure for Grades 4 to 6 across Learning About AI, Learning About How AI Works, and Learning for Life with AI, confirming a three-factor structure in samples of 335 and 579 students with an overall Cronbach's alpha of .934. The authors are explicit about the genre's limit: self-assessment accuracy depends on metacognitive ability that is still maturing in children, so scores index perceived understanding rather than demonstrated competence, and the instrument is intended for formative diagnosis. ### Self-assessment as an outcome variable The measurement face also shows up in evaluation design, where self-assessment instruments are used as outcome measures. One study of AI-assisted assessment of complex reports in higher education explicitly evaluates learning with validated feedback literacy and self-assessment instruments alongside subsequent performance and comparison with a control group. That is defensible when the instrument measures a belief the study is actually about, such as [[self-efficacy]] or perceived competence, and misleading when it is treated as a stand-in for achievement. The recurring critique of [[ai-ed-evaluation|AI intervention studies]] applies with full force here. ## What generative AI changes The two faces converge under generative AI, because the tool attacks the link both of them rely on: that a learner's submitted work is evidence about the learner. Where the work can be produced on demand, a sound essay no longer demonstrates the judgment behind it, and asking a student to assess a draft they did not write is a different task from assessing one they did. This is why [[assessment-validity|assessment validity]] arguments have moved from detection toward redesign, and why self-assessment moves from a study aid to a component of the assessment design itself: the judgment a student makes about a piece of AI output, on the record and with reasons, is evidence in a way that the output is not. It also introduces a naming hazard worth stating plainly. **AI self-evaluation** is not this concept. When a paper reports that a model assesses its own reasoning or grades its own output, that is model evaluation and belongs with [[automated-assessment]] and the study of [[pedagogical-llm-training|model training]] and simulated learners; one tutoring study makes the practical version of the point, arguing that feedback should be conditioned on a separate diagnostic classifier rather than on the model's self-assessed reasoning validity. Learner self-assessment and model self-assessment share a vocabulary and nothing else. ## Implications for practice - **Scaffold the technique.** Rubrics, prompt guidelines, and worksheets are what make a self-assessment cycle developmental. Bare self-assessment, like bare AI feedback, tends to produce reassurance rather than judgment. - **Train calibration separately from feedback use.** Prediction-and-compare activities change the accuracy of self-estimates; feedback-literacy work changes what students do with comments. If accuracy is the goal, run the calibration activity. - **Keep self-grades low-stakes unless the training is real.** A self-grade is a claim about learning made by the person it benefits, and the overestimation profiles in the AI-literacy study are not a small bias to absorb in a summative mark. - **Use self-report where it is honest.** Confidence, perceived competence, and willingness to disclose AI use are legitimate things to measure by asking. Achievement is not. - **Ask for judgment, not opinion.** The most defensible AI-era self-assessment asks a student to appraise a specific piece of work against criteria and justify the call, which is trainable, inspectable, and connected to [[evaluative-judgment]]. - **Watch the cohort effect.** For students whose formative years included substitutive AI use, an accurate self-estimate may be unavailable rather than merely optimistic, which shifts the instructional task from correcting overconfidence to rebuilding the experiential record it needs. ## Connected Concepts - [[evaluative-judgment]] - [[self-report-measures]] - [[peer-assessment]] - [[metacognition]] - [[self-regulated-learning]] - [[feedback-literacy]] - [[formative-assessment]] - [[assessment-validity]] - [[self-efficacy]] - [[scaffolding]] - [[cognitive-offloading]] - [[cognitive-surrender]] - [[trust-calibration]] - [[assessment]] - [[learning-gains]] - [[educational-measurement]] - [[ai-literacy]] ## Connected Articles - [[pedlow-genai-selfassessment-2026]] — Pre- and post-semester self-assessments on ethical GenAI use across nursing, health sciences, engineering and science cohorts (Pedlow et al. 2026) - [[rethinking-ai-writing-feedback-literacy]] — Feedback literacy scripts versus calibration training in AI-assisted writing (2026) - [[scaffolding-srl-feedback-genai-human-peers]] — Three self-assessment cycles comparing scaffolded GenAI feedback with peer feedback, N = 118 (2026) - [[guided-llm-scaffolding-independent-learning]] — Verification-focused scaffolding improved independent performance and self-assessment calibration in statistics (2026) - [[absent-cognitive-baseline-2026]] — Academic self-assessment without the experiential record to calibrate against (2026) - [[ai-literacy-assessment-misalignment]] — Self-reported and objective measures of teacher AI literacy correlate only weakly, r = 0.07 to 0.24 (Zhang et al. 2026) - [[genai-skill-bypass-literacy]] — Rasch analysis of 158 GenAI-literacy self-assessments: an inverted skill profile (2026) - [[self-directed-growth-generative-ai-learning-analytics]] — Self-assessment placed at the center of a self-directed growth framework (2026) - [[tripartite-feedback-framework-ai-assessment-2026]] — Validated self-assessment instruments used as learning outcomes in AI-assisted assessment (2026) - [[ai-feedback-enactment-workflow-2026]] — Enacting AI feedback raised uptake and self-assessment confidence (2026) - [[interactive-learning-dashboards-engagement]] — A Judgment of Learning self-assessment feature inside an interactive dashboard (2026) - [[lodge-adaptive-capabilities-genai-future-2026]] — Structured self-assessments as process evidence for adaptive capabilities (Lodge et al. 2026) - [[critical-media-literacy-education-2026]] — Self-assessed media competence lags perceived importance (2026) - [[age-tiered-ai-literacy-guidebooks-2026]] — Developmentally tiered AI literacy materials, with measurement of acceptance and validity (2026) - [[chatgpt-critical-creative-thinking-review]] — Triangulating AI feedback with peer, instructor, and self-assessment (2026) - [[yasir-llm-tutoring-agents-2026]] — Why feedback should not rest on a model's self-assessed reasoning validity (Yasir et al. 2026) - [[ai-literacy-self-assessment-questionnaire-primary-2025]] — A validated 15-item AI literacy self-assessment questionnaire for Grades 4 to 6, with the metacognitive limits of child self-report (Thianwan & Srikoon 2025) - [[assessing-teachers-ai-literacy-measurement-tools-2026]] — 31 of 33 teacher AI literacy instruments are self-report, two test knowledge objectively, none use performance tasks (Zainal, Mohd Matore & Maat 2026) --- ## [Oral Assessment](https://edtechdev.github.io/aied/concepts/oral-assessment/) > **Oral assessment** — assessment in which a learner must explain, defend or demonstrate understanding in speech, live or recorded: the viva voce, the oral exam, the oral defense, the code-review interview, the clinical consultation. Because the exchange is real-time and the questions need not be known in advance, the format resists the substitution of generated text for understanding in a way that take-home [[authentic-assessment|authentic tasks]] cannot. It also carries a distinctive validity problem: orals measure fluency and confidence alongside knowledge, so what they certify depends heavily on how they are designed, who is asking, and how the learner's speech is treated as evidence. ## Questions to Consider - Orals are promoted in the AI era because a machine cannot sit in the room and answer for you. Yet the corpus that argues this contains no randomized trial of oral against written assessment. What would you need to see before treating that argument as settled? - The same body of work reports two different mechanisms for reduced anxiety: removing the teacher's live observation, and becoming familiar with the format through practice. These imply different designs. Which fits your context, and which would you trust to generalize? - Speech-only scoring can mistake verbal fluency for conceptual understanding. If a learner gestures correctly but says the wrong word, or says the right word without understanding it, what exactly is your assessment measuring? - Every oral design in this corpus hits the same wall: staff and room time. If the constraint is twenty-eight contact-hours per semester for fifteen students, is the honest answer a smaller number of defended assignments, a larger teaching team, or a different format? ## Introduction Oral assessment is one of the oldest assessment formats and, for most of the last century, one of the least fashionable: hard to scale, hard to standardise, hard to defend in an appeals process. Generative AI has reversed its fortunes. When a written submission can be produced on demand by a model, the assessment value of the artifact collapses, and attention moves to the assessment of the person. A learner who must answer an unfamiliar question out loud, in real time, is at least demonstrably present and thinking. That reversal is now visible across the research corpus, but unevenly. Some work treats the oral exam as a policy answer to [[academic-integrity|academic integrity]], some designs an automated system to deliver or score oral performance at scale, and some asks what speech itself can and cannot show about understanding. The pages connected below disagree about how much orals actually prove. Read together, they support a narrower and more defensible claim than the one usually made for them. ## What a live performance can certify [[fenton-oral-exams-ai-authentic-assessment-2025|Fenton (2025)]] makes the strongest form of the argument. Drawing on the literature on [[authentic-assessment|authentic assessment]], the paper's case for the oral exam rests on its interactivity: questions can be withheld until the moment of the exam, follow-up probes can be improvised, and the learner cannot rehearse a response to a question they have not seen. The format also blocks rote memorisation, because recall is checked through explanation rather than reproduction. The paper is a review and position piece in *Educational Researcher*, not an empirical study, and it ends with thirteen numbered implementation recommendations rather than evidence of improved outcomes. The honest version of the integrity claim is narrower than it is often stated. Orals make substitution harder to hide; the corpus does not show that they reduce AI use. In [[code-review-genai-cs1|the CS1 oral code-review study]], weekly fifteen-minute interviews carried seventy per cent of each assignment's grade, and pasted-to-total characters still rose from 61.0 per cent to 68.1 per cent (p < 0.0001) across three semesters, while exam scores moved by a statistically insignificant two per cent. Ninety per cent of students said the reviews motivated them to understand their code better and sixty-five per cent that they helped avoid over-reliance on AI, but the measured behavior did not follow. Oral assessment changed what could be hidden, not what students did. ## The design space the corpus actually covers The formats in the corpus span a wide range of automation and synchronicity, and each solves a different part of the problem. **Live, human, synchronous.** The viva voce examined by [[aivaluate-anxiety-assessment-2026|a viva voce study of thirty-five pre-university students]] is the baseline: same teacher, same questions, one delivery face-to-face and one mediated by an AI system. The doctoral viva appears in [[pgr-students-genai-uses-qualitative-2026|a qualitative study of fifteen postgraduate researchers]], where participants argued for shifting assessment weight toward it on the reasoning that a thesis is easier to fake and that the viva must be given in person. The authors accept that reasoning only in part, since a viva is not inherently proof against AI assistance. **Recorded and asynchronous.** [[asynchronous-oral-assessment-2026|Asynchronous oral assessments]] replace the room with a time-limited, non-revisable webcam response, roughly thirty seconds of preparation and two to three minutes of speech. In the reported pilot, scores exceeded those of in-person multiple-choice tests: midterm median 92.5 against 70 (p < .001) and final median 94.2 against 86.4 (p = .002), with only moderate cross-format correlations. The authors explicitly attribute the difference to the format rather than to learning, and the design carried 7.5 per cent of the course grade. **Automated delivery and scoring.** [[ai-supported-oral-assessment-tvet-2026|A voice-based system trialled in vocational automotive and engineering classes]] assessed thirty-three learners, of whom twenty-one agreed the live voice task felt realistic and none disagreed that speaking in real time suited the task better than a written portfolio. Word counts for identical questions varied five- to eightfold between learners, fluency conferred no accuracy advantage on closed items, and a preliminary agent grade matched the human tutor's in ninety-five per cent of adjacent horticulture and dairy trials. [[socratic-tests-conversational-assessment|A conversational Socratic test]] takes the opposite design route, aiming at conceptual questioning; its evidence is a self-report survey of ninety-eight students, in which 80.6 per cent agreed the AI scaffolded effectively and fifty-two per cent reported lower stress than traditional exams. Both papers measure acceptance rather than learning. **Oral evidence as defense rather than examination.** [[tool-invariant-framework-agentic-ai|A tool-invariant assessment framework]] pairs AI-free in-class quizzes with ten-minute oral defenses of comment-stripped, AI-assisted work, scored on code comprehension, method understanding, terminology, interpretation and verification, with verification required to pass regardless of the total. The design is argued rather than validated; the paper's own arithmetic for fifteen students is two and a half contact-hours per assignment and roughly twenty-eight contact-hours per semester across eleven defended assignments. ## Anxiety, bias, and who the format disadvantages Orals are argued for on inclusion grounds, because a spoken answer is harder to buy than a written one, and against on the same grounds. [[fenton-oral-exams-ai-authentic-assessment-2025|Fenton]] catalogues the challenges directly: scheduling load, anxiety, and bias along gender, ethnicity, language and answering speed, plus the effects of non-anonymous marking. Against that, the evidence the paper cites suggests orals can be as inclusive as written examinations, including for students with dyslexia, and that unfamiliarity rather than the format itself drives most reported anxiety; students in one cited study were less anxious by their later oral exams. Where anxiety genuinely falls, the reason matters. In the AI-mediated viva, self-reported calmness was significantly higher than in the face-to-face version (means 6.50 against 5.86, t(34) = −1.97, p = .028), and usability was rated good. But the face-to-face condition scored significantly higher on helping students understand their own work (p = .004). Removing the teacher's live observation made students calmer and, by their own report, less illuminating to themselves. What counts as oral evidence is also narrower than it looks. [[multimodal-embodied-cognition-oral-explanations-2026|A study of embodied evidence in oral explanations]] argues that scoring speech alone mistakes verbal fluency for conceptual knowledge and disadvantages learners with language-related difficulty, since gesture carries understanding that the transcript loses. Its demonstration is small: two engineering students explaining statistical concepts, with high-confidence gestures clustering on particular ideas, square and box forms at roughly fifty-four to sixty-one per cent, and tighter gesture–speech coordination accompanying more coherent explanations. ## Scoring, validity, and the scale problem Three findings cut against reading oral scores as learning gains. [[asynchronous-oral-assessment-2026|The asynchronous-format study]] disclaims learning gains for its own score advantage, and reports instructor-versus-model re-scoring agreement of ICC 0.73 at midterm and 0.60 at final, which bounds how far machine scoring can be trusted. In [[ai-standardized-patient-scaffolding-medical-2026|an RCT with one hundred third-year medical students]] comparing a simulated-patient system against progressive-disclosure case material, final examination performance rose (71.8 against 55.6 per cent, Hedges' g = −0.81) and OSCE communication ratings rose (3.53 against 2.64 on a five-point scale), while binary diagnostic accuracy was statistically identical (84 against 86 per cent, P = 1.000). An accuracy-only reading would call the intervention null; a communication-only reading would call it transformative. The binding constraint is staffing and hardware rather than pedagogy, and the corpus is unusually candid about the arithmetic: twenty-eight contact-hours per semester for fifteen students, eleven teaching assistants for a class of over one hundred in the CS1 design, and one laptop serving twelve simultaneous learners in the offline vocational deployment. No randomized trial of oral against written assessment exists anywhere in this body of work, and every study that measures the difference is a single-institution design without a control group. Oral assessment is well supported as a response to AI-assisted substitution, and thinly supported as an improvement in learning. ## Connected Concepts - [[assessment]] — the broader field this format sits inside - [[authentic-assessment]] — the design tradition that makes the integrity case - [[assessment-validity]] — what a format can and cannot claim to measure - [[academic-integrity]] — the pressure that returned orals to prominence - [[automated-assessment]] — machine delivery and scoring of oral performance - [[ai-detection]] — the alternative response, and why it is weaker - [[speech-and-voice-technologies]] — the speech pipeline an oral system runs on - [[multimodal]] — gesture and speech as combined evidence - [[anxiety-and-stress]] — the affective cost of live performance - [[feedback]] — what a viva tells a learner about their own understanding - [[higher-ed]] — the setting almost all of this evidence comes from ## Connected Articles - [[fenton-oral-exams-ai-authentic-assessment-2025]] — Reconsidering the use of oral exams and assessments (Fenton 2025) - [[asynchronous-oral-assessment-2026]] — Asynchronous oral assessments: integrity, engagement and professional communication - [[ai-supported-oral-assessment-tvet-2026]] — Designing AI-supported oral assessment in vocational education - [[aivaluate-anxiety-assessment-2026]] — Anxiety and experience in AI-mediated performance-based assessment - [[code-review-genai-cs1]] — Oral code review interviews in an introductory programming course - [[multimodal-embodied-cognition-oral-explanations-2026]] — Gesture as evidence in assessing oral explanations - [[socratic-tests-conversational-assessment]] — Automated conversational testing as oral assessment - [[tool-invariant-framework-agentic-ai]] — Oral defenses of AI-assisted work - [[ai-standardized-patient-scaffolding-medical-2026]] — Spoken clinical interviews under a simulated-patient system - [[pgr-students-genai-uses-qualitative-2026]] — The doctoral viva under generative AI --- ## [Automated Assessment](https://edtechdev.github.io/aied/concepts/automated-assessment/) > **Automated assessment** — the use of AI to evaluate student work, from [[formative-assessment|formative quizzes]] to high-stakes exams. Automated assessment spans multiple modalities — multiple-choice, short answer, essay, code, and performance-based evaluation — and ranges from direct automated grading to confidence-aware systems that report calibrated uncertainty alongside their scores. ## Questions to Consider - Automated assessment ranges from multiple-choice scoring to essays, code, and performance-based evaluation. What's the biggest difference you'd expect between how reliably AI can grade a multiple-choice quiz versus a free-form essay — and why? - A core design idea here is 'confidence awareness': AI graders that report how certain they are, flagging low-confidence cases for human review rather than issuing a single unqualified score. How would a grade that said 'I'm 70% sure about this' change how you'd use or trust an automated score? - Research shows automated scoring can systematically disadvantage non-native speakers — the AI scores the language rather than the understanding. Why might an automated grader be especially prone to this kind of unfairness, even when it achieves high agreement with human raters overall? - One study found that how validation is done can inflate reported performance: a naive cross-validation method reported near-perfect results that dropped dramatically under more rigorous trial-independent validation. What does this cautionary lesson suggest about how you should read any claim that an AI assessment system 'works'? - The page argues calibrated confidence enables human-in-the-loop workflows, supports trust calibration, and strengthens measurement validity. If an automated system routed its most uncertain cases to a human reviewer, what would you want to know about how those cases are chosen before you trusted the split? - Automated grading is described as one of the most mature AIED applications, yet grading without useful feedback has limited educational value. How might a focus purely on producing scores — rather than usable feedback — change what students actually get out of an AI-graded assessment? ## Introduction ### Assessment modalities - **Short answer and essay:** automated grading, [[automated-essay-scoring]], and [[cong-confidence-asag-2026|confidence-aware approaches]] handle free-text evaluation. - **Code assessment:** [[automated-grading-linux-bash-examinations-large-language-models|Bash grading]] and [[code-review-genai-cs1|code review]] demonstrate programming assessment. - **Formative assessment:** [[automated-formative-assessments-a-level-sciences|A-level science automation]] and [[cotal-formative-assessment-scoring-2026|CoTAL]] focus on formative rather than [[summative-assessment|summative]] use. - **AI-generated assessments at scale:** [[assessing-quality-ai-generated-exams-field-2025|Assessing AI-Generated Exams]] shows that iteratively refined, course-tailored AI-generated exams achieve [[item-response-theory|IRT]]-measured quality on par with expert-written standardized-exam questions (difficulty β̄ = −0.45 vs. 0.35; discrimination ᾱ = 1.3 vs. 1.2) across 91 college classes — evidence that automated assessment can move from scoring to full [[automated-question-generation|item generation]]. - **Performance assessment:** [[engagement-assessment-video|Video engagement assessment]] and [[confidence-aware-student-drawing-assessment|drawing assessment]] extend automation beyond text. - **Multimodal exam data and rubrics:** [[multimodal-exam-obe-rubrics-2026|a multimodal examination answer dataset with expert-designed Outcome-Based Education rubrics]] provides a benchmark resource for criterion-level automated assessment across diverse response modalities, supporting [[benchmark|benchmarking]] and [[educational-measurement|measurement]] research on multimodal student work. - **Neurophysiological assessment:** [[eeg-familiarity-automated-assessment-2026|Nanayakkara & Halloluwa (2026)]] benchmark ML/DL models for EEG-based familiarity prediction (faces vs. math equations) as a step toward direct, objective measures of knowledge acquisition. Crucially, they show standard stratified cross-validation inflates performance (up to 0.9853 F1) via temporal leakage, while trial-independent Group K-Fold validation drops the peak to 0.6038 F1 — a cautionary [[research-methods-aied|methodological]] lesson for all automated-assessment benchmarking. ### Automated grading Automated grading is one of the most mature and widely-deployed [[ai-education|AI in education]] applications — [[ai-technologies|AI systems]] that evaluate student work, from multiple-choice scoring to essay assessment and code review. Its grading modalities include: - **Students separate feedback utility from evaluative authority:** [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]] documents a case where an AI score counted toward grades and students knew it — 13 computing students whose scanned handwritten work was evaluated by ChatGPT against a rubric covering clarity, organization, sentence correctness and conciseness. Every participant found the feedback useful, yet all separated that utility from authority, and several rejected the one-way, non-dialogical character of automated grading — "ChatGPT is not like people. You talk to them and explain excuses... I like normal teacher [to] grade my work." The study positions automated evaluation as acceptable for surface revision but not as a final assessor, with conditional trust mediated by instructor oversight. - **Short answer grading:** [[cong-confidence-asag-2026|confidence-aware ASAG]] evaluate free-text responses, with confidence calibration critical — systems must know when grading is reliable. A scoping review of short-answer auto-marking in [[science-education|science]] (2017–early 2024) documents the field's history: [[auto-marking-short-answer-science-2026|Morley et al.]] found BERT-family models (base, RoBERTa, DistilBERT, SciBERT) dominated — used in 20 of 21 studies, peaking in 2021 — before GPT-based approaches were adopted via [[prompt-engineering]] rather than fine-tuning from roughly 2022. Models augmented with domain data (textbooks, [[feedback|rubrics]], further pre-training) consistently outperformed those without, yet few could justify marks in human-comprehensible terms and [[bias-mitigation|bias]] across demographic and linguistic groups was rarely examined — leading the authors to argue auto-markers should *support* rather than replace [[teacher-role|human examiners]]. A PRISMA-guided [[meta-analysis-systematic-review|systematic review]] of 42 empirical grading and feedback studies (2023–2025) generalizes this caution across the field: LLMs match human raters on closed-ended tasks and short-answer questions but cannot fully replace human judgment on complex, open-ended, or subjective work requiring in-depth analysis or [[creativity]], and the highest grading effectiveness is achieved in hybrid systems that pair AI-driven grading with teacher oversight and verification ([[jukiewicz-chatgpt-teacher-assessment-feedback-2026]]). - **Essay scoring:** [[automated-essay-scoring]] systems like [[choi-anchor-aes-prompting-2025|anchor-based AES]] use [[prompt-engineering|prompting strategies]] to approach human-level reliability. [[aiawe-automated-writing-evaluation|AIAWE]] extends automated evaluation to broader writing assessment. - **Code review:** [[automated-grading-linux-bash-examinations-large-language-models|Linux Bash grading]] and [[code-review-genai-cs1|CS1 code review]] demonstrate automated assessment in [[cs-education|computing education]]. - **Formative integration:** [[automated-formative-assessments-a-level-sciences|A-level science automation]] and [[cotal-formative-assessment-scoring-2026|CoTAL]] show how automated grading feeds into [[formative-assessment]] cycles. - **Bias and fairness:** [[ai-scoring-language-bias-physics|Language bias in physics scoring]] documents how automated grading can disadvantage non-native speakers — connecting to [[bias-mitigation]] and [[equity-in-ai-education]]. - **Benchmarking GenAI models for open-ended grading:** [[pecuchova-automated-grading-open-ended-genai-2026|Pecuchova, Benko & Drlik (2025)]] benchmarked eleven GenAI and sentence-embedding models against two expert human graders on 1,885 open-ended responses to 24 software-engineering questions. Only GPTo1 reached almost-perfect agreement (Fleiss' Kappa 0.82, low false positives and false negatives across grade categories), with Claude3 and PaLM2 close behind; context-sensitive [[generative-ai|GenAI]] models robustly handled short, diversely-phrased student answers, while reference-based embedding models (e.g., BERT's 345 false positives) systematically penalized correct-but-divergently-worded responses — evidence that model choice, not just grading format, shapes reliability for free-text [[assessment|open-ended assessment]]. - **Prompt exemplars decide scoring agreement:** [[automated-constructive-assessment-hdr-llm-2026|Takahashi et al. (2026)]] automated hierarchical diagnostic reasoning (HDR), an applied-competence task in which learners find and explain deliberately embedded errors in a short case. Using GPT-4o with no fine-tuning, the team generated new problems and scored student descriptions: generated and human-authored problems carried the same internal consistency (Cronbach's α = 0.78) and showed no significant score-distribution difference across 100 participants, and the format's correlation with a reading-comprehension control stayed low (α = 0.36 and 0.41). Scoring accuracy depended on the prompt rather than the model — without worked examples Q2 precision fell to 0.42, while adding human-scored correct and incorrect exemplars lifted agreement to 98% or higher on every question (Q1 and Q2 at 100%) — against a measured human cost of 138 minutes of expert scoring for 117 responses plus 30 minutes of consensus. - **Mixed-format exam grading (2026):** [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat, Das, Bhaumik & Thambi (2026)]] evaluated ChatGPT-5 against human faculty grading of a 21-item pharmacy exam (16 students) spanning multiple-choice, select-all-that-apply, fill-in-the-blank, listing, short-answer, and essay items. Concordance was substantial-to-near-perfect for objective item types (CCC 0.935–1.000 — aided by correct answers supplied during grading) but collapsed for listing (0.621–0.708), short-answer (≈0 to negative), and essay (0.341–0.854) items; providing a structured rubric did not consistently improve full-exam agreement (71.1% without vs 68.2% with). The study sharpens a methodological point that recurs across automated assessment — moderate percent accuracy frequently coexists with low [[assessment-validity|concordance]] — so accuracy alone overstates reliability, and it recommends [[human-in-the-loop-ai|hybrid grading]] for complex, subjective, or high-stakes items. - **Cross-family grade-band calibration on authentic marked essays (2026):** [[llm-grade-bands-calibration-bias-2026|Kerwat, Donaldson & Mahomed (2026)]] ran eight pre-specified configurations from four model families (GPT-OSS 20B, GPT-OSS 120B, Qwen 32B and Llama 3.1 8B, each at two temperatures) over the same locked cohort of 114 BAWE/Coventry essays carrying institutional Merit and Distinction labels, with one shared prompt that permitted four bands. Exact band agreement ranged from 18.4% to 54.4%: Llama 3.1 8B at temperature 0.5 led (54.4% exact, 98.2% within one band, MAE 0.474, near-zero bias) while GPT-OSS 120B was weakest (18.4%) and systematically *undergraded*, by an average of 1.316 bands at temperature 0.5, with Distinction accuracy trailing Merit accuracy throughout. The locked cohort contained no Pass or Fail references, so this is Merit–Distinction discrimination rather than reliability across the full grade range, and temperature was no cure — it helped one family descriptively, left another unchanged, and significantly worsened GPT-OSS 20B (39.5% to 29.8% exact). Two lessons transfer beyond the study: grade distance and directional bias reveal what agreement statistics hide, and model size is no proxy for [[assessment-validity|assessment validity]]. - **Rubric engineering for open-ended scoring in medical education:** [[olvet-genai-scoring-open-ended-medical-2026|Olvet et al. (2026)]] asked whether GPT-4 could reliably score open-ended questions on pre-clerkship [[medical-education|medical]] exams. After three iterations of human-driven rubric refinement at two US schools, inter-rater reliability with faculty reached substantial-to-almost-perfect agreement for three of four questions using analytic and holistic rubrics (weighted kappa up to 0.94), while the holistic-rubric item stalled at moderate (κw = 0.54). Error-pattern analysis showed discrepancies were traceable to both raters — GPT-4 over-scored when students offered multiple answers or used rubric-absent vocabulary, whereas faculty were often "overly generous" graders — leading the authors to recommend keeping [[human-in-the-loop-ai|humans in the loop]] (e.g., faculty scoring a subset to confirm accuracy). Because roughly 82% of US medical schools grade pre-clerkship work pass/fail, exact AI score agreement is often unnecessary for operational use. - **HITL scoring for a large-scale national writing assessment (2026):** [[human-in-the-loop-ai-scoring-national-assessment-2026|Curi et al. (2026)]] scale prompt-based [[llm]] scoring to a real high-stakes Spanish writing exam (~5,000–6,000 responses/year), achieving 60–80% item accuracy and 90%+ output consistency across a 15-item analytic rubric, with a deterministic NLP checker replacing the LLM on the spelling item (100% consistency). Its distinctive contribution is decision-oriented validation over model-centric metrics: the AI's systematic under-grading bias, converted via [[item-response-theory|IRT]]-and-Bookmark into proficiency levels, produces pass/fail discrepancies in 15.3–16.5% of cases — all of which the [[human-in-the-loop-ai|human-in-the-loop]] workflow routes to expert review, cutting responses needing full human scoring by at least 50% while preserving decision quality. This grounds hybrid grading in operational assessment rather than proof-of-concept datasets. - **Per-dimension agreement shows where an assessor cannot detect change:** [[sophie-clinical-communication-ai-assessment-2026|SOPHIE 2.0 (Hasan et al., 2026)]] validated an LLM judge against the consensus of all human rating sources — standardized patients and third-party raters — on a three-dimension clinical communication rubric, reaching Pearson 0.759 and ICC(A,1) 0.746 from transcripts alone, inside the spread of the individual human raters rather than outside it. Its weakest dimension was Be Explicit (correlations of 0.414–0.598 across candidate judges), and that was also the only dimension whose scores did not move measurably across the two encounters (Δ = 0.014, p = 0.1204) — so per-dimension agreement can indicate in advance which dimensions an automated assessor is unable to show improvement on. ### Confidence-aware assessment A central design goal within automated assessment is **confidence awareness**: AI assessment systems that report calibrated uncertainty alongside their scores, rather than issuing a single unqualified prediction. A confidence-aware grader not only produces a grade or classification but also signals how certain it is, so that low-confidence cases can be flagged for human review and users can calibrate their [[trust]] in the system. This is central to responsible automated assessment and connects closely to [[psychometrically-aware-ai]] and [[trust-calibration]]. **How confidence is modeled** in the knowledge base's research: - **Fused confidence signals for short-answer grading:** [[cong-confidence-asag-2026|Confidence-Aware ASAG]] fuses model-based confidence signals (verbalized, latent, and consistency-based) with dataset-derived aleatoric uncertainty via Random Forest regression. - **Confidence in multimodal student work:** [[confidence-aware-student-drawing-assessment|Confidence-aware assessment of student-drawn scientific figures]] extends confidence modeling to multimodal student responses. - **Psychometric calibration of LLMs:** [[psychometrically-aware-ai|Psychometrically aware AI]] advances the standard of aligning [[llm]] scoring with measurement theory, with calibration as a core requirement alongside [[item-response-theory]] alignment. - **Difficulty and response-time calibration:** [[llm-difficulty-calibration-programming-exams-2026|Programming-exam difficulty calibration]] repositions LLMs as auxiliary evidence sources whose difficulty estimates correlate with student pass rates. - **Trait-adaptive essay scoring:** [[psyscore-essay-scoring-zpd-feedback|PsyScore]] shows a psychometrically-aware framework can adapt essay feedback to learner traits. - **Evaluating visual student work:** [[diagramir-educational-math-diagram-evaluation|DiagramIR]] back-translates LLM-generated math diagrams (TikZ) into an intermediate representation with deterministic checks, beating LLM-as-a-Judge on agreement with human raters and letting small models match large ones at ~10× lower cost — a scalable route to assessing non-text, diagrammatic student output. - **Trust-gated inference in automated teacher assessment:** [[li-explainable-trustworthy-llm-teacher-assessment-2025|Li, Yang & Fang (2025)]] place Monte Carlo dropout calibration directly in the scoring path, so that dropout variance above a learned threshold triggers a reject-and-refer output rather than a score, alongside adversarial debiasing that holds the fairness gap to 1.8% where baselines sit in the 6.4–8.2% range and an expected calibration error of 0.032 on TeacherEval-2023. Their ablation is the calibration argument in miniature: removing the trust-gated module drops inter-rater consistency to 78.6%, and the authors attribute a 41% reduction in human review workload to the gating — uncertainty handling as architecture rather than post-hoc reporting. - **[[explainable-ai|Explainability]] of rubric-based scoring:** [[shap-llm-rationales-teaching-quality-assessment|Bueno et al.]] show that model-agnostic SHAP attributions are more faithful and transferable than LLM-generated rationales for explaining rubric-based scores (e.g., classroom feedback quality), and propose deletion-based + cross-model tests as a principled way to evaluate any scoring model's explanations. **Why calibrated confidence matters:** - **Enables human-in-the-loop delegation:** low-confidence cases route to a human reviewer — supporting [[human-in-the-loop-ai|human-in-the-loop]] workflows rather than blind automation. - **Supports trust calibration:** [[trust-calibration|calibrated confidence]] lets users match their trust to the system's actual reliability, avoiding both over-trust and under-trust. - **Improves measurement validity:** confidence-aware scoring strengthens [[educational-measurement]] and [[assessment-validity]] by making uncertainty explicit. - **Fairness and defensibility:** flagging low-confidence cases for review reduces the risk of confidently wrong scores, especially for atypical or underrepresented responses. # - **Rubric generation and instructor-supervised grading pipelines:** [[harmogen-ai-assessment-rubric-generation|Mendonça et al. (2026)]] show that LLM-generated assessment rubrics (HARMOGEN-R) can match human-created rubrics for technical content within a ±5-point equivalence margin, with structured generation giving greater cross-model consistency. [[ai-assisted-instructor-supervised-grading-feedback|Cruz et al. (2026)]] evaluate an end-to-end GPT-4o grading pipeline where AI grades fell within 0.5 points of the instructor in 83% of 362 submissions (MAE 0.31) — best framed as a scalable supplement to, not a replacement for, instructor judgment. ## Quality and fairness Automated assessment quality depends on [[assessment-validity]] and [[bias-mitigation]]. [[ai-scoring-language-bias-physics|Language bias]] research shows that automated scoring can systematically disadvantage certain student populations. ### Deployment scenarios and what each one requires How much alignment an institution should demand depends on what the system is allowed to decide. In a [[mixed-methods-research|mixed-methods]] study of [[opraise-automated-marking-ai-assessment-2026|761 authentic undergraduate essays at three UK universities]], three frontier models with 27 prompt configurations each reached only 35–65 percent agreement with human markers on the degree band, and accuracy did not transfer between institutions — so the report treats candidate uses as distinct scenarios rather than one decision. **Quality assurance of human marking**, where AI marks in parallel and significant differences trigger human review, is the most conservative. **A marking assistant**, ranking or categorizing submissions by predicted quality or uncertainty, triaging complex cases, or expanding brief human comments into fuller feedback, is a genuine middle ground. **AI as primary marker**, with human review of a sample or of mark distributions, was judged acceptable only if system performance improves, and no [[stakeholders|stakeholder]] group endorsed AI as a sole marker. Three requirements follow for any of them. Because alignment is model- and institution-specific, validation must be local and continuous — headline accuracy from elsewhere is not evidence of readiness. Because marks were compressed toward the middle, AI was least accurate at the grade boundaries and for the strongest and weakest submissions, which is where assessment decisions carry the most consequence. And because disagreement between an AI and a human mark reflects different judgment rather than mere error, discrepancies should trigger human interpretation rather than algorithmic override, with final authority over the mark retained by people. Adoption also carries [[governance]] obligations the technical metrics do not capture: a right to explanation under GDPR Article 22 where automated decisions affect students, Equality Act duties once attainment level and language use are shown to influence marks, revised appeal processes for errors that are not easily caught, and an irreversibility risk — once staffing and investment shift, an institution may struggle to keep collecting the human marking it would need to return to. [[tripartite-feedback-framework-ai-assessment-2026|Venetsanos (2026)]] supplies the criteria those requirements presuppose, arguing that what makes automation defensible is the *epistemic status* of the task rather than what the technology can technically perform. His framework limits AI to bounded verification of factual and procedural claims — all four of unambiguous documented criteria, direct comparison against established knowledge, no assessment of alternative valid approaches, and a single correct answer or pre-specified acceptable set must hold simultaneously, with ambiguity escalating to a human by default — and makes AI involvement at that level conditional on five non-negotiable principles that must all hold at once: the knowledge base must be assessor-curated and retrieval-grounded in module materials rather than the model's parametric knowledge; human assessors review every AI output before it reaches students and hold absolute override with unshared accountability; feedback must carry clear provenance and attribution; and security must be designed against adversarial use from the start, with input sanitisation for instructions hidden in white or small text, encodings, images or document metadata, since successful circumvention of an automated system spreads through student cohorts. The same paper is explicit that its principles are untested for simultaneous feasibility and that assessor time for curation, security infrastructure and high-frequency oversight may shift staff effort rather than reduce it — a caution that sits alongside the local-validation requirement above. ### Security of AI-mediated grading A red-team evaluation of an everyday grading workflow shows the attack surface that planning for adversarial use has to cover. [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] embedded five indirect prompt injections inside a synthetic essay that Microsoft Copilot (GPT-5.2) had graded fail in six of six baseline runs, iterating each across docx, pdf and htm files. Two strategies changed the grade with no visible warning — an instruction-manipulation and role-playing paragraph at 9 of 9 iterations (100%) and the same paragraph hidden behind an image layer at 17 of 18 (94%) — while a paragraph in small white text at the end of the document failed in all nine iterations, and file-metadata injections never worked, which the author attributes to metadata access being disabled in the version tested. The trust problem outlasted the exploits: after one pdf run detected an embedded instruction and stated it would grade only against the official assignment, re-running the same file raised the grade six more times with no warning, and a chat blocked by a detected attack was silently disabled rather than reported to the user. The author argues the resulting grades carry no [[assessment-validity|validity]] claim in either direction, that an adversary needs only one working combination against layered [[guardrails]], and that the teacher remains the only real [[human-in-the-loop-ai|check on the output]] while the manipulation is designed not to be visible — connecting directly to the input-sanitisation and adversarial-security obligation Venetsanos sets out above. ### Connections Automated assessment connects to [[assessment-validity]] (quality assurance), [[formative-assessment]] (use context), [[bias-mitigation]] and [[equity-in-ai-education]] (fairness), [[teacher-role]] (how automation changes instructor work), and [[ai-feedback-quality]] (grading without useful feedback has limited educational value). Confidence-aware assessment is a specific mechanism within the broader agenda of [[psychometrically-aware-ai]] and a contributor to [[trust-calibration|calibrated trust]]. - **Large-scale AI grading of handwritten [[physics-education|physics]] (2026):** A multimodal model (GPT-5.5) graded 10,364 scanned pages across a national Physics Olympiad theory exam, selection camp, and university quantum-mechanics exam, achieving total-score correlations of 0.91–0.97 with official marks and recovering the same top-five Olympiad team. Revised page-by-page, evidence-location instructions improved agreement — evidence that [[multimodal|multimodal AI]] can support high-stakes summative grading with careful rubric and prompt engineering ([[ai-grading-handwritten-physics-2026]]). - **Selective automation for handwritten [[chemistry-education|chemistry]] (2026):** [[cvengros-grading-handwritten-chemistry-ai-2026|Cvengros & Kortemeyer]] graded a 296-student handwritten general-chemistry final page-by-page against rubric images with a multimodal, reasoning LLM, achieving high run-to-run reliability on total scores (ICC(A,1) = 0.967) and strong total-score agreement with TA grading (R² = 0.91). Because item-level reliability varied sharply by format — textual and chemical-reaction answers graded reliably while drawing and graphing scored worse than random — raw agreement was judged inadequate for high-stakes use. They operationalize [[human-in-the-loop-ai|human deferral]] through confidence filters (partial-credit thresholds, an [[item-response-theory|IRT]]-based risk threshold, and problem-type exclusion) that convert raw AI scores into an accept/deferral policy, and show that false positives (AI crediting genuinely wrong answers) tend to go undetected since students rarely contest them — a concrete [[psychometrically-aware-ai|psychometrically grounded]] route to selective automation ([[cvengros-grading-handwritten-chemistry-ai-2026]]). - **Semi-open handwritten mathematics in a university exam (2026):** [[gpt4-handwritten-math-exam-grading-2026|Liu et al.]] graded a German undergraduate mathematics exam with GPT-4 across two extraction routes (pre-cut answer boxes versus whole pages), two optical-character-recognition routes and two rubric formats, reporting Krippendorff's alpha against human grades of 0.22-0.44 with accuracy 0.59-0.62 on the best workflows and per-problem accuracy spanning 0.27-0.91 - not acceptable for high-stakes summative use. Itemising multi-criteria grading rules generally did not improve agreement (it helped the partial-credit problem), free grading without a rubric was far worse (accuracy 0.50-0.51), and a probabilistic confidence filter meant to flag grades for human re-checking did not significantly improve accuracy while wrongly flagging 7-41 percent of responses as reliable. The study's transferable lesson is procedural: scan and transcribe the whole exam sheet with markers rather than cropping answer boxes, because students write across box borders; rewrite grading rules written for teaching assistants, since the model attempted unrequested derivations from a rule meant for a human; and note that the human-grading rules themselves paraphrased at an alpha of 0.88 among five variants, so prompt robustness is not the bottleneck. - **LLM comparative judgment for writing screening (2026):** Mercer & Reed used seven LLMs as pairwise comparative judges to score informational writing from 1,208 students in Grades 3–6 across three screening occasions. LLM-CJ scores correlated strongly with analytic rubric scores (r = .67–.73, strongest for Gemini-3.1 Pro) with AUC .82–.86 for proficiency; findings were stable across models, with little validity gain from costlier, more capable ones, while averaging three writing samples improved accuracy — sampling breadth mattered more than model choice. Predictive-bias patterns for [[multilingual-learning|multilingual learners]] matched human rubric scoring, supporting LLM-CJ as an efficient, low-cost screening approach ([[llm-comparative-judgment-writing-screening-2026]]). - **Quality-dependent alignment with instructor grading (2026):** Comparing ChatGPT, peers, and an instructor grading the *same* undergraduate [[group-work|group projects]], [[usher-faraon-who-grades-best-2026|Usher & Faraon found]] ChatGPT's alignment with the instructor *improved as project quality increased* — its largest overestimation was for low-quality work (≈ +14 points), shrinking to ≈ +2.5 points for high-quality projects. ChatGPT also graded higher on average than both peers and the instructor, a grade-inflation tendency that undermines its reliability as a standalone summative grader, especially for weaker submissions. - **Small-corpus agreement statistics can mislead deployment decisions.** Scoring 60 marketing posts (15 students plus 15 researcher-authored low-quality anchors) against two independent human raters gave absolute agreement ICC(2,1) of .435 for the LLM, .266 for an equal-weight rule-plus-LLM hybrid and .091 for deterministic rules, with MAE of 6.28, 10.53 and 17.22 points respectively on a 0-100 scale; the hybrid was significantly worse than the LLM alone (paired ICC difference -.169, 95% CI [-.260, -.106]). Adding anchors raised inter-rater ICC from .338 to .902 and LLM agreement from .435 to .846, and a near-empty post was awarded 75 against a human mean of 30.5, showing that a single degenerate response can dominate a small evaluation. ([[automated-scoring-marketing-posts-agreement-2026]]) ## Connected Concepts - [[explainable-ai]] - [[remote-proctoring]] - [[assessment-validity]] - [[formative-assessment]] - [[automated-essay-scoring]] - [[bias-mitigation]] - [[equity-in-ai-education]] - [[teacher-role]] - [[llm]] - [[higher-ed]] - [[ai-ed-evaluation]] - [[feedback]] - [[ai-feedback-quality]] - [[psychometrically-aware-ai]] - [[trust-calibration]] - [[trust]] - [[educational-measurement]] - [[item-response-theory]] - [[human-in-the-loop-ai]] - [[adaptive-learning]] - [[personalized-learning]] - [[summative-assessment]] — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams) ## Connected Articles - [[opraise-automated-marking-ai-assessment-2026]] — OpRaise report: AI marking of 761 university essays across three UK universities - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — HITL AI-assisted scoring in a large-scale national writing assessment (Curi et al. 2026) - [[llm-comparative-judgment-writing-screening-2026]] — Validity of Large Language Model Comparative Judgment for Universal Writing Screening - [[usher-faraon-who-grades-best-2026]] — Comparing ChatGPT, peer, and instructor grading across project quality levels (Usher & Faraon 2026) - [[assessing-quality-ai-generated-exams-field-2025]] — Assessing the quality of AI-generated exams: a large-scale field study - [[automated-formative-assessments-a-level-sciences]] — Automated formative assessments in A-level sciences - [[cong-confidence-asag-2026]] — Confidence-aware automatic short-answer grading - [[ai-scoring-language-bias-physics]] — AI scoring language bias in physics - [[llm-difficulty-calibration-programming-exams-2026]] — LLM-based difficulty calibration for programming exams - [[choi-anchor-aes-prompting-2025]] — Anchor-based automated essay scoring prompting - [[cotal-formative-assessment-scoring-2026]] — CoTAL formative assessment scoring - [[automated-grading-linux-bash-examinations-large-language-models]] — Automated grading of Linux Bash exams with LLMs - [[confidence-aware-student-drawing-assessment]] — Confidence-aware assessment of student-drawn figures - [[diagramir-educational-math-diagram-evaluation]] — DiagramIR: automatic evaluation of generated math diagrams - [[shap-llm-rationales-teaching-quality-assessment]] — SHAP vs LLM rationales for rubric-based teaching quality assessment - [[eeg-familiarity-automated-assessment-2026]] — Automating Learner Assessment: EEG-Based Familiarity Prediction - [[harmogen-ai-assessment-rubric-generation]] — HARMOGEN-R: AI assessment rubric generation - [[ai-assisted-instructor-supervised-grading-feedback]] — AI-assisted instructor-supervised grading and feedback - [[ai-grading-handwritten-physics-2026]] — AI grading of handwritten physics assessments (Olympiad) - [[auto-marking-short-answer-science-2026]] - [[pecuchova-automated-grading-open-ended-genai-2026]] - [[cvengros-grading-handwritten-chemistry-ai-2026]] - [[falahat-chatgpt-grading-pharmacy-exams-2026]] - [[olvet-genai-scoring-open-ended-medical-2026]] - [[jukiewicz-chatgpt-teacher-assessment-feedback-2026]] - [[student-perspectives-ai-writing-grading-2026]] — Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education - [[tripartite-feedback-framework-ai-assessment-2026]] — Tripartite framework: sorting feedback by epistemic status and the five boundary principles for AI involvement (Venetsanos 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Red-team evaluation of prompt injection hidden in student submissions against AI-mediated grading (Humble 2026) - [[li-explainable-trustworthy-llm-teacher-assessment-2025]] — Explainable-by-design LLM framework for automated teacher assessment with trust-gated inference (Li et al. 2025) - [[gpt4-handwritten-math-exam-grading-2026]] — handwritten semi-open mathematics grading with a confidence filter for human re-checking - [[llm-grade-bands-calibration-bias-2026]] — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing - [[sophie-clinical-communication-ai-assessment-2026]] — Scalable AI-based clinical communication training and automated assessment - [[ai-assisted-physics-lab-report-assessment-2026]] — AI-Assisted Assessment of Experimental Physics Laboratory Reports: Potential, Limitations, and Support for Teaching Practice - [[automated-constructive-assessment-hdr-llm-2026]] — Automating Constructive Assessment with Large Language Models: Toward Scalable and Repeated Evaluation of Practical Competence --- ## [Automated Essay Scoring](https://edtechdev.github.io/aied/concepts/automated-essay-scoring/) > **Automated Essay Scoring (AES)** — the use of AI to evaluate and score written essays, spanning traditional statistical approaches, fine-tuned language models, and increasingly accessible [[llm]]-based prompting strategies. AES [[research-methods-aied|research]] in this knowledge base covers scoring accuracy, fairness and bias, psychometric validity, and practical [[accessibility]] for educators. ## Questions to Consider - Automated Essay Scoring has moved from traditional statistical models to LLM-based prompting. A key tension on this page is between accuracy and accessibility — fine-tuned models score well but are impractical for most educators. Which would you prioritize in your own context, and why? - A major finding is that simply including exemplar essays in the prompt brings LLM-human agreement close to human-human reliability — and a cheaper model can match a more expensive one. What does this suggest about how much of AES quality is the model versus how it's prompted? - The page warns that AI scoring can systematically underestimate students from linguistically diverse backgrounds. Before reading, if you saw an AI give a lower score to a non-native speaker's essay, would you have assumed it was a 'bias problem' or just 'the score'? What would change how you respond? - A self-referential approach assesses L2 writers by comparing their writing to their own prior work rather than to native-speaker norms. How does the choice of comparison baseline change what a score means — and which students might it treat more fairly? - AES intersects with formative assessment when used for feedback rather than grading. When would a machine's feedback on an essay be genuinely useful to a developing writer, and when might it flatten the kinds of [[qualitative-research|qualitative]] feedback a human editor would give? ## Introduction Automated Essay Scoring has a long history in educational technology, from early statistical models to modern LLM-based approaches that can evaluate essays holistically without large pre-scored datasets. The key tension in AES research is between accuracy and accessibility — while fine-tuned models achieve strong results, they are resource-intensive and impractical for most educators. - **[[zhang-races-consistent-essay-scoring-llms-2026|Zhang et al.]]** RACES uses reward alignment to make LLM essay scoring both accurate and consistent, addressing a core AES validity concern. ## Key research themes **Prompting-based AES** has emerged as the most accessible approach. The **[[choi-anchor-aes-prompting-2025|Choi et al. anchor paper study]]** shows that including exemplar essays in prompts brings LLM-human agreement close to human-human reliability, with GPT-4o mini achieving comparable results to GPT-4o at lower cost. This connects to broader [[prompt-engineering]] research and makes AES feasible for [[teacher-role|teacher]] use. **Psychometric and trait-level scoring** moves beyond holistic scores. **[[psyscore-essay-scoring-zpd-feedback|PsyScore]]** provides a psychometrically-aware framework for trait-adaptive scoring with [[sociocultural-learning|ZPD]]-grounded feedback. **[[icle-plus-plus-essay-scoring|ICLE++]]** models fine-grained traits for holistic essay scoring, advancing the precision of automated evaluation. **Bias and fairness** is a critical concern. **[[ai-scoring-language-bias-physics|Feser & Tschisgale]]** found that AI scoring systematically underestimates students from linguistically diverse backgrounds, highlighting the need for [[bias-mitigation]] and [[equity-in-ai-education]] considerations in AES deployment. **L2 and self-referential assessment** explores non-native writing contexts. **[[self-referential-l2-writing-llm-assessment|Profile-based L2 assessment]]** uses a self-referential approach comparing student writing to their own prior work rather than native-speaker norms. **[[explainable-ai|Interpretability]] and feature weighting** opens the blackbox of how LLMs actually score. **[[llm-essay-scoring-feature-weighting-2026|Wang et al. (2026)]]** compared three LLMs (Qwen, GPT, Gemini) with human raters on non-native English essays across sixteen textual features, finding strong overall alignment but distinct weighting: LLMs emphasized grammatical accuracy, lexical sophistication, and syntactic complexity, while human raters prioritized content completeness and visual presentation. Critically, LLMs shifted their weighting by proficiency level — placing more weight on language errors for low-proficiency students and increasingly rewarding linguistic sophistication for high-proficiency students — whereas human raters maintained a more stable framework. **[[llm-essay-assessment-framework-reliability-2026|Liu, Ye, and Yan (2026)]]** extend this with a five-model evaluation framework (GPT-4.1, Llama 4 Maverick, Gemini 2.5 Flash, Claude Sonnet 4, DeepSeek R1) on 60 long essays, using causal discovery to reveal distinct evaluative heuristics: most models prioritized lexical precision and fluency, while others emphasized syntactic complexity or cross-domain integration, and some showed inconsistency, score compression, or systematic underestimation. Together these studies establish that AES validity depends not only on overall agreement but on *how* models weight features and whether that weighting is stable across learner subgroups — directly informing [[assessment-validity]] and [[bias-mitigation]] auditing. **Item-type boundaries and the limits of essay grading.** In a mixed-format university exam, [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat, Das, Bhaumik & Thambi (2026)]] found ChatGPT-5's concordance with faculty was substantial-to-near-perfect on objective items (CCC 0.935–1.000) but dropped sharply on open-ended responses — near-zero to negative for short-answer and only 0.341–0.854 for essay questions — and a structured rubric did not consistently improve essay agreement. This bounds AES validity: model fluency helps on well-specified items but does not carry over to holistic essay scoring, where contextual interpretation of partial-credit responses still favors [[human-in-the-loop-ai|human judgment]]. **High-stakes deployment evidence.** Field evidence from a real high-stakes deployment ([[human-in-the-loop-ai-scoring-national-assessment-2026|Uruguay's *Acredita EB*, 2024–2025; Curi et al.]]) shows prompt-engineered GPT-5 reaching 60–80% agreement with trained human raters across a 15-item analytic Spanish rubric, roughly five percentage points below human inter-rater agreement for most items and never more than 15 points below, with run-to-run consistency above 90% for almost all items. Vocabulary, syntax and spelling were the weakest dimensions, and the spelling item had to be handed to a deterministic grammar checker (LanguageTool, 68% accuracy, 100% consistency) because token-level orthography is where the model is least stable. Prompts built for one exam edition transferred to the next with only topic-specific edits, and the AI was systematically stricter than humans — under-grading rather than over-grading. For AES design, the lesson is that prompting-based scoring can approach human agreement even in a national exam, while its remaining weakness sits at the level of low-level language conventions rather than holistic [[writing-education|writing]] quality. **The human [[benchmark]], and what it can and cannot license.** [[opraise-automated-marking-ai-assessment-2026|The OpRaise study]] is the strongest test in this knowledge base of whether LLM marking is ready for routine use, and its answer turns on the benchmark rather than the model. It compared three frontier systems (Claude Opus 4.6, GPT-5.4, Gemini 3 Flash), each under 27 prompt configurations crossing rubric specificity, calibration and scoring strategy, against the moderated marks of 761 authentic Psychology essays from 125 students at three UK [[higher-ed|universities]]. Human marks were adopted as ground truth because academic judgment is the socially accepted standard, and the authors note that human markers agree only moderately with each other — which bounds how strong AI–human agreement could reasonably be demanded to be. Against that benchmark, agreement on the UK degree band ranged from **35 to 65 percent by institution** (63 percent at Cambridge, 53 percent at Nottingham, 35 percent at Manchester Metropolitan) and did not transfer between them, so the report's central recommendation is local validation on an institution's own [[assessment]] materials. Two findings generalize beyond this corpus. First, reliability and agreement pull apart: every model re-marked nearly identically (ICC 1.00, 1.00, 0.97) and the models agreed with each other more closely than with humans (three-model ICC 0.91), yet all three agreed on the degree band for only **56 percent** of submissions — self-consistency is not validity. Second, AI marks compressed toward the middle of the scale (a compression score of 0.47–0.82, with the crossover where AI and human agree on average sitting in the upper 50s to low 60s), making AI least accurate precisely at the grade boundaries that separate Firsts from Upper Seconds and passes from fails. Vocabulary range, connectives, sentence complexity and text length predicted AI marks with small but significant effects while their relationship with human marks was broadly negligible — a direct demonstration of the [[ai-scoring-language-bias-physics|linguistic-bias]] concern, and one that no prompting strategy tested removed. **AES as a training signal rather than a judge.** Scoring models are usually studied as evaluators, but SWIM uses one as a reward function: [[swim-student-writing-simulation-2026|Do, Kontak and Sachan (2026)]] freeze a multi-trait AES verifier and score it against generated student essays, turning the predicted trait profile into a dense per-essay reward (the mean trait-normalized distance from the target profile) for GRPO, deliberately not the Quadratic Weighted Kappa metric itself, because QWK is defined over a batch of target-prediction pairs and is not a per-sample signal, while exact-match rewards are too sparse in the multi-trait setting. The gains were re-checked against a DeBERTa scorer and an evaluation-only verifier the policy never trained against (0.598 versus 0.479 for SFT, and 0.647 versus 0.501), which is the check that separates genuine proficiency control from adaptation to the reward model. For AES research this is a second validity demand: a scorer used as a training target is being optimized against, so its own trait weighting and language biases propagate into everything the generator learns - the same feature-weighting asymmetries (grammatical accuracy, lexical sophistication and syntactic complexity weighted more heavily by models than by human raters) documented elsewhere on this page, now shaping training rather than only marks. ### Connections to related concepts AES sits at the intersection of [[automated-assessment]], [[writing-education]], and [[generative-ai]]. It connects to [[formative-assessment]] when used for feedback rather than grading, to [[feedback|Feedback Loop]] when integrated into iterative writing processes, and to [[ai-literacy]] when educators understand and calibrate AES tools. The [[assessment-validity]] and [[educational-measurement]] concepts are essential for ensuring AES scores are meaningful and fair. - **Agreement, error and what a hybrid scorer adds.** In a small open-ended marketing-writing corpus the LLM out-scored both deterministic rules and an equal-weight hybrid on absolute agreement with human raters (ICC(2,1) .435 versus .266 and .091), with the hybrid significantly worse than the LLM alone, while score dispersion and a single near-empty response showed how strongly such estimates depend on corpus composition ([[automated-scoring-marketing-posts-agreement-2026]]). ## Connected Concepts - [[bias-mitigation]] - [[equity-in-ai-education]] - [[automated-assessment]] - [[language-learning]] - [[educational-measurement]] - [[k-12]] - [[prompt-engineering]] - [[writing-education]] - [[ai-literacy]] - [[assessment-validity]] ## Connected Articles - [[opraise-automated-marking-ai-assessment-2026]] — OpRaise report: AI marking of 761 university essays across three UK universities - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment - [[zhang-races-consistent-essay-scoring-llms-2026]] — RACES: reward-aligned consistent essay scoring with LLMs - [[ai-scoring-language-bias-physics]] - [[choi-anchor-aes-prompting-2025]] - [[icle-plus-plus-essay-scoring]] - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans (Mathew et al. 2026) - [[psyscore-essay-scoring-zpd-feedback]] - [[self-referential-l2-writing-llm-assessment]] - [[aiawe-automated-writing-evaluation]] - [[llm-essay-scoring-feature-weighting-2026]] — Feature weighting patterns in LLM-based essay scoring (Wang et al. 2026) - [[llm-essay-assessment-framework-reliability-2026]] — Framework for evaluating LLMs in essay assessment (Liu, Ye & Yan 2026) - [[falahat-chatgpt-grading-pharmacy-exams-2026]] - [[know-when-to-trust-ai-scoring-reliability-2026]] — Know When to Trust: self-confidence, weighted probabilistic scoring and ensembling improve LLM scoring agreement - [[swim-student-writing-simulation-2026]] — a frozen AES verifier used as a dense training reward for a writing generator - [[llm-grade-bands-calibration-bias-2026]] — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing --- ## [Automated Question Generation](https://edtechdev.github.io/aied/concepts/automated-question-generation/) > **Automated question generation (AQG)** — the use of AI, especially NLP and [[llm|large language models (LLMs)]], to generate [[assessment|educational assessment]] items (multiple-choice, short-answer, fill-in-the-blank, coding, and performance questions) automatically from source material or learning objectives. AQG enables assessment at scale — producing [[formative-assessment|formative]] quizzes, adaptive exercises, and practice items — but quality varies dramatically across item types and requires validation to avoid hallucinated or poorly calibrated questions. It is a core component of [[automated-assessment]] and a key enabler of [[adaptive-learning]] and [[personalized-learning]]. ## Questions to Consider - Automated question generation produces assessment items from source material at scale — but this page warns that quality varies dramatically across item types. Before reading, which item type would you guess AI generates most reliably: multiple-choice, short-answer, or code — and why? - The central challenge is quality control: LLMs can generate factually incorrect questions. One pipeline reduced hallucination by 62% by adding a generate-then-validate-refine loop. Why do you think asking the AI to validate its own questions would meaningfully improve them rather than just rubber-stamping its own output? - [[research-methods-aied|Research]] shows generated questions may skew toward lower-order thinking (recall) unless explicitly designed for higher-order outcomes. If you used AI to build practice items, how would you know whether they were training genuine understanding or just memorization? - Difficulty calibration matters: AI difficulty estimates correlate strongly with student performance, but the page cautions against high-stakes misuse. When would a question that AI judges 'right difficulty' still be the wrong question to ask a particular learner? - [[accessibility]]-aware generation builds questions tuned for Deaf and Hard of Hearing learners, refined in partnership with the target community. What does this example suggest about why question generation can't be treated as a purely technical or content-only problem? ## Introduction Automated question generation matters because assessment items are expensive to create by hand, and AI can produce them rapidly and at scale. However, the knowledge base's research shows that generated items must be validated for correctness, relevance, and difficulty, and that different item types (MCQs, short-answer, code) vary in how reliably they can be generated. AQG therefore sits at the intersection of [[generative-ai|generative AI]], [[educational-nlp|educational NLP]], and [[educational-measurement]]. ## Approaches to question generation The knowledge base's research illustrates several approaches: - **Generate-then-validate pipelines:** [[generate-then-validate-question-gen|Generate-Then-Validate]] introduces a generation → validation → refinement loop that reduces LLM hallucination by 62% compared to direct generation, achieving 89% accuracy on [[stem-education|STEM]] datasets and a 23% improvement in relevance. The validation step filters invalid or low-quality items, and failed items trigger re-generation with corrective prompts. - **Knowledge-tracing-based generation:** [[kt4eqg-personalized-question-generation|KT4EQG]] generates personalized exercise questions guided by [[knowledge-tracing|knowledge tracing]], tailoring items to each learner's knowledge state rather than generating generic questions. - **Misconception-bound distractors as diagnostic labels:** [[colearn-agentic-tutor-co-learning-loop-2026|CoLearn (He et al., 2026)]] attaches each distractor to a single mined misconception and marks exactly one correct option, so an item is generated with its diagnostic targets built in and can then be graded deterministically with no [[llm]] call — about 7.6 seconds and roughly \$0.005 per round on the authors' live deployment, against about 32 seconds and \$0.015 for an LLM-graded short-answer round. The binding is only as good as the misconception labels behind it, since mined misconceptions matched the target misconceptions at F1 ≈ 0.56, and its item choices served a genuinely weak skill 0.72 of the time. - **Cognitive-depth-aware generation:** [[llm-educational-question-cognitive-depth|Evaluating the cognitive depth of LLM-generated questions]] examines whether generated items tap [[critical-thinking|higher-order thinking]] (creation, evaluation) or only memorization, connecting to Bloom's taxonomy and [[educational-measurement]]. - **Error-embedded case problems for repeat assessment:** [[automated-constructive-assessment-hdr-llm-2026|Takahashi et al. (2026)]] generated new hierarchical diagnostic reasoning (HDR) problems — short cases carrying deliberately embedded errors that students must find and explain — and found GPT-4o-authored items matched human-authored ones on internal consistency (Cronbach's α = 0.78 for both) and difficulty, with Fisher's exact tests finding no significant score-distribution difference across 100 participants. The format pairs a [[critical-thinking|higher-order thinking]] demand with a constrained response, so answers converge and grading becomes reproducible. The generative motivation is item reuse: reusing one case invites memory and answer reuse, while structurally equivalent but contextually different generated cases suppress item bias in re-measurement. - **[[pedagogy|Pedagogical]] pipelines:** [[slidesqaqa-pedagogical-question-generation|Slide-deck Q&A generation]] uses a multi-stage pipeline for pedagogically sound question generation from course materials. - **Accessibility-aware generation:** [[llm-question-generation-deaf-hard-of-hearing-2026|Chen et al.]] design an LLM-powered question-generation system for [[inclusive-learning|Deaf and Hard of Hearing learners]], introducing Visual and Emotion question strategies that target moments of visual or emotional difficulty in video, and iteratively refining questions with the target community to ensure linguistic accessibility. - **RAG-based, human-in-the-loop systems:** [[code-gen|CODE-GEN]] combines [[rag|retrieval-augmented generation]] with [[human-in-the-loop-ai|human-in-the-loop]] review for generating [[automated-assessment|multiple-choice assessments]]. - **[[benchmark|Benchmarks]] and evaluation:** [[nsmq-riddles-science-math-benchmark|NSMQ Riddles]] provides a benchmark of scientific/mathematical riddles for evaluating question-generation and reasoning systems. ## Validation and quality The central challenge in AQG is **quality control**: - **Assess the solving process, not the stem:** [[proiqa-math-item-quality-assessment-2026|ProIQA]] argues that item quality review should follow what an expert does — simulate the solution — and builds an LLM-generated reasoning tree per item, verified for mathematical correctness at 90.38–97.80%, then encodes its dependency structure with a graph [[machine-learning|neural network]] alongside a stem-only view. It reports average gains over the second-best method of 7.5% in concept assessment, 6.3% in difficulty estimation and 19.5% in competency assessment, and its error analysis names a failure mode worth watching: a logically correct but structurally shallow reasoning tree leaves a hard item looking easy. - **Hallucination risk:** LLMs can generate factually incorrect questions. [[generate-then-validate-question-gen|Generate-Then-Validate]] shows a dedicated validation phase sharply reduces this, and [[hallucination-risk|hallucination risk]] is a recognized concern throughout. - **Difficulty calibration:** generated questions must be calibrated to appropriate difficulty. [[llm-difficulty-calibration-programming-exams-2026|Difficulty-calibration research]] shows AI difficulty estimates correlate strongly with student performance (e.g., rho ≈ −0.87), enabling better item selection — while cautioning against high-stakes [[ai-misuse-learning-harm|misuse]]. [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] extend this to K-5 math and reading items (N = 5170) calibrated under the Rasch IRT model: GPT-4o's zero-shot difficulty ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but varied by grade, while a feature-based approach — LLM-extracted cognitive and linguistic features fed into tree-based models — reached correlations up to r = 0.87. The study's structured feature extraction (e.g., syntax complexity, [[cognitive-offloading|cognitive load]], distractor trickiness) and its practical seven-step workflow offer a template for calibrating generated items, while its early-grade range-restriction finding and generalizability caveats caution against high-stakes use. - **Large-scale psychometric field validation:** [[assessing-quality-ai-generated-exams-field-2025|Assessing AI-Generated Exams]] validates an iterative-refinement AQG pipeline (generate→judge→revise, Self-Refine style) in 91 real college classes (~1,686 students). Bayesian hierarchical 2PL [[item-response-theory|IRT]] analysis shows AI-generated questions perform on par with expert-written standardized-exam items — somewhat easier (β̄ = −0.45 vs. 0.35) but slightly more discriminating (ᾱ = 1.3 vs. 1.2), with higher peak test information (reliability 0.79 vs. 0.72) — demonstrating that AQG can produce course-tailored, psychometrically sound assessments at scale. - **Task-dependence:** generation reliability varies by item type. [[cong-confidence-asag-2026|Short-answer grading]] and [[self-referential-l2-writing-llm-assessment|analytic writing assessment]] show that open-response and writing items are harder to generate and grade reliably than structured items. - **Cognitive quality:** [[llm-educational-question-cognitive-depth|cognitive-depth evaluation]] shows generated items may skew toward lower-order thinking unless explicitly designed for higher-order outcomes. ## Role in adaptive and personalized learning AQG is a key enabler of [[adaptive-learning|adaptive]] and [[personalized-learning|personalized]] learning: it produces the large item banks that adaptive tutors draw from, and — combined with [[knowledge-tracing|knowledge tracing]] or [[student-modeling|student modeling]] — can generate items tailored to individual learners' knowledge states ([[kt4eqg-personalized-question-generation|KT4EQG]]). [[taklif-ai-interest-based-personalized-assignments|Interest-based personalization]] shows AQG can also adapt questions to student interests, not just difficulty. ## Implications for AI in education - **Generate then validate:** always pair generation with a validation/refinement stage to control hallucination and ensure relevance. - **Match item type to reliability:** use AQG for structured item types (MCQ, fill-in-the-blank, code) where it is most reliable, and apply careful validation to open-response and writing items. - **Design for cognitive depth:** prompts and pipelines should target higher-order thinking, not just recall, to support genuine learning. - **Calibrate difficulty:** use AI difficulty estimates to select appropriately challenging items, with strong validation before high-stakes use. - **Personalize via learner models:** combine AQG with knowledge tracing and interest models to generate adaptive, individualized items. ## Connected Concepts - [[llm]] - [[generative-ai]] - [[educational-nlp]] - [[automated-assessment]] - [[automated-essay-scoring]] - [[assessment]] - [[formative-assessment]] - [[adaptive-learning]] - [[personalized-learning]] - [[knowledge-tracing]] - [[student-modeling]] - [[rag]] - [[human-in-the-loop-ai]] - [[educational-measurement]] - [[item-response-theory]] - [[hallucination-risk]] - [[ai-ed-evaluation]] - [[benchmark]] - [[scaffolding]] - [[intelligent-tutoring]] - [[ai-education]] ## Connected Articles - [[assessing-quality-ai-generated-exams-field-2025]] — Large-scale field validation of AI-generated exam quality via IRT - [[generate-then-validate-question-gen]] — Generate-Then-Validate question generation - [[kt4eqg-personalized-question-generation]] — Personalized question generation via knowledge tracing - [[llm-question-generation-deaf-hard-of-hearing-2026]] — LLM-powered question generation for Deaf and Hard of Hearing learners - [[llm-educational-question-cognitive-depth]] — Cognitive depth of LLM-generated questions - [[slidesqaqa-pedagogical-question-generation]] — Slide-deck Q&A pedagogical question generation - [[code-gen]] — CODE-GEN: RAG-based human-in-the-loop question generation - [[nsmq-riddles-science-math-benchmark]] — NSMQ Riddles benchmark - [[taklif-ai-interest-based-personalized-assignments]] — Interest-based personalized assignments - [[llm-difficulty-calibration-programming-exams-2026]] — LLM-based difficulty calibration - [[self-referential-l2-writing-llm-assessment]] — Self-referential analytic writing assessment - [[cross-dataset-bloom-question-classification]] — Cross-dataset Bloom question classification - [[llm-chatbots-cs-multiple-choice]] — LLM chatbots and CS multiple-choice items - [[zerkouk-comprehensive-review-its-2025]] — Comprehensive review of intelligent tutoring systems - [[socratic-tests-conversational-assessment]] — Socratic tests: conversational assessment - [[llm-turing-test-italian-legal-exams-2026]] — LLM Turing test in legal exams - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[proiqa-math-item-quality-assessment-2026]] — ProIQA: Process-Based Math Item Quality Assessment - [[pivot-generative-video-tutors-stem-2026]] — From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning - [[colearn-agentic-tutor-co-learning-loop-2026]] — CoLearn: An Agentic Tutor that Learns its Learner in a Human-AI Co-Learning Loop - [[automated-constructive-assessment-hdr-llm-2026]] — Automating Constructive Assessment with Large Language Models: Toward Scalable and Repeated Evaluation of Practical Competence --- ## [Assessment Validity](https://edtechdev.github.io/aied/concepts/assessment-validity/) > **Assessment validity** — whether assessments measure what they claim to measure. [[ai-education]] raises fundamental validity questions: do [[automated-assessment|AI-graded]] assessments assess student learning or [[prompt-engineering|AI prompting skill]]? Does AI use invalidate traditional assessment assumptions? ## Questions to Consider - Validity asks whether an assessment measures what it claims to measure. Before reading, if you saw a student submit a polished essay you suspected was AI-assisted, would you think the bigger problem was cheating, or that the task was no longer measuring what you thought it was measuring? - This page poses a sharp question: when students use AI, does the score reflect student knowledge or AI-prompting skill? Can you think of an assessment you've designed or taken where the score might now be telling you more about the tool than about the learner? - A key finding is that the same learner input can receive semantically different replies depending on which underlying LLM is used — introducing 'construct-irrelevant variance' that threatens reliability and [[bias-mitigation|fairness]]. If two students get different AI support purely because of the model behind it, how fair is the resulting comparison? - The page argues that even when an LLM scores well, transferring human score interpretations requires similarity in the latent structure of responses — and LLMs diverge from humans here. What does this suggest about trusting an AI that 'passes' an exam designed for humans? - Rather than trying to detect AI use, the knowledge base argues for redesigning assessments so they stay valid for AI-capable students. Why might redesigning the task be a more validity-preserving strategy than policing whether AI was used? - AI now serves as test-taker, test-maker, rater, and analyst — making the interpretive chain opaque. When every role in an assessment is filled by AI, what does it even mean to say an assessment is 'valid' for the human learner in the middle of it? ## Introduction ### Validity challenges - **Construct validity:** When students use AI on assessments, does the score reflect student knowledge or AI capability? [[genai-performance-vs-learning|Performance vs. learning]] [[research-methods-aied|research]] addresses this directly. - **Cross-LLM construct-irrelevant variance in conversation-based assessment:** [[semantic-variability-llm-conversation-assessment-2026|Hao (2026)]] shows that even for a single conversational turn, the semantic content of [[llm]]-generated replies varies across models and conversational-context conditions. Within-model similarity consistently exceeds between-model similarity (0.715–0.795 vs. 0.443–0.604), and adding chat history meaningfully changes response content (median cross-history similarity ~0.40–0.45). Because the same learner input can receive semantically different replies depending on the underlying model, prompting and context alone cannot preserve response consistency — introducing potential **construct-irrelevant variance** that threatens validity, reliability, and fairness. Maintaining consistent assessment conditions as LLMs evolve is therefore an *infrastructure* challenge (symbolic rules, response templates, validation layers), not merely a [[prompt-engineering]] one. - **Evidence that is present yet never reaches the scorer:** [[ai-assisted-physics-lab-report-assessment-2026|Abreu, Stari and Martí (2026)]] name a construct-irrelevant threat that precedes any reasoning: in AI review of experimental [[physics-education|physics]] laboratory reports, an equation, graph, table or unit may be correctly included in a report yet not be retrieved from the processed content the model works from, and because a changed sign, value or unit alters the physical interpretation, a disagreement with the instructor need not indicate a reasoning error. Their remedy is procedural — high-resolution PDFs rather than photographed scans, equations written in an equation editor, legible axes and units inside figures, and instructions requiring a concrete citation for every score — because document quality sets the ceiling on whatever validity claim the resulting score supports. - **Latent-structure validity across humans and LLMs:** [[assessment-latent-structure-human-llm-2026|Strugatski et al. (2026)]] add a deeper validity condition: even when an LLM scores well, transferring human score interpretations requires similarity in the *latent structure* of responses. Comparing six [[multimodal]] LLMs to human cohorts on [[chemistry-education|chemistry]] and [[quantitative-research|quantitative]]-reasoning instruments, they find LLM–human factor structures consistently diverge (LLM–human congruence below the human–human baseline), so performance on a human-normed exam is weak evidence about LLM abilities on the constructs the items were designed to measure. - **Consequential validity:** Do AI-mediated assessments have fair consequences? [[ai-scoring-language-bias-physics|Language bias studies]] show that AI scoring can disadvantage non-native speakers. - **Rule scope is itself a validity property:** [[wright-transcription-not-generation-2026|Wright (2026)]] argues that the prohibitions written across [[higher-ed|higher education]] since 2023 bar "generative AI" without the technical precision to separate the generation of assessed content from the conversion of a format, since optical character recognition, handwritten text recognition and speech-to-text are recognition [[ai-technologies|technologies]] that infer what is already present rather than producing new content. Where a rule attaches to platform identity rather than function, it is over-inclusive by definitional accident: it captures transcription-only use of a multi-functional tool and sanctions a student who did not do the thing the rule was designed to prevent, falling hardest on [[learners]] who rely on those tools for accessibility — an inference the assessment cannot support about the construct it claims to measure. He proposes function-based drafting plus four operational criteria (fidelity, non-augmentation, traceability, attestation) for borderline cases, and treats the reliance on detector output and style heuristics as evidence of how weak the rule's own evidential basis is. - **Agentic completion removes the human-production assumption:** [[ai-agents-complete-lms-assessment-validity-2026|Hadjisolomou & El-Haddad (2026)]] extend the validity analysis from generative *assistance* to agentic *completion*: autonomous [[agentic-ai|AI agents]] can now log into an LMS, read materials, and complete unproctored asynchronous work end-to-end (demonstrated on a live course — a quiz scored 10/10 in under 5 minutes, and a fabricated-but-credible discussion-board reflection). Placing a "human-production assumption" at the base of Kane's argument-based inference chain, they show agent completion silently removes the backing for the *scoring* inference on which generalization, extrapolation, and decision inferences all rest — so every unproctored asynchronous score, including honestly earned ones, loses its interpretive support because authorship is unverifiable. Their decisive move is classifying this as a **validity failure rather than an integrity one**: an institution can punish misconduct and still lack grounds for the scores it reports. Detection is structurally insufficient (classifiers unreliable and biased; LMS monitoring sees the same clicks a student would), so the remedy is assessment redesign for verified human presence (a short oral component, process-visible drafts, class-session-specific references), with an equity-preserving menu of options rather than a proctoring mandate. - **Construct-irrelevant variance in human grading of [[generative-ai|GenAI]]-assisted work:** [[luo-dawson-value-judgments-grading-2026|Luo & Dawson (2026)]] provide a direct empirical demonstration that human grading of GenAI-assisted work is shot through with construct-irrelevant variance. In scenario-based interviews with 33 university teachers, grading decisions were driven by person-oriented (student honesty, diligence), capability-oriented (independence from AI, GenAI skill, disciplinary mastery), relation-oriented (trust built with students), and justice-oriented (fairness, beneficence) values — all of which can vary grades on factors unrelated to the outcomes being assessed. Marking down GenAI-assisted work is justified, they argue, *if and only if* the AI use prevented students from demonstrating the assessed outcomes; otherwise value-driven grading threatens validity. The study grounds the validity framing in real [[teacher-role|teacher]] practice and calls for "two-way [[explainable-ai|transparency]]" — teachers clarifying how GenAI use will affect grades, not just students declaring use. - **AI-driven grade inflation as a validity threat:**not just students declaring use. - **Variation-at-scale demands construct-equivalent variant generation:** [[varia-construct-equivalent-assessment-variant-generation-2026|VARIA (Lee 2026)]] subjects the "variation-at-scale" premise of AI-Integrated Authentic Assessment (AIAA) — replacing surveillance proctoring with per-student task variation — to an empirical, falsifiable check. Because each examinee receives a unique-but-equivalent performance task, the integrity guarantee is conditional on LLMs generating variants that are simultaneously surface-distinct, construct-equivalent, rubric-applicable, and difficulty-matched. [[benchmark|Benchmarking]] three frontier model families across four prompting strategies on 600 variants, VARIA finds frontier generators satisfy the joint integrity criteria only at the margin (joint score 0.81–0.88) while non-frontier references collapse (0.50–0.55), and no single prompting strategy dominates all four properties — so "variation-at-scale cannot be solved by prompting alone" if the diversity threshold is set aggressively, and institutions must validate their specific model-prompt pair. - **Under-grading bias in HITL AI scoring is a validity feature, not a bug:** [[human-in-the-loop-ai-scoring-national-assessment-2026|Curi et al. (2026)]] show that in a large-scale national writing assessment the systematic conservative (under-grading) bias of an LLM scorer, while problematic as a final decision, is precisely what enables safe delegation: AI-marked "passing" responses can be accepted with confidence while AI-marked "failing" responses (15.3–16.5% of cases) are routed to expert review — keeping the validity threat of AI error off the final outcome. Their [[item-response-theory|IRT]]-and-Bookmark pipeline makes the [[educational-measurement|proficiency-level]] impact of AI scoring legible rather than treating raw score agreement as the sole validity signal. - **AI-driven grade inflation as a validity threat:** [[chirikov-ai-grade-inflation-2026|Chirikov (2026)]] identifies a novel, technology-driven mechanism of [[summative-assessment|grade]] inflation operating *upstream of grading* — on the production of graded work. In a difference-in-differences study of 500,000+ grades across 319 courses (2018–2025), courses with more AI-exposed tasks (writing, coding) saw the share of A grades rise by 13 percentage points after ChatGPT's release, with grade-distribution compression. A triple-differences analysis shows the effect concentrates in homework-heavy courses — evidence that AI **task displacement** (AI performing graded tasks before instructors observe them) inflates grades without a corresponding rise in skill. This reduces the comparability of grades across courses and erodes the informational value of transcripts in ways difficult to detect from grade distributions alone. - **Program-level pass/fail susceptibility, and the marking criteria that set it:** [[ivory-psychology-assessment-integrity-2026|Ivory et al. (2026)]] judged ChatGPT output for every coursework assessment in a three-year psychology degree (40 assessments, 16 types) with a binary pass/fail, the decision markers actually make on a first read: 36 of 40 passed. Two validity implications follow. First, the operative variable is the pass boundary, not detection — marking that rewards structural fluency and the correct choice of analysis while condoning incorrect values lets fabricated statistics and hallucinated references clear a pass, so the assessment certifies tool output as student understanding because of how it is marked rather than because anyone was fooled. Second, multiple-choice items leaked their own answers: ChatGPT answered statistics questions correctly without the figure or output table shown to students, because the item stem and options implied the answer, which makes answer-option design a validity property of the item rather than a property of the AI. - **Accuracy vs. agreement as distinct validity signals:** [[falahat-chatgpt-grading-pharmacy-exams-2026|Falahat, Das, Bhaumik & Thambi (2026)]] graded a 21-item pharmacy exam with ChatGPT-5 and found that moderate percent accuracy frequently coexisted with low concordance-correlation coefficients — a methodological distinction between *scoring accuracy* and *agreement* that limits AI's reliability as a grading substitute even where raw accuracy looks acceptable. Agreement was strong on objective items (CCC 0.935–1.000) but near-zero on short-answer and modest on essay (0.341–0.854), and rubric provision did not consistently close the gap. - **Validity of AI-generated items:** [[assessing-quality-ai-generated-exams-field-2025|Assessing AI-Generated Exams]] shows that AI-generated questions, validated via Bayesian [[item-response-theory|IRT]], achieve difficulty and discrimination on par with expert-written standardized-exam items (reliability 0.79 vs. 0.72) — supporting the validity of course-tailored AI-generated assessments when backed by psychometric evaluation. - **Authentic assessment:** [[authentic-assessment]] and [[ai-assessment-scale-reform|the AI Assessment Scale]] propose validity-preserving assessment redesigns. - **Confidence and calibration:** [[automated-assessment|Confidence-aware systems]] improve validity by flagging uncertain assessments. - **[[embodied-learning|Embodied]] and multimodal evidence:** speech-only assessment can mistake verbal fluency for conceptual knowledge; [[multimodal-embodied-cognition-oral-explanations-2026|Morphew et al.]] show that computer-vision gesture analysis coupled with LLM speech analysis increases construct validity and [[equity-in-ai-education]] by capturing understanding expressed through gesture, not just words — reducing bias against learners who express understanding non-verbally. - **Homework stops certifying capability when a model can solve the task:** [[ai-particle-physics-education-redesign-2026|Mikhasenko et al. (2026)]] document a concrete invalidation of homework as a capability measure in a Bochum particle-[[physics-education|physics]] course: once a [[generative-ai|generative model]] can produce a correct solution, a submitted derivation no longer establishes that the student can solve the problem independently. Their response separates the two functions — research-shaped, AI-permitted homework kept as exploratory, bonus-bearing work, and a tools-free written examination made the sole determinant of the final grade — a validity-driven division of labor rather than a detection regime. - **Decision-oriented validity for AI scoring at scale:** [[human-in-the-loop-ai-scoring-national-assessment-2026|Curi et al. (2026)]] illustrate a decision-oriented view of validity in the Acredita EB national writing assessment: rather than asking only whether [[automated-assessment|AI scores]] agree with humans, they ask whether AI errors can change a certification decision. Automated [[item-response-theory|IRT/Bookmark]] cut scores closely reproduced the operational cut scores, and systematic AI under-grading was neutralized by routing failing AI results to [[human-in-the-loop-ai|human review]] — notably against a reference standard that is itself contested, since ten expert raters scoring the same 50 texts never reached unanimity on any rubric item. - **Student-side validity concerns about AI as grader:** [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]] adds the learner's voice to the validity question: 13 computing students whose handwritten writing task was scored by ChatGPT questioned whether an AI scorer can validly interpret intent, effort and institutional grading norms — one asking pointedly, "If ChatGPT [is] checking the exams, why are we going to university?" - **Permission does not by itself degrade the task, and use rates are the wrong validity signal:** [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al. (2026)]] surveyed 85 student teachers whose assessments explicitly permitted generative AI and found assessment engagement high and statistically unaffected by whether students used it (mean 4.21/5, SD 0.46; Mann–Whitney U = 781.5, r = 0.07, p = 0.536). With 62.4% (53) declining to use generative AI at all and adopters' use confined largely to proofreading (43.8%) and clarity checks (34.4%), the adoption rate carries little information about whether the assessment still warranted an inference about learning — and non-use driven by fear of wrongful plagiarism accusation (41.5% of non-adopters) is not evidence of validity either. - **Sampling-frame validity when a school-level test is read as system evidence:** [[el-salvador-ai-tutoring-selection-claim-2026|Restrepo Morales et al. (2026)]] give the inference from scores to claims about a *system* a quantitative treatment. A pilot assessment of 1,198 volunteer students in 171 Salvadoran schools on the PISA-based Test for Schools produced results announced as comparable to Germany and Sweden, but the instrument produces estimates for a school and is not administered under the national sampling frame with its response-rate standards, exclusion limits and weighting — so a comparison between a school-level estimate and a country mean carries at least three unreported sources of uncertainty (school-level sampling error, country-mean sampling error, and the linking error of the scale equating). Rather than test the learning claim, the authors bound it: with the top 5.5% of the Salvadoran [[math-education|mathematics]] distribution matching the German mean under zero learning, and voluntary participation understood to bias [[learning-gains|achievement]] testing upward, the published evidence cannot separate a real effect from a selected sample. This is a failure of the comparison's validity inference, not of the students or of the instrument. A conceptual proposal raises a validity boundary that applies to every AI-mediated assessment on this page: [[ai-agents-joyful-assessment-third-space-2026|El Khoury and Ma (2026)]] argue that speed is not validity, and that AI-generated prompts, examples and transcript-based reports still have to be tested for accuracy, cultural responsiveness, [[accessibility]], interpretability and alignment with course outcomes. They also widen what counts as evidence: where dialogue becomes the assessed artifact, the transcript is a process record in which judgment, empathy, clarification and shared decision-making unfold in context, and the design question becomes whether that evidence supports the inference drawn from it — the same question automated scoring faces, answered by alignment among task, feedback and evidence of learning rather than by technological novelty. ### Redesign over detection The knowledge base argues that maintaining assessment validity requires redesigning assessments for AI-capable students, not [[ai-detection|detecting AI use]]. A parallel validity problem runs through research measurement: many AI-in-education claims rest on [[self-report-measures|self-report data]], which cannot support an inference about learning or competence however well the instrument itself is validated. [[beyond-detection-authentic-assessment-ai-2025|Beyond detection approaches]] and [[assessment]] represent validity-forward thinking. **Take-home products and the [[qualitative-research|qualitative]]/quantitative split.** [[brunnstrom-ai-interaction-literacy-srl-2026|Brunnström and Palmqvist (2026)]] give the redesign argument a specific proposal grounded in the SOLO taxonomy. Working through a cognitive-science take-home exam question with a [[conversational-ai|chatbot]], they find GenAI is strongest exactly where the taxonomy is lowest — producing comprehensive, fluent, factoid-type content at the *quantitative* (multistructural) level — while the *qualitative* level (relating, evaluating, generalizing) emerged only through repeated learner-driven calibration. Their recommendation is therefore differentiated by construct: take-home assessments should emphasize evidence of qualitative understanding, while quantitative recall-based knowledge is better assessed in class where GenAI is unavailable. The demonstration also underlines that a polished submitted artifact is weak evidence in either direction, and that valid redesign must account for how demanding legitimate AI-supported learning turns out to be ([[summative-assessment]], [[academic-integrity]]). A design-level counterpart comes from the same population the redesign argument usually ignores. [[zou-is-this-a-trap-student-teachers-genai-2026|Zou et al. (2026)]] found that reflective and personalized tasks were widely judged too personal for AI help, while a knowledge-heavy assessment in another course prompted strong intent to use generative AI for references — so what students did with AI tracked the construct the task asked them to evidence. The lesson for validity is precise rather than celebratory: design can remove the payoff from generic substitution, but the same study warns against reading an opt-out cohort as proof that it did, since a substantial share of non-use was attributed to fear of wrongful plagiarism accusation rather than to the task resisting AI. [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] supplies a redesign argument on the integrity side that turns integrity evidence into validity evidence. Extending Eaton's (2023) postplagiarism framing into assessment design, he argues that detection- and verification-based integrity models are misaligned with work in which human judgment and machine generation are entangled, and reframes integrity as a [[pedagogy|pedagogical]] practice enacted through [[evaluative-judgment]] — the capacity to weigh options, justify academic choices and assume responsibility under epistemic uncertainty. The design consequence is that integrity should be evidenced rather than inferred: annotated decision trails, verification of GenAI-contributed claims, oral defense and version history are offered as integrity artifacts, and he is explicit that the difference from [[authentic-assessment]] is epistemic — authenticity asks whether a task mirrors real-world practice, integrity-oriented design asks whether learners can justify decisions and assume responsibility against disciplinary standards, so integrity becomes an assessable criterion embedded in the task architecture. He names the validity risk of his own proposal rather than leaving it implicit: requiring documented reasoning privileges learners fluent in reflective discourse and risks "the replacement of one compliance regime with another", because judgment as evidence "remains relational and situated rather than mechanically verifiable" — the same interpretive-reliability problem this page records for every AI-mediated assessment. [[weidlich-inference-at-risk-assessment-validity-2026|Weidlich (2026)]] reframes redesign as a question of which inference a change is meant to protect. Working from Kane's argument-based validity, he separates five pressures that debate tends to collapse into one: construct underrepresentation, construct-irrelevant variance, attribution of performance, conditional extrapolation, and unsupported score use. AI use is not automatically a threat, because tool use is germane where the intended construct is AI-supported professional judgment and bypasses the target performance where it is unaided reasoning, so construct and AI conditions have to be specified together. His caution for redesign follows, since interventions trade one pressure for another: oral defenses may strengthen attribution while reducing reliability or accessibility. ### Connections Assessment validity connects to [[authentic-assessment]], [[automated-assessment|Automated Grading]], [[automated-assessment|Confidence Aware AI Assessment]], [[formative-assessment]], [[academic-integrity]], and [[rct]] (which relies on valid outcome measures). AI challenges validity at the epistemic level: [[end-of-assessment-ai-disruption-transformation-2026|Hathcoat, Slotnick & Miller (2026)]] argue that when LLMs serve as test-takers, test-makers, raters, and analysts, the interpretive chain becomes opaque and the object of measurement loses definition — reframing validity as requiring AI-fluent "cyborg" judgment, and [[can-ai-evaluate-assessment-llm-meta-assessment-2026|Green et al. (2026)]] show AI scores can align with human raters (87% checklist) while the underlying rationale diverges, especially on measurement quality and weak reports. **When AI generates the assessable artifact.** [[kumar-genai-computing-education-systematic-review-2026|Kumar, Wongsirichot and Nanthaamornphong (2026)]] give the validity problem its sharpest disciplinary case: in [[cs-education|computing education]] the AI produces the artifact being graded — the source code — so tool use, learning and assessment fuse into a single interaction, and a submitted codebase no longer separates learning from delegation. Across 72 studies they find efficiency gains that do not transfer to unaided performance (21 studies), [[prior-knowledge]] moderating whether AI help becomes skill (6 studies), and detection research effectively absent from the evidence base (3 studies) while assessment redesign is comparatively well evidenced (25 studies). Their conclusion for practice is construct-specific and low-tech: add an oral component or other process-visible element to at least one high-stakes assessment per course — the single highest-leverage intervention in the review — and require critical [[student-engagement|engagement]] with AI output as a graded, observable component rather than an optional disposition ([[academic-integrity]], [[assessment]]). ## Validity under imperfect information The sharpest recent reframing treats the generative AI problem as an evidentiary one. The student knows how a piece of work was produced; the institution observes the artifact and, at best, partial traces of the process — a product–process gap that [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi (2026)]] formalize as assessment validity under imperfect information. On this account the question is not whether a rule was broken but whether the assessment still generates credible evidence of student reasoning, effort, and judgment, and each [[governance|institutional]] mechanism — prohibition, monitoring, disclosure, redesign — is evaluated by which student response it makes most attractive. The validity lens also dissolves a false separation: integrity and validity are the same problem seen from different ends, because a finding of misconduct is itself a validity claim about what the work evidences. A learner-side framework names the inference targets an artifact cannot settle by itself. [[gifted-potential-ai-assisted-work-attributional-validity-2026|Sak (2026)]] proposes *attributional validity* for judgments of gifted potential from AI-assisted work: the relation from capability to product is many-to-one, because learner competence, model capability, the learner's direction, human–AI fit and context combine differently to yield equivalent artifacts, so the product cannot identify the capacity behind it. The framework separates four targets of inference — independent competence, intellectual [[agency]], hybrid capability, and developmental carryover — and names three errors that follow when one target's evidence is read as another's: overattribution, developmental illusion, and presumed capability equivalence. Its demand on the record mirrors the evidentiary discipline this section draws from the misconduct literature: name the capability a judgment is about, document how the work was produced (model and version, interface, permitted functions, adult mediation, access conditions), and require later evidence of [[transfer-of-learning|transfer]] before reading assisted performance as a change in the learner. [[teichmann-detecting-undetectable-misconduct-2026|Teichmann (2026)]] draws the procedural consequence: where prohibited use cannot be detected, a regime that still accuses on detector scores cannot warrant the inference it draws, and so produces unfairness without effectiveness. He argues for replacing the forensic question — did the student use AI? — with a validity question — did the student demonstrate the capability the task was designed to certify? — which relocates institutional effort into program-level assessment across linked tasks, oral and supervised elements at certification points, and [[evaluative-judgment]] as an explicit object of assessment. The same section of the evidence base supplies the case-file counterpart, and it is more uncomfortable than the detection literature alone. [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] coded every generative-AI misconduct case at one regional Australian university over three years — 1,162 cases carrying 1,855 discrete evidence items — and rated each item for probative value on relevance, credibility and inferential force. The strongest categories were the ones that do not depend on probabilistic classification of text (student admissions, observed prohibited exam behavior, and independently verified fabricated references), while [[ai-detection|detector]] output drew the weakest ratings of any category: similarity or Turnitin reports were wholly weak in inferential force and standalone detector outputs wholly low in credibility, falling from 8.8% of items in 2024 to 0.5% in 2025. The validity failure is procedural as much as instrumental — no stage of the pipeline set a minimum evidentiary threshold or required investigators to weigh probative value before progressing an allegation, so evidence quality showed no reliable relationship to case outcomes, which is exactly the evidential standard issue catalogued on [[legal-issues-and-risks]]. Their proposal to apply a credentials framework at the point of *allegation* rather than only at determination is a validity demand: an allegation is a claim about what the work evidences, and it should be warranted by materials capable of supporting it. [[hadra-ai-detector-accuracy-efl-2026|Hadra, Cambridge and Mesbah (2026)]] show how far the instrument is from that standard, with macro accuracy of 0.69 (Originality) against 0.61 (Turnitin) across 192 texts, near-total failure on hybrid human–AI writing (sensitivity 0.02 and 0.31), declining accuracy with text length and on scientific writing, and a borderline tendency to misclassify EFL student writing as AI — so a detector score cannot establish the fact a misconduct finding asserts. ### The human-marking benchmark and its ceiling [[opraise-automated-marking-ai-assessment-2026|The OpRaise comparison of three frontier models against 761 authentic essays across three UK universities]] makes the benchmark part of the validity argument. Human marks were used as ground truth on the explicit ground that academic judgment is the socially accepted standard, while the authors acknowledge that human markers agree only moderately with one another — which caps the AI–human agreement that could reasonably be demanded, and means a correlation cannot be read as ready-or-not without a reference point. Within that frame the failures appeared as systematic structure rather than random error: marks compressed toward the middle of the scale, so the best and worst essays were misjudged most; agreement was weakest at grade boundaries; and AI marks tracked vocabulary range, connectives and sentence complexity while human marks were broadly insensitive to them. The practical lesson is double-edged, because the same study found reliability to be excellent — identical re-marks across time and high agreement between models. A validity case for automated marking therefore cannot rest on stability or on average agreement; it has to show the absence of systematic deviation, which is exactly what this evidence did not find. - **Agreement and benchmark error masquerade as learning.** Two 2026 studies show distinct routes by which a score can look valid while establishing something else. An open-ended marketing-writing study found LLM-human absolute agreement of only ICC(2,1) .435 and a hybrid that was significantly worse than the LLM alone, with anchor [[writing-education|composition]] moving agreement from .338 to .902 ([[automated-scoring-marketing-posts-agreement-2026]]). An expert audit of six physics benchmarks attributed 95.20% of audited rejections to defective items or graders rather than model error, moving CritPt mean@5 from 32.29% to 87.50% ([[frontier-models-physics-benchmark-audit-2026]]). In both cases the threat is construct-irrelevant variance outside the model being scored. ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[authentic-assessment]] - [[automated-assessment]] - [[formative-assessment]] - [[academic-integrity]] - [[rct]] - [[bias-mitigation]] - [[equity-in-ai-education]] - [[ai-ed-evaluation]] - [[educational-measurement]] - [[legal-issues-and-risks]] - [[llm]] - [[feedback]] - [[self-report-measures]] - [[prior-knowledge]] - [[student-engagement]] - [[assessment]] ## Connected Articles - [[ivory-psychology-assessment-integrity-2026]] — Program-level pass/fail susceptibility and the marking criteria that let fabricated values pass (Ivory et al. 2026) - [[zou-is-this-a-trap-student-teachers-genai-2026]] — Permission alone doesn't degrade validity: engagement unaffected by GenAI use; reflection resists generic substitution - [[ai-agents-joyful-assessment-third-space-2026]] — AI agents, joyful assessment, and third space - [[opraise-automated-marking-ai-assessment-2026]] — OpRaise report: AI marking of 761 university essays across three UK universities - [[kumar-genai-computing-education-systematic-review-2026]] — When AI generates the graded artifact: computing education's validity problem and redesign evidence - [[brunnstrom-ai-interaction-literacy-srl-2026]] — Take-home exams: assess the qualitative phase, move recall in-class (Brunnström & Palmqvist 2026) - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — HITL AI-assisted scoring in a large-scale national writing assessment (Curi et al. 2026) - [[varia-construct-equivalent-assessment-variant-generation-2026]] — Construct-equivalent assessment variant generation (Lee 2026) - [[chirikov-ai-grade-inflation-2026]] — AI task displacement as a mechanism of grade inflation (Chirikov 2026) - [[ai-agents-complete-lms-assessment-validity-2026]] — AI agents completing LMS tasks; human-production assumption & agentic validity (Hadjisolomou & El-Haddad 2026) - [[semantic-variability-llm-conversation-assessment-2026]] - [[assessment-latent-structure-human-llm-2026]] — Do assessment instruments measure the same thing for humans and LLMs? (Strugatski et al. 2026) - [[multimodal-embodied-cognition-oral-explanations-2026]] — A Multimodal Framework for Embodied Cognition in Oral Explanations - [[assessing-quality-ai-generated-exams-field-2025]] — Assessing the quality of AI-generated exams: a large-scale field study - [[genai-performance-vs-learning]] - [[ai-scoring-language-bias-physics]] - [[end-of-assessment-ai-disruption-transformation-2026]] - [[can-ai-evaluate-assessment-llm-meta-assessment-2026]] - [[luo-dawson-value-judgments-grading-2026]] — Value judgments in grading GenAI-assisted work: honesty, trust, validity, and two-way transparency (Luo & Dawson 2026) - [[falahat-chatgpt-grading-pharmacy-exams-2026]] - [[teichmann-detecting-undetectable-misconduct-2026]] — The misconduct procedure as a validity problem - [[mohamed-temimi-assessment-imperfect-information-disclosure-2026]] — Assessment validity under imperfect information: a response-region model - [[weidlich-inference-at-risk-assessment-validity-2026]] — Which inference is at risk? Separating five validity pressures and checking what a redesign weakens (Weidlich 2026) - [[ai-particle-physics-education-redesign-2026]] — AI in Particle Physics Education: Research Problems and Foundational Skills - [[student-perspectives-ai-writing-grading-2026]] — Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education - [[el-salvador-ai-tutoring-selection-claim-2026]] — Bounding the learning claim of El Salvador's AI tutoring pilot (Restrepo Morales et al. 2026) - [[munoz-misconduct-allegation-evidence-2026]] — What misconduct allegation files actually contain as evidence, and why evidence quality does not predict outcomes (Munoz et al. 2026) - [[hadra-ai-detector-accuracy-efl-2026]] — Detector accuracy 0.69 and 0.61 on 192 texts; hybrid-writing failure and EFL misclassification risk (Hadra et al. 2026) - [[wright-transcription-not-generation-2026]] — Over-inclusive prohibitions: transcription is not generation, so the rule sanctions what it was not designed to catch (Wright 2026) - [[sharma-judgment-visible-genai-assessment-2026]] — Integrity made visible through evaluative judgment rather than detection (Sharma 2026) - [[llm-grade-bands-calibration-bias-2026]] — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing - [[ai-assisted-physics-lab-report-assessment-2026]] — AI-Assisted Assessment of Experimental Physics Laboratory Reports: Potential, Limitations, and Support for Teaching Practice - [[gifted-potential-ai-assisted-work-attributional-validity-2026]] — Four targets of inference and three attribution errors when a product is read as evidence of learner capacity (Sak 2026) --- ## [Psychometrically Aware AI](https://edtechdev.github.io/aied/concepts/psychometrically-aware-ai/) > **Psychometrically aware AI** — AI assessment systems aligned with measurement theory — is the standard advanced in [[llm-psychometric-calibration-cdp]], [[llm-item-difficulty-prediction]], [[automated-assessment|Confidence Aware AI Assessment]], and [[item-response-theory]]: calibrated, uncertainty-aware AI assessment preserves reliability and validity rather than substituting raw model confidence for psychometric evidence. ## Questions to Consider - An AI grades a student's answer and reports a confident-sounding score. On what basis would you trust that number — and does your answer change when you learn the model wasn't calibrated against any measurement standard? - The page warns against substituting raw model confidence for psychometric evidence. Think of a time you believed a confident AI output that turned out wrong. What made its confidence unearned, and what would 'uncertainty-aware' output have looked like instead? - [[research-methods-aied|Research]] found that on the same assessment instrument, human and LLM response structures diverge — meaning a model can score well yet be measuring something different from what the exam intends. If you were a [[teacher-role|teacher]] using an AI grader, how would you ever detect that the test 'means' something different for the machine than for your students? - Item-difficulty prediction uses LLMs to estimate how hard a question is. Before reading, consider: is 'how hard is this question?' a fact about the question, or about the people (or models) answering it — and what does that ambiguity imply for using AI to calibrate exams? - Calibration, reliability, and validity are measurement concepts with precise meanings. Which of these have you actually thought through in your own assessment practice, and where might you be relying on an AI's output that has never been checked against them? - For an [[administrator]] or developer: if an AI assessment tool you're considering reports only raw accuracy, what specific questions would you now ask its vendor before deploying it with real students? ## Introduction As AI systems increasingly score responses, predict difficulty, and provide [[feedback]], a key risk is that they report confident-sounding outputs that have not been validated against measurement principles. Psychometrically aware AI addresses this by grounding AI [[assessment]] in established psychometrics — calibrating outputs, quantifying uncertainty, and preserving [[assessment-validity]] and [[educational-measurement]] standards rather than relying on raw accuracy or [[self-report-measures|self-reported]] confidence. ### How psychometrically aware AI appears in the research - **Calibration and confidence:** [[automated-assessment|Confidence-aware assessment]] and [[llm-psychometric-calibration-cdp|LLM psychometric calibration]] ensure that AI reports meaningful, uncertainty-aware scores rather than overconfident point estimates. - **Difficulty prediction:** [[llm-item-difficulty-prediction|Item-difficulty prediction]] shows how LLM-based estimates must be validated against psychometric models (see [[item-response-theory]]). [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] provide a large-scale demonstration: across 5,170 K-5 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot difficulty ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but were uneven across grades, while a feature-based approach (LLM-extracted features into tree-based models) reached correlations up to r = 0.87. The study's interpretable [[explainable-ai|feature importance]] (grade level and word count top predictors) and its practical seven-step workflow illustrate how psychometrically aware AI can be operationalized — while its early-grade range-restriction finding and generalizability caveats underscore the need to validate LLM estimates against fitted psychometric parameters. - **Measurement validity:** The concept connects to [[assessment-validity]] and [[educational-measurement]], the frameworks that define what valid, reliable AI assessment looks like. - **Latent-structure validity:** [[assessment-latent-structure-human-llm-2026|Strugatski et al. (2026)]] show that a psychometrically aware stance must also verify that an assessment measures the *same latent construct* in LLMs as in humans. Because LLM and human response factor structures diverge on the same instruments, even well-scoring models may not be measuring the construct the exam purports to measure — a caveat for any AI assessment that borrows human validity evidence. - **Latent-ability pipelines and standard setting:** [[human-in-the-loop-ai-scoring-national-assessment-2026|Curi et al. (2026)]] provide a concrete template for psychometrically aware scoring in a national exam: rubric item scores are never summed directly but fed into an [[item-response-theory|IRT]] model whose latent-ability estimates are cut with the Bookmark standard-setting method into Proficient / Close to Proficiency / Insufficient, with passing requiring at least two Proficient sections and the remaining one at least Close to Proficiency. The authors reproduced that pipeline in automated form (a 67% probability of answering at least 7 rubric items correctly for the lower cut and at least 10 for the upper), letting AI and human item scores be compared against identical decision criteria rather than on raw agreement alone. ### Connections Psychometrically aware AI sits at the intersection of [[educational-measurement]], [[assessment-validity]], [[item-response-theory]], and [[automated-assessment|Confidence Aware AI Assessment]]. It is central to [[ai-ed-evaluation]] (whether AI assessment is trustworthy) and connects to [[llm]]-based [[automated-assessment]] and [[automated-assessment|Automated Grading]]. Its emphasis on validity also speaks to the [[limitations-in-aied-research|measurement limitations]] of [[ai-education|AIED]] research. ## Connected Concepts - [[educational-measurement]] - [[assessment-validity]] - [[item-response-theory]] - [[automated-assessment]] - [[ai-ed-evaluation]] - [[llm]] - [[limitations-in-aied-research]] - [[ai-education]] ## Connected Articles - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment - [[assessment-latent-structure-human-llm-2026]] — Do assessment instruments measure the same thing for humans and LLMs? (Strugatski et al. 2026) - [[llm-psychometric-calibration-cdp]] — Aligning LLM assessment with psychometric calibration - [[llm-item-difficulty-prediction]] — LLM prediction of item difficulty - [[cong-confidence-asag-2026]] — Confidence-aware automatic short-answer grading - [[multimodal-item-parameter-estimation-2026]] — Multimodal item-parameter estimation - [[competency-based-education-genai-production-2026]] — Competency-based education with GenAI - [[end-of-assessment-ai-disruption-transformation-2026]] - [[ai-grading-handwritten-physics-2026]] — AI grading of handwritten physics assessments (Olympiad) - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[process-grounded-language-cognitive-diagnosis-2026]] — Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis - [[ai-literacy-measurement-conceptual-landscape-llm-2026]] — Comparing AI literacy instruments: jangle and jingle pairs across 55 constructs --- ## [Educational Measurement](https://edtechdev.github.io/aied/concepts/educational-measurement/) > **Educational measurement** — the psychometric theory and methods for quantifying and validating learning and its constructs — runs through the knowledge base's [[item-response-theory]], [[knowledge-tracing]], and [[assessment-validity]] pages. The [[llm]] era forces measurement to reconcile classical psychometrics with new AI-generated response streams: automated scoring, AI-predicted difficulty, and [[multimodal]] traces must be validated against established measurement principles to preserve reliability and validity. ## Questions to Consider - An LLM predicts that an exam item is easy, but empirical test data says students find it hard. Which do you [[trust]], and what would convince you to trust the machine's estimate? - Educational measurement is about turning observations of learning into defensible [[quantitative-research|quantitative]] claims. When an AI grades an essay or scores a response, is a 'score' automatically a measurement — or does something have to be validated first? What? - Some [[research-methods-aied|research]] suggests assessment instruments may not measure the same thing for humans as they do for LLMs — the latent structure diverges. If the constructs genuinely differ across humans and AI, what does that imply about AI-generated grades or difficulty ratings? - AI can score, generate items, and predict difficulty at unprecedented scale. Is 'more measurement' the same as 'better measurement'? What makes a score reliable and valid, and can those standards be preserved when the measurement is done by a generative model? - Benchmarks and AI-generated scores are everywhere now. What would you need to see before you'd treat an AI-based assessment as evidence about a learner's actual understanding rather than just a number? ## Introduction Educational measurement is the discipline of turning observations about learning — responses, behaviors, scores — into defensible quantitative claims. It encompasses construct definition, item/test design, scaling, reliability, and validity. In [[ai-education|AI in education]], measurement questions are everywhere: does a [[benchmark|benchmark score]] measure what we think? Is an AI-generated grade reliable and valid? Do AI-predicted item difficulties agree with empirically estimated ones? ## How educational measurement appears in the research - **AI's decade-long reshaping of the field:** [[xiong-ai-educational-measurement-review-2026|Xiong and Li (2026)]] map AI's impact across three eras ([[formative-assessment|Formative]] 2015–2018, Expansion 2019–2022, Generative 2023–present) via an Efficiency–Enhancement–Transformation framework, spanning AI on scoring and [[automated-question-generation|item generation]], psychometric modeling, assessment innovation and process data, and [[bias-mitigation|fairness]]/ethics/equity. They argue for a new paradigm integrating measurement theory with AI methods and for reconceptualizing constructs in the context of human–[[student-ai-interaction|AI interaction]] — the same boundary that [[assessment-latent-structure-human-llm-2026|latent-structure comparison]] probes empirically. - **AI-predicted difficulty and calibration:** [[llm-difficulty-calibration-programming-exams-2026|LLM difficulty calibration]] and [[llm-item-difficulty-prediction|item-difficulty prediction]] use LLMs to estimate item difficulty, which must be validated against psychometric estimates (see [[item-response-theory]]). [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] add a large-scale K-5 test of this premise: across 5,170 math and reading items calibrated under the Rasch IRT model, GPT-4o's zero-shot difficulty ratings correlated moderately-to-strongly with true difficulties (r = 0.83 math, r = 0.81 reading) but were uneven across grades and no better than a grade-mean dummy regressor for grades K and 1. A feature-based approach — LLM-extracted cognitive and linguistic features fed into tree-based models — reached correlations up to r = 0.87, with grade level and word count the top predictors. The study illustrates both the promise and the limits of AI-predicted difficulty as a measurement input, and offers a practical seven-step workflow for testing professionals. - **Psychometric awareness in AI [[assessment]]:** [[psychometrically-aware-ai|psychometrically aware AI]] is the standard that AI-based assessment be aligned with measurement theory — calibrated, uncertainty-aware, and validity-preserving (see [[automated-assessment|Confidence Aware AI Assessment]]). - **Automated scoring and validity:** [[ai-scoring-language-bias-physics|AI scoring and language bias]] and [[multimodal-item-parameter-estimation-2026|multimodal item-parameter estimation]] examine how automated scoring and multimodal data affect measurement quality. - **Validity frameworks:** [[assessment-validity]] and [[educational-nlp]] supply the standards and tools for validating LLM-based measurement. - **Latent-structure comparison:** [[assessment-latent-structure-human-llm-2026|Strugatski et al. (2026)]] extend educational measurement to the LLM setting by testing whether assessment instruments show the *same factor structure* for humans and LLMs. Using EFA, factor congruence, and resampling, they show LLM–human latent structures systematically diverge across [[chemistry-education|chemistry]] and quantitative-reasoning instruments, implying the constructs measured differ across populations — a necessary check before human validity evidence is assumed to transfer to AI. - **AI as a rater with measurable error:** [[cvengros-grading-handwritten-chemistry-ai-2026|Cvengros & Kortemeyer]] treat AI-assigned points as fallible observations with multiple error sources (items, runs, tasks) under generalizability theory and Kane's argument-based validity. Reliability analysis of a 296-student handwritten general-chemistry exam showed high stability of *total* scores across five AI runs (ICC(A,1) = 0.967, Kendall's W = 0.959, 95% repeatability coefficient 5.33 of 60 points) with lower item-level stability (ICC(A,1) = 0.836) — unsystematic errors partially cancel when summed (a Spearman–Brown aggregation effect that lifts total-score agreement to R² = 0.91), while a small positive intercept with slope < 1 revealed a "timid grader" score-compression bias. This frames run-averaging and aggregation as measurement design decisions, and motivates confidence filters calibrated to [[item-response-theory|IRT]] risk as a validity safeguard before automated measures are trusted. - **Calibration that scales with new data rather than the whole history:** [[bayesian-consensus-irt-item-banks-2026|Jewsbury et al. (2026)]] confront the operational consequence of AI item generation — banks larger, sparser and more frequently updated than conventional ones, where refitting all accumulated response data at each update is costly and can exceed available memory. Their *consensus calibration* combines independently calibrated quarterly periods by mapping each period's posterior draws onto a common metric with a per-draw robust Haebara link (so linking uncertainty propagates rather than being treated as a fixed transformation), then aggregating as a product of Gaussian posteriors from which each period's *estimated, hierarchical* prior is subtracted and a consensus prior reinstated. On operational Duolingo English Test data across four periods the aggregate matched a pooled single-run benchmark on posterior means (r = .998 for difficulty, .991 for log-discrimination) and standard deviations (r = .970 and .920), with residual under-dispersion of 0.91–0.98 in the consensus/pooled SD ratio across exposure tertiles — concentrated in the sparsest item tertile, where prior over-counting bites hardest. The measurement lesson is twofold: an estimated prior cannot be handled by the standard subset-prior devices of parallel Bayesian computation, and posterior *dispersion* — the uncertainty adaptive selection and scoring consume — is as much a target of the correction as the posterior mean. - **Identifying a system's dispersion from two published numbers:** [[el-salvador-ai-tutoring-selection-claim-2026|Restrepo Morales et al. (2026)]] show that an [[learning-gains|achievement]] distribution's standard deviation can be recovered from published summary statistics alone — the mean and the share of students at or above a fixed proficiency cut score — because under a distributional assumption σ = (c − μ) / z(1 − p). Applied to PISA 2025 El Salvador ([[math-education|mathematics]] mean 346, 12.28% at or above the level-2 cut of 420.07), the implied σ is 63.8 (reading 76.8, science 63.8), well below the international benchmark of 100, as expected for a distribution pressed against the floor of the scale; the three estimates were obtained from independent pairs of figures and lie within about 13 points of one another, a mild internal consistency check. The share at Level 2 is reported for every PISA system and underpins the SDG 4.1.1 indicator, so the identification needs no microdata, and the same paper demonstrates the denominator problem in effect sizes by reporting one mathematics gap three ways — 129 points, 1.29 international standard deviations, or 2.02 standard deviations of the Salvadoran distribution itself — on the argument that standardized effect sizes divided by their own sample's dispersion are not legitimately comparable across studies. ## Measurement instruments in the knowledge base A central function of educational measurement is the development, validation, and use of **instruments** — the concrete scales, tests, and coding schemes that operationalize constructs. The knowledge base's articles document a wide range of instruments for AI-in-education constructs, which can be categorized by what they measure and by their measurement approach. ### AI / GenAI literacy instruments AI literacy is the construct with the richest instrument coverage in the knowledge base. Two broad families exist: **performance-based tests** (objective, less susceptible to self-report bias) and **self-report scales** (subjective, capturing perceived competence). The trade-offs of that second family are the subject of [[self-report-measures]]: self-report reaches attitudes and perceptions cheaply, but it cannot carry a claim about competence or behavior, and the knowledge base documents a 40% overestimation gap when parallel self-report and performance measures are compared. - **Performance-based (objective) measures.** The flagship is [[jin-glat-genai-literacy-assessment|GLAT (Generative AI Literacy Assessment Test)]], a 20-item multiple-choice instrument built on a 25-concept blueprint across four dimensions (Know & Understand, Use & Apply, Evaluate & Create, [[ethics]]) and validated with CTT + 2PL IRT on 355 students (RMSEA = 0.03, CFI = 0.97, α = 0.80, ω = 0.81). Critically, GLAT scores predicted AI-assisted task performance where self-report did not — evidence that **performance-based measurement outperforms self-report** for AI literacy. Related work in [[ai-literacy-assessment-misalignment]] quantifies the gap between self-reported and performance-based AI literacy (teachers overestimate by ~40%), and [[tracing-genai-literacy-interaction-patterns]] traces actual student–AI interaction patterns rather than relying on reported use. - **Self-report scales.** [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|SAIL]] operationalizes AI literacy across three domains (AI Concepts; Application and Technical Skills; AI Digital Citizenship) and four scaffolded levels; [[ai-literacy-heptagon-2026|the AI Literacy Heptagon]] structures seven dimensions (technical, application, [[critical-thinking|critical thinking]], ethics, social impact, integration, legal/[[regulation|regulatory]]) with four Bloom-aligned proficiency levels. [[genai-skill-bypass-literacy]] maps divergent AI-literacy pathways for students vs. staff, and [[panciroli-ai-literacy-episodes-situated-learning]] grounds literacy assessment in [[situated-learning|situated learning]] episodes. Instrument development for younger learners is also advancing: the AI Literacy Self-Assessment Questionnaire (AIL-SAQ) of [[ai-literacy-self-assessment-questionnaire-primary-2025|Thianwan and Srikoon (2025)]] is a 15-item self-report scale for students in Grades 4 to 6, organized around Learning About AI, Learning About How AI Works, and Learning for Life with AI, validated with exploratory factor analysis on 335 students and confirmatory factor analysis on 579 more, with an overall Cronbach's alpha of .934. - **[[discipline-specific-aied|Domain-specific]] AI literacy.** [[teacher-education-ai-literacy-sdt-2026]] develops [[teacher-role|teacher]] AI-literacy measures within a [[self-determination-theory]] framework (382 teachers, factor-validated); [[conceptualizing-preservice-teachers-ai-readiness-2026]] measures pre-service [[teacher-ai-competency|teacher AI readiness]] via intelligent-[[tpack]]; [[ai-literacy-career-adaptability-business-2026]] assesses student AI readiness and [[career-development-and-readiness|career adaptability]] in [[business-education|business education]]; and [[llm-critical-thinking-teamwork-review]] reviews instruments for LLM-supported critical-thinking and teamwork outcomes. Extending into educator measurement, the Teachers' AI Literacy Scale (TAILS) operationalizes the ED-AI framework's six dimensions to measure AI literacy specifically within [[teacher-education|language teacher education]] (validated with factor analysis), filling a gap in assessments that target students or general users. ### Attitudes, acceptance, and motivation instruments - **Technology acceptance.** Instruments grounded in TAM/UTAUT measure perceived usefulness, ease of use, and behavioral intention to use AI. See [[technology-acceptance-model]] and its application in [[acceptance-ai-english-tools-2026]] (AI-assisted English learning tools, psychometric validation across disciplinary/proficiency groups) and [[tian-genai-learning-adoption-pathways-2026|GenAI adoption pathways]]. - **Self-efficacy and motivation.** [[self-efficacy]] instruments and motivation scales (e.g., SDT-based measures of [[agency|autonomy]]/competence/relatedness in [[teacher-education-ai-literacy-sdt-2026]]) capture the motivational antecedents and consequences of AI use. These connect to [[student-engagement]] and [[prior-knowledge]] measurement. - **Project-based learning perceptions.** Zhu and Kong (2026) develop and validate a context-grounded AI [[project-based-learning|project-based learning]] scale (AI-PBLS) measuring students' perceptions of PBL when using AI for [[problem-solving|problem solving]] in AI literacy courses (EFA and CFA on 1,027 [[k-12|secondary]] and university students, 446 complete), and their SEM application demonstrates how empowerment and ethical awareness mediate PBL-to-satisfaction relationships — a robust instrument for assessing perceived PBL experiences in AI applications. ### Assessment-quality and validity instruments - **Automated scoring and rubric instruments.** [[harmogen-ai-assessment-rubric-generation|HARMOGEN-R]] generates assessment rubrics; [[ai-assisted-instructor-supervised-grading-feedback]] evaluates AI-grading quality against Elaborated-[[feedback]] criteria; [[ai-assessment-scale-reform]] addresses how AI disrupts traditional assessment scales. - **Validity-strengthening designs.** [[roe-assessment-twins-2026|assessment twins]] pair a [[generative-ai|GenAI]]-vulnerable task with a less-vulnerable equivalent assessing the same outcomes, mapping threats across Messick's six strands of validity evidence. - **Discourse and engagement coding.** [[icap-cognitive-engagement-llm-agents]] extends the [[icap-framework|ICAP]] framework into a 7-point cognitive-engagement coding scheme, comparing human annotation (κ = 0.906–0.998) with LLM-based labeling (κ = 0.541–0.609) — a measurement-instrument study showing automated coding still trails trained humans. Automated coding is also advancing via [[prompt-engineering|context-aware prompting]], which models contextual dependencies and fuses cognitive and social abilities to code [[collaborative-learning|collaborative problem-solving]] skills from process data at superior performance over strong baselines — enabling large-scale, real-time assessment while addressing the labor-intensity of manual coding. - **Skills extraction.** [[principal-trait-analysis-human-ai-skills-2026]] derives "skills" in human–AI collaboration via principal-trait analysis — a data-driven measurement of collaboration competency. ### Measurement approach matters The knowledge base's evidence repeatedly shows that **how** a construct is measured changes the conclusions. Self-reported AI literacy diverges sharply from performance-based measures ([[ai-literacy-assessment-misalignment]]); LLM annotation of engagement diverges from trained human coding ([[icap-cognitive-engagement-llm-agents]]); and latent structures differ between humans and LLMs ([[assessment-latent-structure-human-llm-2026]]). Rigorous instrument validation — reliability, structural validity, external/predictive validity — is therefore not a formality but the foundation of trustworthy AI-in-education evidence, connecting to [[assessment-validity]] and [[psychometrically-aware-ai]]. A recent material-level study shows how much validation work sits behind a simple acceptance claim: [[age-tiered-ai-literacy-guidebooks-2026|Wang, Chuang and Wu (2026)]] used a split-half design (EFA on a development sample, CFA on a held-out half) for two age-tiered AI literacy guidebooks, reporting KMO = .820, 67.3% variance explained, acceptable student fit (CFI = .943, RMSEA = .059, SRMR = .059), composite reliability of .84 to .90, and measurement invariance across editions — then documenting where the instrument strains, with HTMT values up to .950 and a constrained playfulness-intention correlation that had to be tested for unity. Their supplementary checks for differential item functioning are a model of the honesty this page argues for: invariance held, but low-variation response patterns among younger students meant one group difference weakened once those respondents were excluded, and the paper reports that alongside the headline result rather than instead of it. - **Agreement coefficients are corpus properties, and scores are joint products.** An automated-scoring validation of 60 marketing posts reported absolute agreement ICC(2,1) of .435 for an LLM, .266 for a rules-plus-LLM hybrid and .091 for deterministic rules, with uncertainty estimated from 2,000 writer-cluster bootstrap samples and MAE of 6.28-17.22 points; adding researcher-authored anchors moved inter-rater agreement from .338 to .902, showing how strongly these estimates depend on the reference set ([[automated-scoring-marketing-posts-agreement-2026]]). An expert re-grading audit of [[physics-education|physics]] benchmarks makes the complementary point at the instrument level: 95.20% of audited rejections were benchmark or grader errors rather than model failures, so a measured score must be reported as a property of the item bank, the rubric and the grader together ([[frontier-models-physics-benchmark-audit-2026]]). **Reliability is partly recoverable from the model itself.** [[know-when-to-trust-ai-scoring-reliability-2026|Organisciak and Acar (2026)]] evaluate three cheap upgrades to LLM-based scoring on more than 20,000 responses to the Alternative Uses Test: reading the model's own token-level confidence, taking a probability-weighted mean over its top-n predicted scores instead of a single best guess, and ensembling across models. Each improved agreement with human scores (correlation rising from r = 0.781 to 0.823 and RMSE falling from 0.599 to 0.498 in the best configuration), and their diagnostics locate a specific failure of single-pass scoring: the model's expressed confidence was negatively related to its accuracy (β = −0.602). This is a measurement-approach result rather than a model result — the same model, scored differently, is a more reliable instrument — and it sits alongside the corpus-level agreement evidence above. A 2026 appraisal of 33 teacher AI literacy instruments shows where instrument development is mature and where it is not. [[assessing-teachers-ai-literacy-measurement-tools-2026|Zainal, Mohd Matore and Maat (2026)]] graded the instruments against a decision matrix adapted from COSMIN and Terwee et al. (2007) and found internal consistency the strongest domain (28 of 33 at Grade A, 84.8%) and fairness the weakest, with only five instruments (15.2%) reporting measurement invariance or differential item functioning evidence. Structural validity was strong, with 24 instruments (72.7%) at Grade A through CFA, PLS-SEM or IRT modeling, yet content validity rested mostly on qualitative review, with 21 instruments (63.6%) at Grade B for lacking quantitative expert agreement statistics. ## Issues and limitations: what measurement can miss or get wrong Educational measurement is powerful but fallible. Understanding its failure modes is essential to reading AI-in-education evidence critically — and to recognizing where an apparent learning gain or construct claim may be an artifact of measurement rather than a real effect. - **Reliability limits.** Measurement is never perfectly reliable; error variance is always present. When instruments have low internal consistency or test–retest stability, observed differences may be noise. In AI contexts, new failure modes compound this: LLM-generated responses can be scored with high machine agreement yet diverge from human scoring ([[icap-cognitive-engagement-llm-agents]]), and automated scoring can be *internally* consistent while systematically wrong — precision without validity. The [[limitations-in-aied-research|measurement limitations]] of AIEd research document unreliable instruments as a cross-cutting weakness. - **Rater agreement sets the practical ceiling.** A large-scale validation on Uruguay's [[human-in-the-loop-ai-scoring-national-assessment-2026|*Acredita EB*]] found that human rater agreement itself sets the practical ceiling for AI scoring: among ten experts independently scoring 50 texts, no rubric item reached unanimous agreement with consensus, the most divergent rater typically fell below 80% agreement while the best exceeded 90%, and Cohen's Kappa was only slight or fair for several skewed items that almost all responses satisfy. The authors therefore judge automated scores against a reference standard that is itself imperfect — ten raters, a single operational score for most responses, and several items in the conventional 70% acceptability zone — a caution for any measurement claim built on single-rater operational labels. - **Validity — measuring the wrong thing.** Validity asks whether an instrument measures the construct it claims to. Common failures include **construct under-representation** (an AI-literacy test that samples only technical knowledge, missing ethics) and **construct-irrelevant variance** (an item that rewards reading fluency rather than the target skill). [[ai-scoring-language-bias-physics|AI scoring and language bias]] shows how surface features — language, phrasing, style — can drive automated scores in ways unrelated to the intended construct. [[assessment-validity]] is the guardrail against these threats. - **The self-report gap.** Self-report measures capture *perceived* competence, not actual competence. The knowledge base repeatedly shows self-reported AI literacy diverging sharply from performance-based measures ([[ai-literacy-assessment-misalignment]], ~40% overestimation by teachers) and that self-report fails to predict real AI-assisted performance where performance tests succeed ([[jin-glat-genai-literacy-assessment|GLAT]]). Measures that rely on self-report can systematically overstate constructs and conceal true skill gaps. - **Constructs that don't transfer across populations.** [[assessment-latent-structure-human-llm-2026|Strugatski et al.]] show assessment instruments can have a *different factor structure* for humans and LLMs — meaning the same items may not measure the same latent construct across populations. Even within humans, instruments validated on one group (e.g., Western, resourced [[higher-ed]]) may not generalize to others ([[global-south]]), a concern for the generalizability of AI-in-education measures. - **What measurement can miss.** Some of the constructs most central to AI-in-education learning are the hardest to measure well — and therefore the most easily missed or distorted by instruments: - **Process and strategy.** Standard outcome measures capture the *product* of learning, not the *process*. They can miss how students actually engage with AI — [[cognitive-offloading|over-reliance]], unreflective acceptance, or critical verification. [[tracing-genai-literacy-interaction-patterns]] uses interaction traces precisely because self-report and outcome tests miss these dynamics. - **Longitudinal and durable learning.** A single post-test may show inflated performance from AI assistance while missing the erosion of durable, unassisted knowledge (the [[genai-performance-vs-learning|performance–learning gap]]). Measurement at one time point can be actively misleading about learning. - **Equity and access.** Instruments that assume uniform device/connectivity access, or that are normed on privileged samples, can miss — or systematically under-measure — the capabilities of [[equity-in-ai-education|under-resourced]] [[learners]], misattributing access gaps to ability gaps. - **[[affective-computing|Affective]] and motivational states.** Engagement, motivation, and self-efficacy are often measured by self-report, inheriting the self-report gap above; they may not capture the situated, momentary dynamics that drive learning. - **Automated measurement can be confidently wrong.** The combination of high machine confidence and opaque scoring is a distinctive AI-era risk: an LLM grader or annotator can produce highly self-consistent scores that are systematically biased, and the appearance of rigor (large N, high inter-LLM agreement) can mask invalidity. The [[ai-ed-evaluation]] and [[psychometrically-aware-ai]] frameworks are the antidote — requiring calibration, uncertainty awareness, and validity evidence before automated measures are trusted. In short, educational measurement can **miss** what it does not sample (process, durability, access, affect) and can **get wrong** what it samples poorly (self-perception, surface features, cross-population constructs). Reading AI-in-education findings therefore requires asking not just *what* was measured but *how* — and what the instrument may have failed to capture. ## Connections Educational measurement is the foundation for [[item-response-theory]], [[assessment-validity]], [[knowledge-tracing]], and [[student-modeling]]. It connects to [[learning-analytics]] (measurement of learning data), [[educational-nlp]] (measuring language), and [[psychometrically-aware-ai]] (AI aligned with measurement theory). Its validity and reliability concerns underpin [[ai-ed-evaluation]] and the [[limitations-in-aied-research|measurement limitations]] of the field. For the constructs it measures, it intersects with [[ai-literacy]], [[technology-acceptance-model]], [[self-efficacy]], [[motivation]], and [[student-engagement]]. - **Causal modeling of support interventions (2026):** a structural causal modeling protocol moves educational assessment beyond associative item-response-theory belief updating toward interventional and counterfactual reasoning (e.g., the effect of hints), with structural equations elicited from experts using purely logical information — illustrated on compulsory-school algorithmic-skills tasks ([[causal-modeling-competency-assessment-2026]]). ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[item-response-theory]] - [[assessment-validity]] - [[psychometrically-aware-ai]] - [[knowledge-tracing]] - [[student-modeling]] - [[educational-nlp]] - [[learning-analytics]] - [[ai-ed-evaluation]] - [[automated-assessment]] - [[limitations-in-aied-research]] - [[ai-literacy]] - [[technology-acceptance-model]] - [[self-efficacy]] - [[benchmark]] - [[self-report-measures]] - [[motivation]] - [[student-engagement]] ## Connected Articles - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment - [[semantic-variability-llm-conversation-assessment-2026]] - [[causal-modeling-competency-assessment-2026]] — Causal Modeling of Support Interventions for Student Competency Assessment - [[assessment-latent-structure-human-llm-2026]] — Do assessment instruments measure the same thing for humans and LLMs? (Strugatski et al. 2026) - [[jin-glat-genai-literacy-assessment]] — GLAT: IRT-validated GenAI literacy test (Jin et al. 2025) - [[cdpk-pedagogy-benchmark-llms]] — LLM pedagogical-knowledge benchmark (CDPK + SEND) - [[melo-llm-classroom-observation-teach-2026]] — LLM classroom observation reliability and accuracy (Melo et al. 2026) - [[icap-cognitive-engagement-llm-agents]] — Measuring cognitive engagement with an extended ICAP framework - [[roe-assessment-twins-2026]] — Assessment twins for strengthening assessment validity under GenAI - [[ai-literacy-heptagon-2026]] — The AI Literacy Heptagon - [[ai-literacy-assessment-misalignment]] — Self-reported vs performance AI literacy misalignment - [[teacher-education-ai-literacy-sdt-2026]] — Teacher AI literacy through self-determination theory - [[acceptance-ai-english-tools-2026]] — AI acceptance measures for English learning tools - [[llm-difficulty-calibration-programming-exams-2026]] — From evaluated models to evaluation aids - [[llm-item-difficulty-prediction]] — Cognitive evaluation of LLM item-difficulty prediction - [[multimodal-item-parameter-estimation-2026]] — Multimodal item-parameter estimation - [[ai-scoring-language-bias-physics]] — AI scoring and language bias in physics - [[hashmi-socratic-physics-chatbot-2025]] — Socratic physics chatbot - [[sc2r-counterfactual-recourse-educational-2026]] — From Student Risk Prediction to SC2R: Counterfactual Recourse - [[end-of-assessment-ai-disruption-transformation-2026]] - [[eeg-familiarity-automated-assessment-2026]] — Automating Learner Assessment: EEG-Based Familiarity Prediction - [[ai-writes-code-student-writes-model-2026]] — Model authorship: theory & measurement for learning-by-construction with GenAI - [[assessing-student-drive-framework-2025]] — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise) - [[xiong-ai-educational-measurement-review-2026]] — Decade thematic review of AI in educational measurement - [[questionnaire-teachers-genai-uses-validation-2026]] — Questionnaire on teachers' uses of generative AI (Pérez-Montesdeoca et al. 2026) - [[ai-literacy-course-satisfaction-pbl-scale-2026]] — AI-PBLS scale; empowerment and ethical awareness mediating PBL-to-satisfaction in AI literacy courses (Zhu & Kong 2026) - [[context-aware-prompting-cps-skill-identification-2026]] — Context-aware prompting for automated collaborative problem-solving skill coding - [[determinants-chatgpt-use-higher-education-2026]] — ML/SHAP determinants of future ChatGPT use in higher education - [[personalized-neural-cognitive-architecture-search-2026]] — AutoML personalized neural cognitive architecture search for learner profiles - [[language-teachers-ai-literacy-edai-2026]] — Teachers' AI Literacy Scale (TAILS) psychometric study (ED-AI framework) - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[cvengros-grading-handwritten-chemistry-ai-2026]] - [[age-tiered-ai-literacy-guidebooks-2026]] — Split-half EFA/CFA validation with invariance, HTMT and DIF checks, including an honest account of playfulness-intention construct overlap - [[proiqa-math-item-quality-assessment-2026]] — ProIQA: Process-Based Math Item Quality Assessment - [[competent-generative-ai-use-measures-review-2026]] — Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use - [[durable-skills-measurement-ai-teammates-2026]] — Toward Scalable Measurement of Durable Skills - [[bayesian-consensus-irt-item-banks-2026]] — Bayesian consensus calibration: divide-and-conquer recalibration of a continuously evolving IRT item bank (Jewsbury et al. 2026) - [[el-salvador-ai-tutoring-selection-claim-2026]] — Bounding the learning claim of El Salvador's AI tutoring pilot (Restrepo Morales et al. 2026) - [[know-when-to-trust-ai-scoring-reliability-2026]] — Know When to Trust: model self-confidence, probabilistic scoring and ensembles as scoring reliability levers - [[air-scale-motivations-ai-reading-2026]] — The AIR Scale: development and validation of an instrument for motivations to use AI while reading - [[llm-distractor-generation-student-reasoning-2026]] — distractor quality criteria and the reasoning strategies behind them - [[ai-literacy-self-assessment-questionnaire-primary-2025]] — A validated 15-item self-assessment questionnaire for upper-primary AI literacy (Thianwan & Srikoon 2025) - [[assessing-teachers-ai-literacy-measurement-tools-2026]] — Systematic appraisal of 33 teacher AI literacy instruments across COSMIN-style quality domains (Zainal et al. 2026) - [[llm-grade-bands-calibration-bias-2026]] — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing - [[mental-health-literacy-students-llms-2026]] — Mental Health Literacy Across Psychology Students and Large Language Models --- ## [Item Response Theory](https://edtechdev.github.io/aied/concepts/item-response-theory/) > **Item response theory (IRT)** — a family of psychometric models that estimate latent ability from item responses by modeling the relationship between a learner's ability and the probability of answering each item correctly. IRT models item difficulty and discrimination, enabling measurement precision and adaptive testing. In the AI era, IRT meets [[llm|LLMs]] in [[llm-item-difficulty-prediction]] and [[llm-psychometric-calibration-cdp]]: AI predicts and calibrates item difficulty, potentially improving measurement precision and feeding [[adaptive-learning]]. ## Questions to Consider - Item response theory treats ability and item difficulty as jointly estimated from response patterns, rather than treating a raw test score as the measure. How might two students with the same number correct actually differ in ability? - IRT lets you compare learners on a common scale and estimate precision per person. Why might knowing an item's difficulty and discrimination matter more than just knowing whether a student got it right? - One study used IRT person-fit statistics to distinguish human from AI-generated responses on multiple-choice tests — flagging AI responses as 'aberrant.' How could the same measurement machinery that assesses learning also police academic integrity? - [[research-methods-aied|Researchers]] use IRT to validate that AI-generated exam questions match expert-written ones in difficulty and discrimination. If an AI writes an item that 'looks' good, why is empirical calibration against fitted IRT parameters still necessary? - As AI predicts and calibrates item difficulty, what could go wrong if a model's estimate of difficulty isn't validated against real student response data? - IRT connects to adaptive testing and knowledge tracing — using your responses to choose what to ask next. How does estimating your ability from each answer enable a test to become shorter and more precise rather than just longer? ## Introduction IRT treats ability (θ) and item parameters (difficulty, discrimination, sometimes guessing) as jointly estimated from response patterns, rather than treating a raw score as the measure. This makes it possible to compare learners on a common scale, to select items adaptively, and to estimate precision per person rather than globally. ### How IRT appears in the research - **AI-predicted difficulty:** [[llm-item-difficulty-prediction|LLM item-difficulty prediction]] uses language models to estimate item difficulty, which must be validated against empirically fitted IRT parameters. - **Psychometric calibration:** [[llm-psychometric-calibration-cdp|LLM psychometric calibration]] aligns model-based assessment with IRT-based measurement so that AI-generated responses preserve measurement properties. - **Knowledge tracing and student modeling:** IRT is closely related to [[knowledge-tracing]] and [[student-modeling]] — models that track learner knowledge over time — sharing the goal of estimating unobservable learner states from observable responses. - **Bayesian hierarchical field validation:** [[assessing-quality-ai-generated-exams-field-2025|Assessing AI-Generated Exams]] uses a Bayesian hierarchical 2PL IRT model (with pre-test anchor items to place 1,686 students on a common θ scale) to show that AI-generated questions match expert-written standardized-exam items in difficulty and discrimination — a large-scale demonstration of IRT as the validation backbone for [[automated-question-generation]]. - **Separating human from GenAI responses with person-fit statistics:** [[irt-human-genai-mcq-responses|Strugatski and Alexandron (2026)]] apply person-fit statistics (PFS) within IRT to distinguish human from [[generative-ai]] responses on multiple-choice assessments. PFS flag GenAI responses as 'aberrant' responders in two authentic contexts (a [[chemistry-education|chemistry]] test and a national exam), show that different [[conversational-ai|chatbots]] produce distinct response patterns (a heterogeneous group of 'intelligences'), and reveal that newer GenAI versions become more human-like — positioning IRT as a robust framework for [[academic-integrity|integrity]] screening in high-stakes testing. - **LLM difficulty estimation against Rasch IRT parameters:** [[razavi-powers-item-difficulty-llm-2026|Razavi and Powers (2026)]] evaluate whether GPT-4o can estimate the difficulty of K-5 math and reading assessment items (N = 5170) calibrated under the Rasch IRT model. A zero-shot direct estimation approach correlated moderately-to-strongly with true Rasch difficulties (r = 0.83 math, r = 0.81 reading) but was uneven across grades and often no better than a grade-mean dummy regressor for grades K and 1, likely due to range restriction in lower-grade item difficulties. A feature-based strategy — LLM-extracted cognitive and linguistic features fed into tree-based models — outperformed direct estimation (correlations up to r = 0.87), with grade level and word count the top predictors. The study underscores that LLM difficulty estimates must be validated against empirically fitted IRT parameters, and that structured feature extraction can sharpen prediction where holistic zero-shot judgment falls short. - **Item-writing flaws as a pre-deployment screen for IRT parameters:** [[item-writing-flaws-irt-difficulty-2026|Schmucker and Moore (2026)]] test whether Item-Writing Flaw (IWF) rubrics — a domain-general, textual evaluation requiring no student data — predict empirically estimated IRT difficulty and discrimination. Across **7,126 multiple-choice questions** in [[stem-education|STEM]] (physical science, [[math-education|mathematics]], life/earth sciences), they used automated, LLM-assisted coding to show that IWF rubrics carry predictive validity for empirical IRT parameters, offering a scalable pre-deployment screen that complements or partially substitutes resource-intensive pilot testing. - **IRT-based risk filtering for selective AI grading:** [[cvengros-grading-handwritten-chemistry-ai-2026|Cvengros & Kortemeyer]] fit a two-parameter logistic IRT model to AI-graded handwritten-chemistry data and define the "risk" of accepting an AI judgment as the absolute deviation between the AI's normalized score and the IRT-expected probability of credit (Risk = |s−p|); accepting only items within a chosen tolerance of this Bayesian expectation flags "surprising" AI scores for [[human-in-the-loop-ai|human review]], turning IRT from a pure score-aggregation tool into an operational acceptance/deferral mechanism for [[automated-assessment]] — one that achieved alignment with human grading similar to simpler partial-credit thresholds but with lower human workload, though its logic is less transparent to non-technical audiences. - **Divide-and-conquer calibration for continuously evolving banks:** [[bayesian-consensus-irt-item-banks-2026|Jewsbury et al. (2026)]] treat IRT recalibration as a scaling problem rather than a fitting problem. When AI-based item generation and feature-based parameter prediction make a bank larger, sparser and continuously updated, refitting the full response history at every update grows steadily costlier; their *consensus calibration* instead calibrates each time period once and combines a new period with already-computed earlier posteriors. Two features separate it from existing IRT divide-and-conquer work: the periods do not share a latent metric, so each is linked to a reference metric by a robust Haebara criterion solved *separately for every posterior draw* (carrying linking error into the linked posteriors), and each period is its own hierarchical fit contributing an estimated prior, so the naive product of posteriors must have that prior divided out and a consensus prior reinstated — reducing to the Bayesian committee machine rule when the priors are fixed. Against a pooled benchmark on four quarterly periods of the Duolingo English Test, posterior means agreed at r = .998 (difficulty) and .991 (log-discrimination) with posterior SDs at r = .970 and .920, leaving mild under-dispersion (SD ratio 0.91–0.98) that was largest in the lowest per-period exposure tertile. It is IRT calibration re-engineered for the delivery conditions AI-generated item banks create. - **How rarely IRT anchors instrument validation:** an appraisal of teacher AI literacy instruments quantifies IRT's absence rather than its use. [[assessing-teachers-ai-literacy-measurement-tools-2026|Zainal, Mohd Matore and Maat (2026)]] graded 33 instruments against a decision matrix adapted from COSMIN and Terwee et al. (2007); structural validity was strong, with 24 (72.7%) at Grade A through CFA, PLS-SEM or IRT modeling, yet none used IRT or Rasch as its primary evidence, and only five instruments (15.2%) reported measurement invariance or differential item functioning evidence. The authors argue for IRT and performance tasks alongside self-assessment to separate validated capability from reported confidence. ### Connections IRT is a foundation of [[educational-measurement]] and [[assessment-validity]], underpins [[adaptive-learning]] (adaptive item selection) and [[student-modeling]], and connects to [[psychometrically-aware-ai]] (AI assessment aligned with measurement theory) and [[knowledge-tracing]]. It features in [[llm-difficulty-calibration-programming-exams-2026|LLM difficulty calibration]] for programming assessment. ## Connected Concepts - [[educational-measurement]] - [[assessment-validity]] - [[knowledge-tracing]] - [[student-modeling]] - [[psychometrically-aware-ai]] - [[adaptive-learning]] - [[automated-assessment]] - [[intelligent-tutoring]] ## Connected Articles - [[item-writing-flaws-irt-difficulty-2026]] — Impact of item-writing flaws on IRT difficulty and discrimination (Schmucker & Moore 2026) - [[causal-modeling-competency-assessment-2026]] — Causal Modeling of Support Interventions for Student Competency Assessment - [[assessment-latent-structure-human-llm-2026]] — Do assessment instruments measure the same thing for humans and LLMs? (Strugatski et al. 2026) - [[assessing-quality-ai-generated-exams-field-2025]] — Large-scale IRT field validation of AI-generated exams - [[jin-glat-genai-literacy-assessment]] — GLAT uses IRT/2PL validation (Jin et al. 2025) - [[llm-item-difficulty-prediction]] — LLM prediction of item difficulty - [[llm-psychometric-calibration-cdp]] — Aligning LLM assessment with psychometric calibration - [[llm-difficulty-calibration-programming-exams-2026]] — LLM difficulty calibration in programming exams - [[multimodal-item-parameter-estimation-2026]] — Multimodal item-parameter estimation - [[huang-interpretable-knowledge-tracing-2026]] — Interpretable knowledge tracing - [[zhang-ct-ai-training-test-2026]] — Computational Thinking in AI Training Test (CTAT) - [[irt-human-genai-mcq-responses]] — Using IRT to separate human and GenAI MCQ responses - [[cogevolution-student-cognitive-evolution-agent-2026]] — CogEvolution: generative agent simulating students' cognitive evolution - [[razavi-powers-item-difficulty-llm-2026]] — Estimating item difficulty using LLMs and tree-based ML - [[cvengros-grading-handwritten-chemistry-ai-2026]] - [[process-grounded-language-cognitive-diagnosis-2026]] — Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis - [[bayesian-consensus-irt-item-banks-2026]] — Bayesian consensus calibration of a continuously evolving IRT item bank (Jewsbury et al. 2026) - [[assessing-teachers-ai-literacy-measurement-tools-2026]] — Field audit showing IRT/Rasch rarely used as primary validation evidence in teacher AI literacy instruments --- ## [Self-Report Measures](https://edtechdev.github.io/aied/concepts/self-report-measures/) > **Self-report measures** — instruments in which the person being studied is also the source of the data: questionnaires and surveys, interviews, diaries, self-assessed competence, self-estimated usage, perceived learning, and satisfaction. They are the workhorse of [[ai-education|AI in education]] research — in this knowledge base 139 article pages carry a survey method tag, more than any design except experiment and [[benchmark]] — and they are the only practical way to reach attitudes, beliefs, intentions, and [[self-efficacy]]. Their limit is categorical rather than statistical: a self-report cannot measure behavior or learning, only what someone says about behavior or learning, and the two come apart in this research base often enough to be a finding in its own right. ## Questions to Consider - If a study says students "reported high engagement," what exactly was measured — and what would you need to see to conclude they actually engaged? - In one study, self-reported and objective measures of teacher [[ai-literacy|AI literacy]] correlated at only r = 0.07 to r = 0.24 across four dimensions. Where would your own [[self-assessment]] most likely diverge from a test of the same skill, and why? - Satisfaction is easy to measure and easy to improve — a system tuned to please users will score well on it. Why might that make satisfaction a poor outcome measure for learning, and what would you measure instead? - "Nearly all students use AI for coursework" comes from asking students. What could asking rather than logging get wrong in either direction? - A survey gets 112 responses at a 31% response rate, or 90 responses from 572 invitations. Before accepting its percentages, what do you want to know about who did not answer? - When a study correlates two self-reported variables measured at one moment, how many separate explanations can you generate for the correlation — and what design would rule them out? - If you were advising someone on a small evaluation, which claims would you let them support with a survey, and which would you insist need log, performance, or observational evidence? ## Introduction Self-report spans far more than the questionnaire. It includes [[quantitative-research|cross-sectional surveys]] using Likert and other rating formats, interviews and focus groups conducted by [[qualitative-research|qualitative researchers]], diaries and experience sampling, self-assessed skill or confidence, recalled estimates of how often a tool was used, and perceived learning or satisfaction treated as outcomes. What unites them is the epistemic position of the respondent: they are reporting on themselves rather than being observed. This is not a weakness to be apologized for. Many constructs of interest — motivation, anxiety, [[trust-calibration|trust]], perceived usefulness, [[student-engagement|engagement]] intentions — are internal states that no log file records, and asking is the only defensible way to reach them. The problem arises when a self-report is used to support a claim it cannot carry, which is a validity question at heart: whether the instrument measures the construct it names, and whether the claim being made is about the sort of thing self-report can see. That concern belongs with [[educational-measurement]] and [[assessment-validity]]; this page is about what happens when AI in education research leans on self-report specifically, and where the evidence base shows it breaking down. **How this page relates to its neighbors.** [[research-methods-aied]] catalogs designs and their strengths and limitations, including survey and SEM studies. [[quantitative-research]] covers the numerical methods family. [[educational-measurement]] covers instrument development, [[item-response-theory]], and psychometric validation. This page is the narrower slice those pages point at: the measurement instrument and data source itself — what self-report can and cannot establish, the recurring gap between what people report and what they do, and how to read a self-report finding critically. ## What self-report can and cannot establish Self-report is well suited to attitudes, beliefs, intentions, perceptions of a tool, affective states, and self-judged confidence or difficulty. It is poorly suited to behavior, and it cannot measure learning at all: a person's belief that they learned something is itself a perception, and the knowledge base treats perceived learning as a distinct quantity from [[learning-gains|learning gains]]. The consequence is a set of claims that look similar in a results section but differ in strength: - "Students found the tutor useful" — a perception, appropriately measured by asking. - "Students used the tutor regularly" — a behavior, only approximated by asking. - "Students learned more with the tutor" — an outcome, not establishable by asking at all. - "Teachers are confident using AI" — a self-belief, not evidence of competence. Two implications follow for anyone reading or designing this research. First, whether a self-report instrument is even measuring its named construct is an empirical question, answered by validation rather than by the plausibility of the items. Second, the direction of the temptation in AI in education is consistent: tools are evaluated by how users feel about them, and feeling is precisely the part that self-report captures most cheaply. [[self-assessment]] is a member of that wider family rather than a synonym for it: where self-report measures reach attitudes, trust, and satisfaction, self-assessment turns the learner's estimate specifically onto their own skill, confidence, or learning. A validated instrument can make that limit precise rather than vague. The AI Literacy Self-Assessment Questionnaire (AIL-SAQ) of [[ai-literacy-self-assessment-questionnaire-primary-2025|Thianwan and Srikoon (2025)]] is a 15-item scale confirmed with a stable three-factor structure across two samples (n = 335 exploratory, n = 579 confirmatory) and an overall Cronbach's alpha of .934. Its authors are explicit that it records perceived understanding, attitudes, and awareness rather than demonstrated skill, and that self-assessment accuracy depends on metacognitive ability still maturing in children, so a child's self-estimate is a weaker signal than an adult's. ## The perception–behavior gap The strongest self-report finding in this knowledge base is not that self-report is biased in general but that reported and observed behavior diverge in specific, documented ways. [[ai-literacy-assessment-misalignment|A study that built parallel self-report and objective measures of teacher AI literacy]] found **weak agreement between the self-report and objective factors (r = 0.07 to r = 0.24)** in a sample of 288 teachers, with confirmatory factor analysis supporting the construct validity of both measures while showing a **low correlation between the self-report and objective factors**. The two instruments were credible; they simply measured different things. In the full sample of 288 teachers the correlations between the objective and self-reported factors ran from r = 0.07 to r = 0.24, and latent profile analysis found 43 teachers who rated themselves high while scoring lower on the objective measure against 59 who showed the reverse pattern. [[jin-glat-genai-literacy-assessment|GLAT]] reaches the same conclusion from the other direction: its 20-item performance-based test predicted performance on [[generative-ai|GenAI]]-supported learning tasks in a within-subject study of 83 students, **while self-reported ChatGPT literacy did not**. The authors' framing is blunt — instruments in this area overwhelmingly rely on self-reported surveys, "which capture perceived rather than actual competence and are prone to bias and overestimation." Behavioral estimates show the same split. [[predicting-attrition-competitive-programming|A large-scale competitive-programming study]] reported a **disconnect between self-reported confidence and actual practice behavior**, the former a poor proxy for the latter. [[engagement-intensity-learner-modeling|A learner-modeling study]] is candid that usage frequency was self-reported on a Never-to-Daily scale rather than observed, and that with predictors and outcomes both self-reported, response styles such as acquiescence or extremity bias could produce the associations. Sometimes asking and observing are set up head to head. [[student-llm-interaction-taxonomy-review-2026|A scoping review of 46 categorizations from 33 studies]] found the literature split about evenly between self-report and interaction-log data, and concluded that categories "often reflect the measurement approach as much as the interaction itself" — with self-report studies capturing perceptions and intentions while log-based studies capture observable conversational behavior, and the two rarely integrated. [[tracing-genai-literacy-interaction-patterns|Process-data work on GenAI literacy]] makes the constructive version of the point: whether a student prompts iteratively, refines output, and manages [[hallucination-risk|hallucinations]] is observable in interaction logs and not in a questionnaire. A 2026 structured review and exploratory [[meta-analysis-systematic-review|meta-analysis]] of measures for competent [[generative-ai]] use puts a pooled number on that gap from the other direction: pooling three directly reported same-sample subjective–objective correlations (combined reported N = 2,765) gave r = .055 (Hartung–Knapp 95% CI [−.047, .156]), and adding a fourth study's cross-factor correlations reached only r = .079. All three primary effects came from a single research program, the largest contributor's reported correlation and p-value could not be reconciled, and the review concludes that self-report cannot stand in for objective performance scores — while noting that the performance instruments are themselves narrow, covering foundation knowledge rather than the oversight and reliance behaviors that matter at work ([[competent-generative-ai-use-measures-review-2026|Verí (2026)]]). The instrument literature itself can be read as evidence of the gap's scale. [[assessing-teachers-ai-literacy-measurement-tools-2026|Zainal, Mohd Matore and Maat (2026)]] appraised 33 instruments for teacher AI literacy and found that 31 (93.9%) were self-report scales of perceived confidence, only two (6.1%) tested knowledge objectively, and none used performance-based tasks between 2019 and 2025. Their conclusion is the one this page keeps reaching from other directions: self-report scores index confidence rather than capability, so they cannot stand in for competence when groups or programs are compared. Outcome measures inherit the same gap. [[pramod-agentic-ai-motivational-pathways-2026|Pramod and Patil (2026)]] model the path from [[agentic-ai|agentic AI]] through motivation and social presence to what they label learning performance with a coefficient of 0.671 — the strongest relationship in the study — and their own limitations section states that this dependent variable reflects learners' perceptions and not exam results, assignment performance or learning analytics. That is the pattern to read carefully in pathway models generally: a large coefficient on a perceived outcome quantifies how consistently students believe something helped, and says nothing yet about whether it did. ## Satisfaction and perceived learning as outcomes Satisfaction is the most frequently self-reported outcome in this corpus, appearing on 63 article pages. It is also the weakest as a proxy for learning, and the knowledge base contains explicit arguments to that effect. - [[sequenced-ai-feedback-learning|Work on sequenced AI feedback]] states the design principle directly: user satisfaction and behavioral engagement are not reliable proxies for learning gains, so learning must be measured directly. - [[puech-pedagogical-steering-llm-productive-failure-2025|On pedagogical steering]] notes that current language models are instruction-tuned to be helpful assistants that maximize user satisfaction, while a tutor's goal is to maximize learning — the two objectives can conflict, which makes satisfaction a potentially misleading target rather than merely a weak one. - [[preferred-scaffolding-ai-mathematical-modeling|A scaffolding study]] argues that perceived usefulness, ease of use, and immediate satisfaction should not be treated as sufficient indicators of [[scaffolding]] effectiveness. - [[nie-personavlm-long-term-personalization-2026|On personality-aligned student modeling]] observes that optimizing for user satisfaction is not the same as optimizing for learning outcomes, and that the two can come apart. The pattern worth carrying away: satisfaction is responsive to the wrong things when learning is the goal. A system that answers quickly, agrees readily, and reduces effort will be rated highly, and those same properties are the ones the knowledge base associates with reduced [[cognitive-offloading|productive struggle]] and inflated performance on AI-assisted work. Reported satisfaction and measured learning are not enemies; they are simply not substitutes, and treating the first as evidence of the second is the most common slippage this page documents. ## What makes a self-report instrument credible Four properties separate a survey that can support a claim from one that cannot. - **Construct validity.** Items must be shown to measure the named construct, ideally with factor analysis. The parallel-measure study above validated both its instruments and still found they diverged — validity of an instrument says nothing about whether it captures behavior. - **Reliability and structure.** Factor-validated instruments recur in this corpus, from [[teacher-education-ai-literacy-sdt-2026|teacher AI literacy measures grounded in self-determination theory]] to acceptance scales modeled with [[technology-acceptance-model|TAM]] and SEM in [[acceptance-ai-english-tools-2026|studies of AI-assisted language tools]]. - **Invariance, when groups are compared.** A scale that behaves differently across populations cannot support a comparison. [[same-ai-different-pathways|Multigroup SEM work]] reports measurement invariance testing (configural, metric, scalar) alongside its indirect effects, and [[genai-thoughtless-use-self-directed-learning-2026|a Switzerland–China comparison]] found only intrinsic value reached metric invariance and not scalar invariance — a limitation stated rather than hidden, and one that constrains what its cross-country differences can mean. - **A performance alternative where competence is the construct.** GLAT's 20 multiple-choice items exist precisely because self-report cannot carry a competence claim. ## Sampling, response, and non-response Because surveys are cheap to administer, the sampling problems are usually where the generalizability claim fails. This corpus reports them explicitly and instructively: - [[genai-student-experiences-uk-he-survey-2026|A UK survey of student GenAI use]] reports a convenience sample at 7 institutions, sharply varying [[governance|institutional]] response rates, and non-response bias that may skew toward students with strong views — and notes that ~32% of respondents described conduct that arguably broke policy, in a format where social desirability and contextual pressure make **under-reporting** the likely error direction. - [[ai-assisted-instructor-supervised-grading-feedback|A study of AI-assisted grading feedback]] rests on 112 responses, a 31% response rate. - [[ai-pedagogical-orientation|A faculty study]] draws on 90 [[stem-education|STEM]] faculty from 572 invited awardees — 16%. - [[student-genai-use-views-writing|A sociology survey]] obtained 504 respondents from 844 invitees, and the authors note response rates varied by question. - [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|A Delphi study]] completed its first round with 17 respondents, a limitation it states plainly. Self-selection compounds this. [[chatgpt-inoculation-training-verification-2026|An inoculation study]] recruited 782 participants and retained 100 valid cases from a convenience sample of self-reported ChatGPT users; [[genai-assisted-problem-posing-physics-2026|a physics study]] had students self-select into an extra-credit task, plausibly skewing views toward the agreeable and excluding skeptics. Two further self-report-specific errors recur. **Introspection limits and social desirability bias** are named as intrinsic to self-report in [[engagement-assessment-video|work on engagement assessment]] and in [[ccct-cooperative-learning-technique|studies whose qualitative data rely on self-report]]. **Common-method variance** arises when predictors and outcomes come from the same respondents at the same time: [[ai-use-critical-thinking-medical-students-2026|a moderated-mediation study of AI use and critical thinking]] operationalized usage, cognitive load, and [[self-regulated-learning|self-regulated learning]] all through standardized self-report instruments at a single time point, and raises common-method bias as a consequence; [[genai-motivation-engagement-2026|a PLS-SEM study of motivation and engagement]] likewise rests entirely on self-report and cannot establish temporal ordering among its mediators. ## Remedy: adding non-self-report evidence The knowledge base's constructive answers are consistent, and none of them requires abandoning surveys. - **Pair self-report with trace or log data.** [[mejeh-fromm-srl-adaptive-learning-feedback-2026|A 194-student, eight-week study of adaptive learning technology]] combined self-report with trace data under hierarchical linear modeling, which is what allowed it to follow self-regulated learning phases as they unfolded rather than as recalled. - **Observe practice directly when the claim is about practice.** [[hawkins-feedback-literacy-ai-essay-writing|A screen-recorded assessment task plus video-stimulated interviews]] bridged the gap between self-reported AI use and observed practice, capturing [[metacognition|metacognitive]] awareness the survey could not. - **Score behavior against a rubric rather than asking about it.** [[adaptive-pretesting-retention|Retention work]] treats practice effort as a behavioral indicator derived from interaction logs, explicitly not a measure of internal [[motivation|motivational]] state. - **Triangulate.** [[genai-over-reliance-learning-2026|Mixed-method designs]] and [[mixed-methods-research|mixed-methods research]] generally trade sample size for the ability to check one data source against another; [[ai-vocational-education-training-review|a systematic review of vocational education]] is instructive for the imbalance, noting that among the quantitative studies it examined, 2 relied exclusively on self-report while 5 used only objective measures. - **Measure the outcome you care about.** If the claim is about learning, [[learning-gains|learning gains]] must be measured, not inferred from satisfaction. ## How to read a self-report claim - Ask what kind of claim it is. Perceptions are fair game; behavior is approximated; learning and competence are not reachable this way. - Check who answered and who did not. Response rate, sampling frame, and self-selection decide what the percentages describe — a study of one institution's students describes that setting, not students in general. - Check the instrument. Was it validated for this construct and population, and if groups are compared, was invariance tested? - Watch for single-source designs. If predictors and outcomes both come from the same questionnaire at one time point, the associations are consistent with common-method variance as well as with the theory proposed. - Treat satisfaction as a design signal, not a learning outcome. High satisfaction plus untested learning is the standard configuration in AI tool evaluation, and this corpus shows the two diverging. - Prefer studies that show their working: stated response rates, stated limitations, and at least one source of evidence that does not come from asking. ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[educational-measurement]] - [[self-assessment]] - [[research-methods-aied]] - [[quantitative-research]] - [[qualitative-research]] - [[mixed-methods-research]] - [[assessment-validity]] - [[item-response-theory]] - [[ai-literacy]] - [[trust-calibration]] - [[learning-gains]] - [[student-engagement]] - [[self-efficacy]] - [[learning-analytics]] ## Connected Articles - [[pramod-agentic-ai-motivational-pathways-2026]] — Engagement predicts perceived rather than measured performance in an agentic AI path model (Pramod & Patil 2026) - [[ai-literacy-assessment-misalignment]] — Parallel self-report and objective measures of teacher AI literacy - [[jin-glat-genai-literacy-assessment]] — GLAT: a performance-based alternative to self-report - [[genai-student-experiences-uk-he-survey-2026]] — Student GenAI use surveyed across UK institutions - [[student-llm-interaction-taxonomy-review-2026]] — Taxonomies shaped by whether data came from self-report or logs - [[tracing-genai-literacy-interaction-patterns]] — Process data characterizing GenAI literacy - [[engagement-intensity-learner-modeling]] — Self-reported usage as a learner-modeling signal - [[predicting-attrition-competitive-programming]] — Self-reported confidence versus practice behavior - [[ai-use-critical-thinking-medical-students-2026]] — Common-method bias in a single-source mediation model - [[genai-motivation-engagement-2026]] — Self-report-only PLS-SEM of motivation and engagement - [[fouad-bentley-trust-utility-gap-physics-2026]] — Survey of 81 physics undergraduates on AI use and trust - [[hawkins-feedback-literacy-ai-essay-writing]] — Screen recording and stimulated recall as a self-report remedy - [[mejeh-fromm-srl-adaptive-learning-feedback-2026]] — Self-report plus trace data on self-regulated learning - [[adaptive-pretesting-retention]] — Behavioral effort indicators derived from logs - [[ai-vocational-education-training-review]] — Review reporting the self-report/objective-measure imbalance - [[student-genai-use-views-writing|Student use of and views on GenAI for writing]] — Survey plus interviews in one sociology department - [[competent-generative-ai-use-measures-review-2026]] — Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use - [[pause-ai-cognitive-offloading-self-reflection-2026]] — PAUSE: A Privacy-Preserving Self-Reflection Tool for AI-Associated Cognitive Offloading - [[air-scale-motivations-ai-reading-2026]] — The AIR Scale: a validated self-report measure of motivations for AI use in reading - [[ai-literacy-self-assessment-questionnaire-primary-2025]] — A 15-item self-assessment instrument for upper-primary students, explicit about what perceived competence can establish - [[assessing-teachers-ai-literacy-measurement-tools-2026]] — Review finding 31 of 33 teacher AI literacy instruments self-report and none performance-based --- ## [AI Detection](https://edtechdev.github.io/aied/concepts/ai-detection/) > **AI detection** — the [[ai-technologies|technologies]] and methods used to identify AI-generated content in academic submissions, and the broader question of how institutions should respond to the risk that students use large language models (LLMs) to produce work that is not their own. It spans classifier-based approaches, latent-prompt and likelihood techniques, watermarking, and stylistic analysis — and, increasingly, debates about the limits of detection and the value of redesigning assessment rather than policing it. ## Questions to Consider - If an AI detector flags a student's essay as AI-generated, how confident would you be that the flag is correct — and what evidence would you want to see before acting on it? - One argument is that AI detection is not just unreliable but conceptually unsound: a binary 'human vs. AI' ignores that student work is usually created with, not by, AI. If work is a hybrid, what does 'detecting AI' even mean? - Detection tools can be biased against non-native writers, producing false positives that unfairly penalize students. How would you weigh the risk of a [[legal-issues-and-risks|false accusation]] against the value of catching genuine misuse? - Detection can undermine integrity rather than safeguard it, fostering a climate of suspicion that erodes trust. How does being watched change how you, or a student, behave in an assessment? - Research suggests detection should be a limited, situational tool rather than a strategy of first resort, and that assessment design should recognize AI's role. What alternatives to detection might better verify what a student actually learned? - AI detectors can't be independently verified in real submissions — there's no ground truth for whether a flagged text was actually AI-generated. How comfortable are you acting on an unverifiable probability in an integrity investigation? ## Introduction AI detection sits at the intersection of [[academic-integrity]], [[generative-ai]], [[llm|large language models]], and [[assessment]]. It arose as institutions confronted students using LLMs to draft essays, code, and short answers. The field has two intertwined strands: **technical detection** (how reliably can AI-generated content be identified?) and **institutional response** (what should detection lead to, given its limits and fairness concerns?). ## Detection approaches The knowledge base's research illustrates the main technical families: - **Zero-shot likelihood / latent-prompt methods:** [[detecting-llm-generated-text-latent-prompt|EchoPrompt]] is a training-free zero-shot detector that exploits the latent prompt dependency inherent in machine-generated text. By restoring a generic assistant-response prefix and measuring likelihood-gain differences between instruction-tuned and base models, it achieves state-of-the-art detection without training, remaining robust across domain shift and paraphrasing attacks. This contrasts with purely probability-based statistical detectors that ignore the generation mechanism. - **LLM self-detection:** [[llm-detecting-llm-generated-content-education|Leinonen & Denny (2026)]] test whether LLMs can reliably detect their own generated content across programming, reflective writing, and short-answer tasks. Detection proves **highly task-dependent**: reliable for programming and longer reflective responses, but poor for short answers, where LLMs often judge their own output as *more* human-like than authentic student work. Minor prompt variations sharply reduce accuracy. - **Classifier-based and watermarking approaches:** statistical classifiers and watermarks are widely deployed in commercial tools, though their reliability is contested as LLM outputs become more sophisticated. ## The limits and risks of detection Research consistently cautions against standalone reliance on detection: - **Validity and fairness failures:** detection tools can be biased against non-native writers, producing false positives that unfairly penalize students, a concern connecting to [[bias-mitigation]] and [[equity-in-ai-education]]. - **Notable error rates and trust erosion:** unreliable detection undermines student [[trust]] and the integrity of the assessment process. - **Task-dependence:** as the self-detection study shows, accuracy varies sharply by task type, so no single detector is dependable across all assessments. - **The integrity catch-22, quantified.** [[karr-ai-detection-humanization-2026|A controlled study of 642 published abstracts (Karr et al. 2026)]] shows the [[educational-policy-ai|policy]] failure is not just conceptual but measured: guideline-compliant light AI editing is flagged at 38–80%, unmodified recent originals at 9–15% (non-STEM far above STEM), and humanizer-assisted AI text evades detection in >96% of cases. Because detectors key on surface style (long-token and academic-word density) rather than authorship intent, honest AI assistance draws sanction while deliberate humanizer evasion escapes — the authors argue detector scores should never be standalone misconduct evidence. The most recent empirical evaluations make those error rates concrete rather than generic. [[hadra-ai-detector-accuracy-efl-2026|Hadra, Cambridge and Mesbah (2026)]] ran Turnitin and Originality over a balanced 192-text corpus of genuine EFL coursework, professional writing, AI output, and 50/50 hybrids: macro accuracy reached only 0.69 and 0.61, both fell below a macro F1 of 0.55, and both were effectively useless on the hybrid texts (Originality's sensitivity 0.02), with accuracy dropping significantly as texts lengthened and again on scientific writing, plus a borderline-significant tendency to misclassify legitimate EFL student work. [[van-vlasselaer-ai-detector-reliability-2026|Van Vlasselaer, Van Droogenbroeck and Spruyt (2026)]] tested four commercial tools against a ground-truth-controlled corpus of 160 master's theses: three of them (Turnitin, GPTZero, Copyleaks) failed almost completely on fully AI-generated papers, while only Pangram performed convincingly — and when applied to 1,163 genuinely submitted theses it flagged 45.5% of them, a figure the authors insist is not a prevalence rate because live submissions have no ground truth. Both studies land on the same procedural conclusion: a detector score can prompt closer review, but it is not a finding. The calibration question also cuts the other way: because GPTZero's high-confidence label carried applicant-level false-positive rates of 0.7%, 0.5% and 1.4% across 2020–2022, [[ai-written-admissions-essays-penalized-2026|Isley, Gaebler and Goel (2026)]] read its flags as a conservative prevalence series across six admissions cycles in one US public policy master's program — the share of applicants submitting at least one flagged essay rose from 21.8% in 2023 to 45.1% in 2024 and 56.1% in 2025 against a signed prohibition, reaching 69.3% of international applicants against 38.6% of domestic ones. Human readers were a weaker instrument again: five admissions officers separating 50 human from 50 AI essays reached AUC 0.70 (95% CI [0.65, 0.75]), above chance but well below the commercial detectors. Even a well-performing interpretable detector keeps its errors at the level of documents. [[detecting-gpt-assisted-writing-stylometric-2026|Kumar, Siddiqui and Fuchsberger (2026)]] trained window-level classifiers on nine interpretable stylometric features — type-token ratio, hapax ratio, word entropy, non-stopword ratio, part-of-speech ratios and sentence-length variability — with 90 participants writing the same prompts independently and then by paraphrasing GPT output, keeping every window from a participant in one fold and testing on 18 unseen writers. Random Forest reached an ROC-AUC of 0.870 and an F1 of 0.842 on 36 held-out documents, yet flagged 4 of 18 independently authored documents as GPT-assisted, a false positive rate of 22.2% (95% CI 9.0–45.2%); [[explainable-ai|SHAP]] attribution put the most weight on hapax ratio. The authors frame the model as decision support that should prompt contextual review rather than as an automatic misconduct screen, and the accuracy it does reach does not lower the stake of its errors — each false positive is an independently authored document accused of GPT assistance, a validity problem rather than a tuning problem. The definitional and procedural problems sit alongside the statistical ones. [[wright-transcription-not-generation-2026|Wright (2026)]] argues that blanket prohibitions on "AI use" are drafted around platform identity rather than function, so they capture non-generative format conversion — speech-to-text transcription, OCR, plain text to LATEX — along with the generative drafting they mean to bar; because detectors read low-perplexity writing as machine authorship, the resulting false positives fall hardest on disabled and [[equity-in-ai-education|equity]]-exposed students. [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] reaches the design-side version of the same conclusion, positioning detection as a supplementary layer of integrity infrastructure at most, since it asks whether GenAI was used rather than how decisions were made. Detection research also carries a validity argument that outlasts questions of accuracy. Weidlich (2026) treats detector output as a conditional, probabilistic signal that may prompt further inquiry but cannot by itself establish misconduct or competence, which makes detection-centered governance an insufficient basis for upholding [[assessment-validity|assessment validity]]. Classification performance varies systematically across tools, task types, disciplines, model versions and human-AI editing practices, with formulaic STEM writing especially susceptible to algorithmic bias. Attempts to restore assessment security through detection then risk introducing construct-irrelevant variance, threatening fairness and the interpretation of scores. ## Why not to use (or try to use) AI detectors [[bassett-ai-detectors-education-2026|Bassett et al. (2026)]] argue that generative AI detection should **not be used in education at all**, on grounds that go beyond "be careful" to "this is conceptually unsound." Their case consolidates the reasons against relying on AI detectors: 1. **Unverifiable probabilistic estimates.** AI detectors output a probability that text was AI-generated, based on linguistic markers (perplexity, burstiness). Unlike other probabilistic tools (spam filters, medical diagnostics), their results **cannot be independently verified**: in real-world conditions, no ground truth exists for whether a flagged text was actually AI-generated, so validation reduces to circular reasoning. Signal-detection metrics (false-positive/negative rates) only apply in controlled tests, not real submissions. 2. **Questionable training and test data.** Detectors are trained and validated on pre-generative-AI human writing (e.g., Turnitin tested on 700,000 pre-2019 papers). The assumption that such text reflects contemporary student writing — which students now produce having been shaped by AI — is unverified, and performance shifts with model, prompt, and platform. 3. **Mutually-exclusive-linguistic-markers is a flawed assumption.** There is no principled reason a human cannot write with the linguistic features attributed to AI (or an AI with human ones), so the marker foundation itself is shaky. 4. **The false dichotomy.** Classifying text as human- vs AI-generated ignores the reality that students' work is frequently created *with*, not *by*, AI — a hybrid continuum. The binary is not merely inadequate but meaningless, making detection conceptually flawed from the outset. 5. **Procedural unfairness and evidential insufficiency.** Academic-integrity investigations must meet the balance-of-probabilities standard; AI-detector scores — alone or combined with linguistic markers, style comparisons, LLM claims, or student silence — do not reach it. Students under investigation also retain a right to silence, which detection-driven processes erode. 6. **Security and privacy risks.** Detectors store student work on servers (sometimes overseas with weaker [[privacy]] protections), creating breach, misuse, and commercial-exploitation risks. 7. **Detection undermines integrity rather than safeguarding it.** Reliance on detectors and surveillance fosters a climate of suspicion, eroding student [[trust]] and the integrity of assessment itself. Bassett et al. conclude that AI detection is an unworkable solution to a problem that cannot be solved through surveillance and punishment: the focus must move to [[assessment|assessment design]] that recognizes AI's role in learning and the reality that unsupervised assessments cannot be secured. This consolidates the knowledge base's [[beyond-detection-authentic-assessment-ai-2025|beyond-detection]] stance with a direct, evidence-based argument for retiring detection tools. ### Detector bias and the mechanism of the arms race Detection is not merely imprecise; its errors are patterned. [[teichmann-detecting-undetectable-misconduct-2026|Teichmann (2026)]] assembles the accumulated case against treating a detector score as evidence: no tool in the most comprehensive early [[benchmark]] reached 80% accuracy; simple paraphrasing or "humanising" roughly halves even that; false positives at realistic base rates exceed true ones; and non-native English speakers are misclassified systematically because the features detectors treat as signals of AI also characterize competent second-language writing. Judgment by humans does not fill the gap — expert and novice markers alike fail to distinguish AI from student prose and are confident when wrong — and the asymmetry of error means the careless-but-honest are caught while the deliberately dishonest evade, since detectors are also opaque (no thresholds, no training data, no independent replication) and therefore cannot be answered or cross-examined in a hearing. Two further points sharpen the practical stakes. First, the underlying statistical signal shrinks as models are optimized toward human prose, so the arms race is one an institution cannot win. Second, the empirical limit case is stark: in a covert field study, 94% of wholly AI-generated submissions passed unnoticed through a live online [[summative-assessment|examination]] system across five psychology modules, and the AI work on average outscored real students. [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi (2026)]] add the counter-intuitive corollary that determines when monitoring pays at all: because sensitivity raises false positives along with true positives, the rational deterrent is discrimination — the gap between flagging hidden use and flagging legitimate work. Where extra sensitivity creates more new false positives than new true positives, more monitoring makes concealment relatively more attractive, penalizing honest students faster than it identifies hidden users. The design consequence is that detectors, rules, and disclosure procedures should not be built separately. ## Beyond detection: assessment redesign A key theme in the knowledge base is that detection should be a **limited, situational tool — not a strategy of first resort**. [[beyond-detection-authentic-assessment-ai-2025|Kickbusch et al. (2025)]] argue that surveillance and detection **misdiagnose the problem**: in an AI-mediated world, authenticity cannot be policed into existence; it must be redesigned. They reconceptualize authenticity as constructed where AI is expected, declared, and scrutinized, and offer discipline-agnostic design-for-learning patterns that position AI as a collaborator rather than a cheating application. This connects detection to [[authentic-assessment]], [[assessment-validity]], [[responsible-assessment-ai-era-stanford-2026|responsible assessment]], and [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene|coauthorship integrity]]. The constructive question shifts from "how do we prevent students from using AI?" to "how do we enable them to use it thoughtfully, responsibly, and effectively in contexts that mirror their future work?" Detection therefore connects to [[ai-literacy]] (helping students [[reducing-ai-misuse|use AI responsibly]]), [[cognitive-offloading|Over-Reliance]] (understanding when AI use undermines learning), and the broader goal of supporting genuine learning rather than policing submissions. It also links to student-side phenomena such as [[student-rationalization-ai-writing|student rationalization of AI writing]] and the identity-detection challenge in [[socially-fluent-ai-identity-detection]]. ## Implications for AI in education - **Detection is situational:** institutions should use detection tools sparingly and with awareness of their error rates, fairness limits, and task-dependence — not as an automatic, standalone gate. - **Assessment design matters more than policing:** investing in [[authentic-assessment|authentic]] and process-based assessment, where AI use is expected and declared, addresses integrity more effectively than detection alone. - **Fairness and equity:** detection tools that penalize non-native writers or produce false positives risk amplifying existing inequities. - **AI literacy is complementary:** helping students understand appropriate versus [[ai-misuse-learning-harm|harmful AI use]] is more productive than relying on surveillance. - **Detection reliability caution.** A [[meta-analysis-systematic-review|systematic review]] of AI and academic integrity concludes that plagiarism/AI-detection tools cannot be relied upon for AI-generated work and should be paired with multiple assessment methods and manual review — reinforcing that detection is a limited, situational tool.([[ssaho-ai-academic-integrity-review-2025]]) - **Beyond detection: dialog over surveillance.** A practitioner account of Grand Canyon University's learning-verification framework ([[best-response-student-ai-dialog-2026|Mandernach 2026]]) argues the best response to [[student-ai-interaction|student AI use]] is dialog, not detection. Because detectors are unreliable (and biased against nonnative writers), GCU stopped asking "did the student use AI?" and instead asks students to demonstrate understanding in a brief conversation — an extension of [[authentic-assessment|assessment redesign]] that treats detection as a dead end and verification as good teaching. ## Connected Concepts - [[academic-integrity]] - [[llm]] - [[generative-ai]] - [[assessment]] - [[assessment-validity]] - [[authentic-assessment]] - [[ai-literacy]] - [[cognitive-offloading]] - [[equity-in-ai-education]] - [[bias-mitigation]] - [[higher-ed]] - [[ai-education]] - [[legal-issues-and-risks]] ## Connected Articles - [[evaluation-age-ai-output-evidence-2026]] — Evaluation in the Age of AI - [[best-response-student-ai-dialog-2026]] - [[ai-tools-academic-work-cheating-2026]] - [[detecting-llm-generated-text-latent-prompt]] — EchoPrompt: Latent Prompt Restoration Detector - [[ivory-psychology-assessment-integrity-2026]] — Detection is the wrong lever: the pass boundary decided whether AI work was graded as achievement (Ivory et al. 2026) - [[llm-detecting-llm-generated-content-education]] — Evaluating LLMs for Detecting LLM-Generated Content - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond Detection: Authentic Assessment - [[responsible-assessment-ai-era-stanford-2026]] — Responsible Assessment in the AI Era - [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene]] — Coauthorship Integrity and Assessment Validity - [[student-rationalization-ai-writing]] — Student Rationalization of AI Writing - [[socially-fluent-ai-identity-detection]] — Socially Fluent AI Identity Detection - [[ssaho-ai-academic-integrity-review-2025]] — Review of AI-based plagiarism/AI-content detection reliability - [[bassett-ai-detectors-education-2026]] — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026) - [[teichmann-detecting-undetectable-misconduct-2026]] — Why detector output cannot ground a misconduct finding - [[mohamed-temimi-assessment-imperfect-information-disclosure-2026]] — Discrimination rather than catch rate, and when monitoring backfires - [[hadra-ai-detector-accuracy-efl-2026]] — Turnitin and Originality on 192 texts: both below a macro F1 of 0.55 and near-useless on hybrid writing - [[van-vlasselaer-ai-detector-reliability-2026]] — Four detectors against 160 ground-truth papers; only Pangram performed, yet flagged 45.5% of live theses - [[munoz-misconduct-allegation-evidence-2026]] — Detector output is the weakest-rated evidence type in 1,162 real misconduct case files - [[wright-transcription-not-generation-2026]] — Blanket "AI use" rules conflate transcription with generation - [[sharma-judgment-visible-genai-assessment-2026]] — Detection demoted to a supplementary layer behind visible judgment - [[weidlich-inference-at-risk-assessment-validity-2026]] — Detection as a conditional signal, and why security responses add construct-irrelevant variance (Weidlich 2026) - [[ai-written-admissions-essays-penalized-2026]] — AI-written admissions essays are widespread but penalized - [[detecting-gpt-assisted-writing-stylometric-2026]] — Nine interpretable stylometric features: ROC-AUC 0.870 but four of 18 independently authored documents flagged (Kumar et al. 2026) --- ## [Remote Proctoring](https://edtechdev.github.io/aied/concepts/remote-proctoring/) > **Remote proctoring** — the monitoring of examinations when students take them away from a supervised physical venue, ranging from a human proctor watching via webcam or control center to fully automated AI-based proctoring systems (AIPS) that use machine/deep learning to verify identity and flag suspicious behavior. It is the primary means of preserving [[summative-assessment]] validity and [[academic-integrity|academic integrity]] in [[online-teaching-and-learning|online and distance learning]], where in-person invigilation is often unfeasible — but it raises serious concerns about [[privacy]], academic surveillance, equity, fairness, and the erosion of [[trust]], costs that can themselves harm the learning environment it is meant to protect. ## Questions to Consider - Before you read, when you picture 'academic integrity in online exams,' what's your first instinct about how to preserve it — and is that instinct about detecting cheaters or about building trust? The page argues these reflect two very different philosophies. - Remote proctoring ranges from a human watching via webcam to AI systems analyzing eye movements, head posture, and facial expressions. What does an AI system *actually* observe when it flags 'suspicious behavior,' and how confident are you that looking away from the screen equals cheating? - The page warns that monitoring can be counterproductive: being watched raises [[anxiety-and-stress|test anxiety]], and stressed students may be *more* likely to cheat. Can you think of a time pressure or surveillance affected your own performance? Does that experience undermine the case for proctoring? - Automated proctoring captures students' living space, face, and voice, often with little genuine choice but to consent. Where is the line between reasonable exam supervision and surveillance that presumes students guilty until proven honest — and who should draw it? - Research shows proctoring can flag benign behavior as suspicious, producing false accusations, and that model accuracy varies across demographics and environments. If you're an [[administrator]], how do you weigh the integrity it recovers against the equity and trust it can erode? - The page frames proctoring as one tool, not a solution — alternatives like oral and process-based assessment exist. Before reading, which approach would you defend for high-stakes assessment in a remote context, and what evidence would change your mind? ## Introduction Remote proctoring exists on a spectrum. **Online proctoring** typically involves a human proctor monitoring a student through a webcam or from a control center. **Automated/AI-based proctoring (AIPS)** replaces or augments the human with [[machine-learning]] and deep-learning systems (CNNs, RNNs, LSTMs) that analyze visual cues — eye movements, head posture, facial expressions, and body language — to detect suspicious behavior in real time. Common platforms include ProctorU and Kryterion. AIPS typically combine four functions: (1) identity authentication (e.g., camera face verification), (2) browsing restrictions, (3) remote authorization/control of the exam, and (4) report generation from recorded sessions. ## Advantages and opportunities - **Recovers validity and integrity in online assessment.** Remote proctoring answers a genuine validity problem. In [[online-teaching-and-learning|online assessment]], [[generative-ai|generative AI]] makes unproctored work unreliable as a measure of learning: unassisted, proctored, closed-book measures are the strongest signal of what students actually know (see [[summative-assessment]], [[generative-ai-reduced-study-time-math|proctored retention evidence]]). Without some form of supervision, online exams can inflate grades by capturing AI-assisted rather than independent performance. - **Scalability and cost.** Automated proctoring reduces the need for dedicated physical venues and human invigilators, making monitoring feasible at scale — an advantage for MOOCs and large online programs where traditional proctoring is logistically and financially impractical. - **accessibility and reach.** Proctoring allows students in remote locations to take exams from anywhere, removing geographic and scheduling barriers to credentialing. - **Detection capability.** Advanced ML/DL systems can detect cheating (eye movements, head posture, facial expressions) more reliably than manual observation, and can monitor continuously rather than intermittently. ## Disadvantages, risks, and harms - **Harms to trust and the student-instructor relationship.** The most consequential cost of remote proctoring is often relational. Continuous surveillance communicates that students are presumed dishonest, which can erode [[trust]], corrode the student–instructor relationship, and undermine the sense of shared purpose that supports academic integrity. Monitoring-as-enforcement can crowd out the trust-based, educative approach to integrity that builds responsible use. - **Academic surveillance.** Remote proctoring is a form of academic surveillance that extends institutional monitoring into the student's private home environment. Beyond the exam itself, these systems capture the student's living space, face, voice, and behavior continuously — a level of scrutiny with few precedents in [[higher-ed|higher education]]. Critics argue this normalizes a surveillance culture in which students are presumed guilty until proven honest, reshaping the relationship between institutions and learners and raising questions about proportionality: whether the integrity gains justify subjecting every student to pervasive monitoring for the misdeeds of a few.([[privacy]]), [[trust]] - **Privacy and consent.** AIPS continuously access facial imagery, voice patterns, gaze, and keystroke dynamics, often through persistent audiovisual surveillance. Data handling must comply with frameworks like GDPR and India's PDP Bill, and requires clear consent and secure biometric-data handling. Students often have little choice but to accept monitoring if they wish to take an exam, raising questions about whether consent is genuinely voluntary.([[privacy]]) - **False positives, false accusations, and anxiety.** Systems may flag benign behavior (looking away, adjusting posture) as suspicious, eroding student trust and producing false malpractice accusations — especially where proctors or test-takers lack proficiency. - **Stress and test anxiety.** Taking a proctored exam is itself a source of significant stress and anxiety. Continuous surveillance, fear of being falsely flagged, and the pressure of being watched can raise test anxiety and impair performance — and, per the evidence, stressed students may be *more* likely to resort to dishonest behavior, meaning the monitoring can be counterproductive. Being monitored is stressful and can itself induce the unethical behavior it aims to prevent. - **Equity and the digital divide.** Device dependency, unstable internet, lighting, and hardware variability disproportionately disadvantage rural and low-bandwidth students; model accuracy can vary across demographics and environments, risking unfair flagging.([[digital-divide]]), [[equity-in-ai-education]] A scoping review of proctoring in [[nursing-education|nursing]] assessment shows this is not an edge case: four of its six included studies reported internet connectivity problems, one reported load shedding alongside limited data bundles, and device incompatibility, browser-extension faults and failed environmental scans were routine — so infrastructural inequity, not only model bias, determines who can be assessed at all.([[harerimana-remote-proctoring-nursing-scoping-2026]]) - **Detection gaps and the arms race.** Identity spoofing (photos/video masking), browser use, and copy-paste remain hard to catch reliably; detection accuracy is bounded by dataset limitations, single-model evaluation, and reproducibility gaps. Proctoring does not fully solve integrity, and can create a false sense of security. - **The governance question.** Whether surveillance is the right response versus [[authentic-assessment|assessment redesign]] (oral, process-based, [[eportfolio|portfolio]]) is an open institutional decision; remote proctoring is one tool, not a complete solution.([[governance]]) The exposure that follows when enforcement goes wrong — accusations resting on scores and event logs, retention and onward processing of captured data, and rules that fall unevenly on disabled or non-native-speaker students — is mapped on [[legal-issues-and-risks]]. ## Evidence base - A decade-long [[meta-analysis-systematic-review|systematic review]] of 80 peer-reviewed studies (2014–2024) finds advanced ML/DL proctoring detects cheating more reliably than traditional methods, but is limited by dataset gaps (35% did not fully disclose data), single-model evaluation (40%), reproducibility issues (30%), sparse ethical reporting (only 25%), and inconsistent metrics (20%). False positives/negatives — flagging normal behavior as suspicious or missing subtle cheating — undermine reliability and trust.([[automated-online-exam-proctoring-decade-review-2026]]) - A companion review documents the cheating methods AI must counter (identity spoofing via photos/video, browser/device use, copy-paste) and the practical barriers: test-taker anxiety, proficiency gaps causing false accusations, and infrastructure (webcam, microphone, internet) that is not universally affordable or available. It reports ~37.8% of college and ~41.8% of high-school students admit to cheating — the motivation for monitoring.([[academic-dishonesty-automated-proctoring-ai-2026]]) - **Instructors see no benefit from proctoring exams taken outside class.** [[biology-degree-integrity-genai-cheating-2026|Chan et al. (2026)]] surveyed 56 instructors (47% response rate) in one [[biology-education|biology]] department after coding all 38 syllabi of its core required courses, and asked them to rate each graded category they use for vulnerability to academic dishonesty (0 = minimally to 4 = highly vulnerable). Proctored and unproctored outside-of-class exams were rated *identically* at a median of 3 (moderately vulnerable), while in-person proctored exams were the only category rated minimally vulnerable (median 0) and were significantly less vulnerable than every other category (p_adj < .01). Because lockdown browsers were available for outside-of-class exams, the result is a perception that the tools did not reduce exposure — consistent with the review evidence above that online proctoring is effective only unevenly, and with the documented anxiety, privacy, and false-accusation costs that make the trade-off contested. The stakes are visible in the same study's point accounting: outside-of-class exams carried a mean of 54.2% of the grade in the in-person courses that used them and 59.2% in online courses, where every offering relied on them. - A scoping review of remote proctoring in [[nursing-education|nursing]] [[assessment|student assessment]] maps a spectrum of modalities rather than a single practice — live human invigilation (ProctorU), AI browser-extension proctoring with webcam, microphone and behavior flagging (Honorlock), webcam-plus-lockdown-browser monitoring (Respondus Monitor), an institutionally deployed mobile invigilation app, and fixed test-center desktops under central surveillance — and argues the diversity reflects disparities in infrastructure and institutional capacity rather than a shared standard. Only six studies met its inclusion criteria (1,567 nursing students in the USA, the UK, Southern Africa and Egypt), and none was published before 2021. What the review establishes is that surveillance produces deterrence *perceptions* and a heavy faculty review burden: the 98–100% agreement that webcam monitoring and lockdown browsers deter cheating is [[self-report-measures|self-report]] from one graduate nurse practitioner program, high-risk AI alerts were rare at no more than 5%, yet frequent minor alerts produced false positives requiring time-intensive review, and South African lecturers reported continued dishonesty under active monitoring. The one comparative performance study points the other way — an in-person proctored cohort scored significantly higher on the HESI Exit Exam and NCLEX readiness than the ProctorU cohort — and the review treats the link between surveillance and more honest learning as an open question rather than a settled benefit.([[harerimana-remote-proctoring-nursing-scoping-2026]]) ## Recommended directions - **Hybrid human–AI oversight.** Pair automated flagging with [[human-in-the-loop-ai|human review]] to reduce false positives and keep judgment contextual. - **Privacy-preserving architecture.** Edge processing, anonymization, and on-device handling reduce the invasiveness of continuous data capture. - **Equity-aware deployment.** Diverse, geographically-inclusive datasets and lightweight models for low-resource environments; accessible alternatives for students without reliable devices/connectivity. - **Educate before you surveil.** Prefer fostering [[ai-literacy]] and [[reducing-ai-misuse|responsible use]] through culture and trust-building, reserving proctoring for the high-stakes cases that genuinely require it. ## Connected Concepts - [[anxiety-and-stress]] - [[academic-integrity]] - [[summative-assessment]] - [[assessment]] - [[automated-assessment]] - [[online-teaching-and-learning]] - [[ai-misuse-learning-harm]] - [[privacy]] - [[equity-in-ai-education]] - [[digital-divide]] - [[student-experience]] - [[trust]] - [[governance]] - [[legal-issues-and-risks]] - [[higher-ed]] ## Connected Articles - [[automated-online-exam-proctoring-decade-review-2026]] — Decade-long systematic review of automated online exam proctoring - [[academic-dishonesty-automated-proctoring-ai-2026]] — Comprehensive review of academic dishonesty in automated proctoring - [[ssaho-ai-academic-integrity-review-2025]] — AI and academic integrity: systematic review - [[conijn-fear-big-brother-proctored-exams-2022]] — The fear of Big Brother: proctoring's negative side-effects on test anxiety - [[biology-degree-integrity-genai-cheating-2026]] — Can students cheat their way to a biology degree? A case study of the vulnerability of biology course grades to academic dishonesty in the era of generative AI - [[harerimana-remote-proctoring-nursing-scoping-2026]] — Under surveillance: mapping remote proctoring practices in nursing student assessment --- ## [Learning Gains](https://edtechdev.github.io/aied/concepts/learning-gains/) > **Learning gains** — measurable improvements in student knowledge, skills, or competencies resulting from educational interventions, including AI-assisted instruction. In [[ai-education|AI in education]] research, learning gains serve as the primary outcome measure for evaluating whether AI tools actually improve learning — not just [[student-engagement|engagement]] or satisfaction. ## Questions to Consider - Have you ever felt you learned a lot from an activity, only to fail a test that measured something different? The page distinguishes immediate AI-supported performance from durable learning — how might those two diverge in your own experience? - A core finding is that generative AI can inflate scores on AI-assisted homework while lowering scores on proctored, closed-book measures. If you were evaluating whether an AI tool really helps students learn, which outcome would you trust and why? - Research shows unguided reliance on AI predicts worse learning gains, while structured use predicts better ones — the same tool, opposite outcomes. What distinguishes 'structured' from 'unguided' use in a real classroom? - Hint buttons correlate with reduced learning: more hints, less learning. Have you ever been tempted to reach for a hint or an answer the moment you were stuck? What does that suggest about how much struggle is actually necessary for learning? - A large meta-analysis pooled many studies and found AI-enabled [[edtech-platform|EdTech]] raised learning by a modest amount, with no advantage for generative AI over earlier adaptive tools. How should this cautious, pooled estimate change how you read exciting claims about a single AI product's effectiveness? ## Introduction Learning gains are the ultimate test of any educational technology. In the knowledge base's research, they appear as dependent variables in [[rct|randomized controlled trials]], pre-post comparisons in quasi-experimental studies, and correlational analyses linking AI tool usage to academic outcomes. **Terminology.** The *outcome* sense of achievement — student achievement, academic achievement, learning achievement, prior achievement, achievement gaps — is treated as a synonym for learning gains here, and those phrases link to this page. The *felt* sense ("a sense of achievement") is a motivational experience rather than a measured outcome; see [[motivation]] and [[self-efficacy]]. Achievement-goal theory ("achievement goals", goal orientations) is likewise a motivational construct and links to [[motivation]]. Key findings from the knowledge base: - **[[adaptive-pretesting-retention|Adaptive pretesting]]** research examines whether [[generative-ai|GenAI-enabled]] [[retrieval-spacing-interleaving|pretesting]] produces durable learning gains that persist beyond immediate testing. - **[[genai-meta-analysis-programming-learning|Meta-analyses of GenAI in programming]]** find positive learning gains from structured AI use but negative effects from unguided reliance — a key distinction between productivity and durable learning. - **[[lak2026-hint-button-unproductive-use|Hint button research]]** shows negative associations between hint abuse and learning gains — more hints correlate with less learning. - **[[instructional-guidance-genai-learning|Instructional guidance]]** studies demonstrate that learning gains depend on HOW AI is used, not just WHETHER it's available. - **[[burneo-can-edtech-close-learning-gaps-2026|World Bank meta-analysis]]** pools 191 effect sizes from 14 RCTs to estimate that adaptive and AI-enabled EdTech raises learning by ~0.125 sd on average — above the median for education RCTs — while finding no advantage for generative AI over earlier adaptive tools. - **Gains are content–treatment interactions, not constants.** [[rachatasumrit-example-problem-ratio-2026|Rachatasumrit et al. (2025)]] show the optimal example–problem ratio depends on knowledge content: pure [[mastery-learning|retrieval practice]] yields higher gains for verbatim facts, while example-integrated practice (alternating worked examples and problems) yields higher gains for generalizable skills — direct evidence that "more practice" is not always better and that gains hinge on matching the training schedule to the knowledge component being learned. ### The AI-era measurement problem A central theme in the knowledge base's learning-gains research is that **generative AI can inflate apparent performance without producing learning gains** — and that the choice of outcome measure determines whether this is visible. [[generative-ai-reduced-study-time-math|Research]] and [[stromberg-generative-ai-learning-penalty-secondary-2026|large-scale field data]] show a sharp divergence: AI use improves scores on AI-assisted homework while *lowering* scores on proctored, closed-book, unassisted measures. [[genai-performance-vs-learning|Performance-vs-learning research]] and [[young-people-learning-generative-ai-rapid-review-2026|rapid reviews]] therefore distinguish immediate AI-supported performance from durable learning, and treat unassisted summative measures (see [[summative-assessment]]) as the reliable signal of genuine learning gains. ### What the efficacy research shows Across the knowledge base's [[rct|RCTs]], [[meta-analysis-systematic-review|meta-analyses]], and field studies, a consistent picture of **learning efficacy** (which AI interventions actually produce learning gains, and how large) emerges: - **Meta-analytic evidence is broadly positive but conditional.** [[genai-educational-outcomes-meta-analysis|A comprehensive meta-analysis of 53 studies]] (Dong 2026) finds generative AI generally outperforms traditional approaches on academic achievement, [[critical-thinking|higher-order thinking]], and writing — with [[ai-feedback-quality|AI feedback]] particularly effective — though game-assisted GenAI shows no significant added benefit and gains vary by country and outcome. [[genai-meta-analysis-programming-learning|The GenAI-and-programming meta-analysis]] finds large productivity gains but no significant learning gain (g ≈ 0), separating task-efficiency from durable learning. [[robot-assisted-language-learning-meta-analysis-2026|A language-learning meta-analysis]] finds positive but modest learning gains from AI-enhanced [[embodied-learning|embodied]] robots. For [[ai-literacy|AI literacy]] specifically, [[liu-ai-literacy-interventions-meta-analysis-2026|a three-level meta-analysis of 59 studies]] estimates a large overall effect (g = 0.837) — but the wide prediction interval and the finding that knowledge-focused interventions outperformed those targeting skills, attitudes, or [[ethics]] caution that the *outcome measured* shapes the apparent gain, echoing the broader point that AI-related gains depend on what and how you assess. - **Well-designed [[intelligent-tutoring|AI tutors]] produce real gains.** A two-year cluster RCT ([[one-click-away-khanmigo-two-year-school-experiment-2026|Khanmigo]]) found AI tutoring raised math achievement ~1.3 national percentile ranks per term (~0.06–0.08 SD/school year, ~0.14 SD for a full year), gains resembling practice without AI — demonstrating that *engagement*, not model capability, is the binding constraint. [[making-ai-tutoring-productive-mastery-math-2026|NUMI]] showed AI support improved next-attempt correctness after mistakes with more time per question — a "productive slowdown" that builds durable mastery. [[virtual-tutoring-computer-assisted-learning-takeup-2026|Virtual tutoring]] found the binding constraint is take-up and sustained participation, not tutor quality. - **AI can match human help.** [[chatgpt-hints-human-tutor-learning-gains-2024|ChatGPT-generated help]] produces learning gains equivalent to human tutor-authored help on [[math-education|mathematics]] skills — evidence that generative AI can be as efficacious as human [[scaffolding]] when used appropriately. - **Unguarded AI can harm learning.** The guardrail RCT ([[generative-ai-guardrails-harm-learning|PNAS 2025]]) found an unguarded ChatGPT-style tutor raised assisted practice +48% but *reduced* unassisted exam scores −17%, while a guardrailed (hint-not-answer) tutor eliminated the harm. This is the sharpest demonstration that **learning efficacy is design-contingent**: the same class of tool can be a strong learning gain or a net harm depending on how it is configured. - **Perceived vs. actual efficacy diverge.** [[ai-literacy-assessment-misalignment|Self-reported performance misaligns with measured performance]], and [[absent-cognitive-baseline-2026|the absent cognitive baseline]] shows AI-native students overestimate their learning — so efficacy claims based on self-report are unreliable without objective outcome measures. [[self-report-measures]] collects the cases where reported and measured outcomes come apart. **Takeaway:** the weight of evidence supports **modest, conditional, and design-dependent learning gains** from AI — real when AI is structured to coach rather than answer, guardrailed, and paired with unassisted outcome measures, and absent or negative when it substitutes for the learner's own effort. This is why learning gains as an outcome must be measured with valid, AI-resistant instruments and why [[ai-ed-evaluation]] pairs efficacy claims with [[research-methods-aied|methodological]] scrutiny. ### A field-level map of AI's effects on learning and achievement The knowledge base's corpus — meta-analyses, RCTs, quasi-experiments, and field studies — supports a nuanced, sometimes contradictory picture. Findings cluster into positive, negative, and conditional effects. **Positive effects (AI can produce genuine learning gains):** - **Meta-analytic evidence is broadly positive but conditional.** [[genai-educational-outcomes-meta-analysis|A 53-study meta-analysis]] (Dong 2026) finds generative AI generally outperforms traditional approaches on academic achievement, higher-order thinking, and writing, with AI [[feedback]] especially effective. A meta-analysis of 29 experiments ([[zhao-genai-higher-order-thinking-meta-2026]]) finds a moderate positive effect on higher-order thinking, strongest for [[problem-solving]], but limited for [[creativity]]. - **Tutoring-specific AI reliably outperforms general-purpose AI.** [[stanford-evidence-base-ai-k12-2026|The Stanford SCALE review]] and the umbrella review ([[genai-higher-education-systematic-review-2026]]) converge: pedagogically designed agents with hints and step-by-step scaffolding produce real gains where open [[conversational-ai|chatbots]] often do not. - **Well-designed AI tutors produce real, measurable gains.** The two-year Khanmigo RCT ([[one-click-away-khanmigo-two-year-school-experiment-2026]]) raised math achievement ~1.3 national percentile ranks per term; [[making-ai-tutoring-productive-mastery-math-2026|NUMI]] improved next-attempt correctness after mistakes; [[virtual-tutoring-computer-assisted-learning-takeup-2026|virtual tutoring]] showed take-up is the binding constraint. - **AI can match human help.** [[chatgpt-hints-human-tutor-learning-gains-2024|ChatGPT-generated help]] produces learning gains equivalent to human tutor-authored help on math skills. - **Feedback-focused AI works.** [[genai-educational-outcomes-meta-analysis|AI feedback]] is among the most effective GenAI applications; [[ai-feedback-critical-thinking-writing-2026|AI feedback for writing]] improves critical thinking when coupled with instruction. In a field experiment, [[gpt4-feedback-student-activation-2026|Geschwind et al. (2026)]] found students receiving individual GPT-4 feedback on open-ended tasks showed the largest content learning gains (~0.11, p < 0.10; rising to 0.16 among those who actually received prior feedback) — an effect driven by reliable, consistent AI provision rather than inherent superiority, since peer outcomes matched AI when high-quality [[peer-assessment|peer feedback]] was actually delivered. LLM critique partners in writing extend this to iterative, collaborative feedback: [[oppenheimer-llms-collaborative-learning-partners-2026|Oppenheimer, Cash & Connell Pensky (2025)]] found significant gains across a semester on argumentative writing, [[prompt-engineering|prompt engineering]], and response-to-AI-feedback quality (all p < .001, roughly a full standard deviation per dimension), with gains appearing even on essays written without LLM support — evidence of durable skill rather than mere tool dependency, though the lack of a control condition limits causal attribution. - **Structured scaffolds yield gains.** [[scaffolding-srl-feedback-genai-human-peers|Scaffolded self-regulated feedback]] and [[learner-ai-interaction-patterns-oop|interaction design]] show small but significant gains when AI is designed to coach. - **AI in collaborative and problem-based contexts helps when structured.** [[ai-enhanced-pbl-chatgpt-scaffolding-2026|AI-enhanced PBL]], [[ai-assisted-collaborative-learning-model-dbr|AI-assisted collaborative learning]], and [[ccct-cooperative-learning-technique|AI-designed cooperative techniques]] report meaningful gains. **Negative effects (AI can reduce or fail to improve learning):** - **Unguarded AI can harm learning.** The PNAS guardrail RCT ([[generative-ai-guardrails-harm-learning]]) found an unguarded ChatGPT-style tutor raised assisted practice +48% but *reduced* unassisted exam scores −17%; [[guardrails]] eliminated the harm. - **Faster completion ≠ learning.** [[generative-ai-reduced-study-time-math|GenAI reduced study time]] on math problems and the knowledge they build; [[genai-performance-vs-learning|performance-vs-learning research]] shows AI inflates AI-assisted performance while lowering proctored, closed-book, unassisted scores. - **Homework outsourcing harms learning.** [[stromberg-generative-ai-learning-penalty-secondary-2026|Field data]] show a generative-AI learning penalty when students outsource homework. - **[[llm]] reliance correlates with lower grades.** [[jost-llm-programming-education-learning-outcomes|Jošt et al.]] found significant negative correlations between LLM use for code generation (rho = −0.305) and debugging (rho = −0.360) and final grades. - **AI can harm [[teacher-role|teaching]] and achievement for some.** [[genai-can-harm-teaching-rct-2026|A teacher-facing GenAI RCT]] found below-median teachers' students lost ground (−0.129 SD). - **Hint abuse correlates with less learning.** [[lak2026-hint-button-unproductive-use|Hint button research]] shows more hints, used unproductively, correlate with less learning. - **Meta-analytic programming gains are illusory.** [[genai-meta-analysis-programming-learning|The GenAI-and-programming meta-analysis]] finds large productivity gains but no significant learning gain (g ≈ 0), separating task-efficiency from durable learning. **Conditional and mixed effects (context determines direction):** - **Design is decisive.** The same tool class can be a strong gain or a net harm depending on configuration ([[generative-ai-guardrails-harm-learning]], [[stanford-evidence-base-ai-k12-2026]]). - **How AI is used matters more than whether.** [[jost-llm-programming-education-learning-outcomes|Use for explanations was benign; use for code generation was harmful]]; [[genai-over-reliance-learning-2026|over-reliance]] erodes gains. - **Duration and self-[[regulation]] moderate effects.** [[zhao-genai-higher-order-thinking-meta-2026|Effects were strongest at 8–16 weeks]] and for learners with higher [[self-regulated-learning]]. - **AI-IBL supports creativity but not necessarily problem-solving.** [[mujib-ai-ibl-creative-math-2026|Mujib et al.]] improved creative performance and attitudes but not critical problem-solving. - **[[student-experience|Student experience]] diverges from measured gains.** [[absent-cognitive-baseline-2026|AI-native students overestimate their learning]], and [[ai-literacy-assessment-misalignment|self-report misaligns with performance]] — perceptions of gains are not reliable evidence of them. ### Meta-analytic learning-gain estimates are inflated — read them with caution The learning-gain numbers that dominate this page's efficacy summary — especially pooled effect sizes from meta-analyses — must be read with a strong caveat: a wave of meta-research shows the field's positive synthesis estimates are inflated by publication bias, construct incoherence, and methodological shortcuts. - [[bartos-ai-learning-meta-meta-analysis-2026|A study-level meta-meta-analysis of 1,840 effect sizes from 67 meta-analyses]] finds the publication-bias-adjusted average AI effect is roughly **one-third** the reported magnitude (SMD ≈ 0.196 vs. a median of 0.67), with a prediction interval spanning large harm to large benefit and no moderator producing consistent gains. Even a single extreme study contributes almost no information given the heterogeneity — meaning *more studies of the current type will not settle the question*; only high-quality, pre-registered, replication-oriented trials will. - **A second-order meta-analysis of the synthesis layer itself.** [[ai-education-effects-second-order-meta-analysis-2026|Emslander et al. (2026)]] invert the usual unit of analysis: treating 45 meta-analyses as their data and correcting for study overlap (818 unique primary studies, weighted by uniqueness rather than discarded), they pool 129 meta-analytic effect sizes to g = .57 [.50, .65], with outcome clusters running from .35 for STEM performance to .68 for general academic outcomes. Two results bear on the numbers above. AI type did not moderate the effect (p = .63) — [[intelligent-tutoring|intelligent tutoring systems]], [[conversational-ai|chatbots]] and ChatGPT are statistically indistinguishable at this layer, so the field's pooled gains are not being driven by generative AI as such — and the one clear moderator was educational level (tertiary .61 vs. .30 for younger samples), which is partly a study-design artifact given that 74% of the corpus tested tertiary students and [[k-12|K-12]] evidence is comparatively thin. Its quality appraisal is the sharpest in the corpus: the 45 meta-analyses averaged 9.3 of 17 on an AMSTAR-adapted scale, none preregistered, and only 23 of 45 had 80% power to detect the effect they themselves reported. Its two publication-bias tests also disagree — a symmetric funnel plot (Kendall's τ = −.01, p = .868) against PET-PEESE (B = 0.9, p = .009) indicating inflation — so the SOMA's own medium effect arrives with an unresolved bias cloud over it. - [[oneill-presumed-effective-meta-analysis-2026|A forensic audit of 14 high-impact AIED meta-analyses]] finds none provided a valid basis for its pooled learning-gain claim: none had a coherent outcome construct, all had unresolved extreme heterogeneity (I² ranged from 77.2% to 94.4% across the 13 meta-analyses that reported it, and 12 of those 13 exceeded 80%), twelve treated dependent effect sizes as independent, and none validly assessed publication bias. A majority of randomly vetted primary studies were mismatched to the meta-analytic claim. The errors reached policy: two of the audited meta-analyses treated study sample size as class size and, on the strength of a miscalculated primary study, concluded that 21 to 40 students is the ideal intervention size and recommended that GenAI interventions be designed for that number. - **A 2026 [[stem-education|STEM]] synthesis that fixes two of the audited defects — and still illustrates a third.** [[ai-supported-instruction-stem-meta-analysis-2026|Doğan and colleagues (2026)]] pooled 35 experimental and quasi-experimental STEM studies with one effect size per study to avoid the dependent-effect-size problem, and ran a full publication-bias battery (funnel plot, Begg's and Egger's tests, trim-and-fill, Rosenthal's fail-safe N = 2404) rather than asserting symmetry — the two shortcuts the audit above found in every meta-analysis it examined. Two cautions survive. First, the authors applied no formal quality appraisal and treated the inclusion criteria as the rigor threshold, so studies of very different designs entered the pool unweighted by quality. Second, the headline heterogeneity depends entirely on which model you read: the same 35 studies are reported as I² = 82.98% under the fixed-effect model and I² = 15.75% under the random-effects model, so a reader who encounters the numbers without the model label can conclude the corpus is homogeneous when it is not. Its pooled estimate (g = 0.670, 95% CI [0.491, 0.848]) is therefore the kind of figure the section above describes: better built than most, and still an upper bound. - [[weidlich-chatgpt-effect-search-cause-2025|A media/methods critique]] shows many "learning gains" were not measured validly — outcomes were often self-reported skills or performance measured *during* AI assistance rather than durable, unassisted learning. **Bottom line for the gains numbers above:** treat large pooled AI effect sizes as upper bounds, not point estimates. Prefer the learning-gain evidence from well-designed [[rct|RCTs]] and field studies with unassisted, standardized outcome measures (the guardrail RCT, the Khanmigo and NUMI experiments, and the World Bank EdTech meta-analysis cited above), and read meta-analytic gains as provisional and likely over-stated until synthesis quality improves. The problem also outlives correction: of the papers published after one audited meta-analysis was retracted, 60% still cited it as authoritative support for large ChatGPT learning gains and none acknowledged the retraction, so a withdrawn estimate keeps propagating. This is why the [[limitations-in-aied-research|Limitations in AIEd Research]] page now documents the meta-analytic evidence crisis in detail. ### Measuring what matters Learning gains connect to [[assessment-validity]] — if [[assessment|assessments]] fail to capture deeper understanding, learning gain measures are misleading. They also intersect with [[cognitive-offloading|Over-Reliance]] and [[cognitive-offloading]], where apparent performance improvements may mask learning losses, and with [[rct]] (randomized trials as the gold-standard design for detecting causal learning gains), and with [[meta-analysis-systematic-review]] (pooling effect sizes across studies to establish the field's efficacy evidence). - **Significant pre/post gains from mistake-based AI [[pedagogy]]:** [[pedagogy-ai-mistakes|Hosseini (2026)]]'s database design course (n=13) showed large, significant learning gains on identical pre/post items (mean 4.25→6.83/7, Cohen's *d*=1.49, *p*<.001), with gains uncorrelated with prior AI or database confidence — the AI-integrated critique-refinement design benefited students regardless of initial perceptions. ## Connected Concepts - [[rct]] - [[meta-analysis-systematic-review]] - [[formative-assessment]] - [[summative-assessment]] - [[cognitive-offloading]] - [[math-education]] - [[human-in-the-loop-ai]] - [[affective-tutoring]] - [[theory-development-aied]] — Theory Development in AI in Education - [[self-report-measures]] - [[research-methods-aied]] — Research Methods in AIED (DBR section) - [[social-emotional-learning]] — Social-Emotional Learning ## Connected Articles - [[genai-performance-vs-learning]] — why assisted performance is not a learning outcome (Yan et al. 2025) - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — Virtual tutoring with CAL: an experiment in take-up and learning - [[making-ai-tutoring-productive-mastery-math-2026]] — Making AI tutoring productive: mastery-based math practice - [[one-click-away-khanmigo-two-year-school-experiment-2026]] — One Click Away: Khanmigo in a two-year school experiment - [[studychat-student-dialogues-chatgpt-ai-course-2026]] — The StudyChat dataset of student–LLM dialogues in an AI course - [[ai-literacy-assessment-misalignment]] — AI Literacy Assessment: Self-Reported vs Performance Misalignment - [[generative-ai-reduced-study-time-math]] — Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build - [[genai-meta-analysis-programming-learning]] — A meta-analysis of the effect of generative AI on productivity and learning in programming - [[liu-ai-literacy-interventions-meta-analysis-2026]] — Meta-analysis of AI literacy intervention effects - [[chatgpt-hints-human-tutor-learning-gains-2024]] — ChatGPT help produces learning gains equivalent to human tutor help - [[generative-ai-guardrails-harm-learning]] — Generative AI without guardrails can harm learning (PNAS 2025 RCT) - [[absent-cognitive-baseline-2026]] — The Absent Cognitive Baseline: Theorizing a Structural Gap in AI-Native College Students' Academic Self-Assessment - [[oecd-digital-education-outlook-2026]] — OECD Digital Education Outlook 2026 - [[ai-changing-teaching-workflows]] — How AI Is Changing Teaching Workflows - [[robot-assisted-language-learning-meta-analysis-2026]] — Meta-analysis of AI-enhanced embodied robot-assisted language learning - [[genai-educational-outcomes-meta-analysis]] - [[young-people-learning-generative-ai-rapid-review-2026]] — Immediate performance vs durable learning distinction - [[stromberg-generative-ai-learning-penalty-secondary-2026]] — The generative AI learning penalty: homework outsourcing harms learning - [[ba-ai-agents-cscl-review-2026]] — AI agents in computer-supported collaborative learning review - [[ai-assisted-collaborative-learning-model-dbr]] — AI-Assisted Collaborative Learning model DBR (critical thinking +24.1%, problem-solving gains) - [[stanford-evidence-base-ai-k12-2026]] — Tutoring-specific vs general AI - [[jost-llm-programming-education-learning-outcomes]] — LLM reliance and grades in coding (negative correlations) - [[genai-can-harm-teaching-rct-2026]] — Generative AI can harm teaching (RCT) - [[zhao-genai-higher-order-thinking-meta-2026]] — GenAI and higher-order thinking meta-analysis - [[mujib-ai-ibl-creative-math-2026]] — AI-supported IBL and creative math performance - [[ai-enhanced-pbl-chatgpt-scaffolding-2026]] — AI-enhanced PBL scaffolding gains - [[ccct-cooperative-learning-technique]] — AI-designed cooperative learning technique - [[scaffolding-srl-feedback-genai-human-peers]] — Scaffolded self-regulated feedback gains - [[learner-ai-interaction-patterns-oop]] — Interaction patterns and learning gains in OOP - [[ai-feedback-critical-thinking-writing-2026]] — AI feedback and critical thinking in writing - [[pedagogy-ai-mistakes]] — The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking (Hosseini 2026) - [[graph-its-adaptive-algorithms-2026]] — Graph-Based Intelligent Tutoring for Dynamic Domains (2026) - [[computational-thinking-aica-2026]] — Computational Thinking Levels and AI Coding Assistants (2026) - [[adaptive-scaffolding-cognitive-engagement-its]] — Adaptive ICAP scaffolding in an ITS (BKT vs DRL) - [[burneo-can-edtech-close-learning-gaps-2026]] — Meta-analysis: adaptive/AI EdTech raises learning ~0.125 sd - [[gpt4-feedback-student-activation-2026]] - [[rachatasumrit-example-problem-ratio-2026]] - [[oppenheimer-llms-collaborative-learning-partners-2026]] - [[weidlich-chatgpt-effect-search-cause-2025]] — ChatGPT in Education: An Effect in Search of a Cause (media-comparison critique of gains measures) - [[bartos-ai-learning-meta-meta-analysis-2026]] — Meta-meta-analysis: bias-adjusted AI learning-gain effects ~1/3 of reported size - [[oneill-presumed-effective-meta-analysis-2026]] — Presumed Effective: audit of flawed AIED meta-analyses - [[ai-supported-instruction-stem-meta-analysis-2026]] — A STEM synthesis that controlled dependent effect sizes and tested publication bias, but applied no quality appraisal (Doğan et al. 2026) - [[ai-education-effects-second-order-meta-analysis-2026]] — Second-order meta-analysis of 45 AI-in-education meta-analyses, with overlap-corrected effects and a quality appraisal of the synthesis layer (Emslander et al. 2026) --- ## [Research Methods in AIED](https://edtechdev.github.io/aied/concepts/research-methods-aied/) > **Research methods in AIED** — the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study [[ai-education|AI in education]]: whether and how AI tools support (or harm) learning, and under what conditions. The knowledge base's corpus spans experimental, survey, qualitative, design-based, computational-benchmark, and review methods. Each has distinct strengths and limitations, and choosing among them involves trade-offs among internal validity (confidence in causal claims), external validity (generalizability), ecological validity (real-world authenticity), and the feasibility of studying fast-moving AI tools. ## Questions to Consider - The page's central tension: the strongest designs for causal inference (randomized experiments) are the hardest to run in real classrooms, while the most authentic settings offer weaker causal control. If you had to decide whether an [[intelligent-tutoring|AI tutor]] helps learning, which of these two failures would you rather live with — and why? - Before you read, can you name the difference between internal, external, and ecological validity? The page argues every design trades these off. How might a study that's rigorously causal still tell you almost nothing useful about a real classroom? - A benchmark shows an AI scores high on accuracy, but the page insists high benchmark accuracy does not entail educational effectiveness. Why might a system that 'passes the test' still fail to help students learn — and what kind of evidence is missing? - Design-based research iterates on a real intervention but can't attribute gains to a specific mechanism, while an RCT isolates causes but runs in artificial conditions. Given the fast pace of AI change, how long do you think a rigorous RCT remains relevant before the tool it tested is obsolete? - Delphi expert consensus establishes agreement among experts, not empirical effect. When is it legitimate to build a competency framework from what experts believe, versus from data about what works — and how would you tell the difference in practice? - The page advocates triangulation — combining benchmark evaluation, experiments, measurement, and qualitative work to judge both whether a tool works and how. Before you read, where in a claim like 'this AI improves learning' would each method be needed to make you confident? ## Introduction The central tension in AIED research is that the strongest designs for causal inference — randomized experiments — are often the hardest to run with authentic AI tools in real classrooms, while the most authentic settings (field deployments, case studies, log-data analyses) offer weaker causal control. No single method resolves this; the field advances by triangulating across methods, and by being explicit about what kind of claim each design can support. Every method also carries cross-cutting limitations — generalizability, measurement validity, the fast pace of AI change, reproducibility, and weak theory use — that readers must weigh; see [[limitations-in-aied-research]]. The page's subject is method rather than findings. [[learning-sciences]] is the substantive field these methods serve: where this page covers how a study should be designed, measured and reported, that page covers what the field has established about how people learn and how learning environments should be designed, and it treats design-based and mixed methods as the learning sciences' signature approaches rather than two options among many. ### Reporting rigor and the TEP-AIED model The reporting quality of AI-in-education studies is itself a research concern. [[tep-aied-model-reporting-2026|The TEP-AIED model (Hwang, Xie, Wah & Gasevic, 2026)]] offers a structured framework for presenting AI-in-education research with rigor, organizing the essential components of a study report — theory/technology/educational problem framing, design, data, analysis, and results — so that readers and reviewers can assess whether claims are supported and whether the work is reproducible. It responds to the field's chronic weaknesses in reporting (vague tool descriptions, unstated model versions, omitted evaluation details) that the [[limitations-in-aied-research|limitations]] page documents. Reporting frameworks like TEP-AIED sit alongside established reporting checklists (e.g., [[rct|CONSORT]]-style guidance for trials, PRISMA-style guidance for reviews) as part of the field's broader move toward [[educational-measurement|methodological transparency]] and reproducibility. [[raise-framework-ai-education-reporting-2026|RAISE (Allison, 2026)]] approaches the same problem from the opposite direction — as a checklist rather than a narrative structure. It sets out **30 items across ten thematic domains** (educational justification and theoretical grounding, AI system specification, AI role and interaction, [[accessibility]] and cultural fit, setting and participants, human involvement, study design and evaluation, ethics and trustworthiness, transparency and reproducibility, and limitations and implications), with an editable version and a companion **Ethics and Risk Matrix** covering learner agency, [[equity-in-ai-education|equity]] of access, data [[governance]] and algorithmic transparency. The two frameworks address each other directly: TEP-AIED characterizes RAISE as comprehensive but faults its breadth, arguing that "its breadth and granularity may make it complex and less accessible for routine empirical applications," while RAISE's own framing is that it mandates no method or model and only requires that choices be made visible. Read together they mark the trade-off in this literature — the fuller audit checklist versus the leaner three-dimension narrative — and they converge on the same non-negotiables: name and version the AI system, disclose prompts and interaction design, define treatment and comparison conditions, report ethical review and risk mitigation, and state whether outcomes measure performance, retention, or transfer. For a study to be adjudicated on those terms, the reporting instrument has to be adopted at design time rather than assembled at manuscript stage, which is the point both frameworks insist on. Corpus-level historical analysis is itself a methodological choice with transparency obligations: [[rismanchian-ai-education-four-decades-aixed-2026|Rismanchian & Doroudi]] locate each paper in their AI×Ed framework based on author judgment of abstracts and full texts, explicitly acknowledge that this is not a systematic or scalable data-driven categorization, and make their full dataset publicly available as supplementary material for replication. ### Experimental and quasi-experimental designs An **efficacy study** tests whether an intervention produces its intended learning effect, typically using experimental or quasi-experimental designs that compare outcomes with and without the intervention. Experiments randomly assign learners to conditions (e.g., AI tutor vs. human tutor, or AI-scaffolded vs. unassisted) to estimate causal effects on outcomes like learning gains, [[student-engagement|engagement]], or motivation. **Randomized controlled trials** are the gold standard for internal validity. [[access-not-enough-ai-tutoring-2026|A randomized field study of human support plus AI tutoring]] and [[genai-can-harm-teaching-rct-2026|an RCT on generative AI in teaching]] use assignment to isolate causal effects. **Quasi-experimental** designs (pre/post, between-subjects, or matched groups without randomization) are more feasible in intact classrooms but weaker on causal claims. - **Strengths:** strongest causal inference; clean outcome measurement; supports effect-size estimation and efficacy claims. - **Limitations:** costly and slow; artificial conditions can reduce ecological validity; fast-changing AI tools make long experiments date quickly; small samples often underpower detection of meaningful effects; [[ethics|ethical]] constraints on withholding potentially helpful tools. - **Exemplars:** [[access-not-enough-ai-tutoring-2026]], [[genai-can-harm-teaching-rct-2026]], [[adaptive-pretesting-retention]], [[agent-voice-accents-k12-group-learning]], [[ai-use-critical-thinking-medical-students-2026]]. ### Survey and structural-equation-modeling studies Cross-sectional surveys measure self-reported attitudes, perceptions, motivation, [[self-efficacy]], and [[technology-acceptance-model|technology acceptance]], often modeled with regression or structural equation modeling (SEM/PLS-SEM) to test hypothesized relationships and mediators. These dominate the knowledge base's corpus, particularly for acceptance, motivation, and psychological-mechanism questions. - **Strengths:** large samples; broad, low-cost coverage; can test complex mediational models of psychological mechanisms; feasible for studying attitudes that are hard to observe. - **Limitations:** cross-sectional data cannot establish causation; common-method/self-report bias; convenience sampling limits generalizability; mediators inferred from covariance, not manipulation. The instrument itself deserves separate scrutiny. What a questionnaire, interview, or diary can and cannot establish — and the documented gap between what people report and what they do — is gathered on [[self-report-measures]]. - **Exemplars:** [[acceptance-ai-english-tools-2026]], [[genai-motivation-engagement-2026]], [[ai-autonomous-learning-accomplishment-2026]], [[genai-over-reliance-learning-2026]], [[ai-use-critical-thinking-medical-students-2026]]. ### Qualitative methods Interviews, focus groups, and thematic analysis produce rich, contextual accounts of how students and teachers experience AI tools, the meanings they attach to them, and the tensions and harms that standardized measures miss. [[hazra-safetutors-pedagogical-safety-2026|Research on AI tutor safety]] and [[ai-changing-teaching-workflows|how AI changes teaching workflows]] rely heavily on qualitative evidence. See the dedicated [[qualitative-research]] concept page for the full treatment of qualitative approaches — thematic analysis, grounded theory, phenomenology/phenomenography, discourse analysis, observations and ethnography, case studies, and interviews/focus groups — each with knowledge base exemplars. - **Strengths:** deep ecological and conceptual insight; surfaces unexpected phenomena, risks, and mechanisms; essential for theory-building and for studying contested constructs like trust, [[agency|autonomy]], and authorship. - **Limitations:** limited generalizability; interpretive and researcher-dependent; small samples; weaker support for causal claims; findings can be hard to synthesize across studies. - **Exemplars:** [[hazra-safetutors-pedagogical-safety-2026]], [[ai-changing-teaching-workflows]], [[scaffolding-critical-engagement-genai-minority-students]]. ### Mixed-methods designs Mixed-methods studies combine quantitative and qualitative strands — often sequentially (e.g., QUAL→QUAN→qual) — so that qualitative data explains or contextualizes quantitative findings. [[genai-over-reliance-learning-2026|A mixed-method study of GenAI and sustainable learning]] pairs three-wave surveys with educator interviews; [[t2i-competence-paradox-2026|the competence-paradox study]] uses instructor focus groups, a student survey, and follow-up interviews. - **Strengths:** triangulation increases confidence; quantitative breadth plus qualitative depth; can explain unexpected results and bridge mechanism and magnitude. - **Limitations:** complex, resource-intensive, and methodologically demanding; integration can be shallow if not carefully designed; still inherits the weaknesses of each strand (e.g., self-report). - **Exemplars:** [[genai-over-reliance-learning-2026]], [[t2i-competence-paradox-2026]], [[same-ai-different-pathways]], [[fouad-bentley-trust-utility-gap-physics-2026]]. ### Design-based research (DBR) DBR iteratively designs, implements, and refines an educational intervention in authentic contexts, cycling between theory, design, and real-world practice. It is prominent in the knowledge base for developing AI learning environments and [[pedagogy|pedagogical]] models. See the dedicated [[design-based-research]] concept page for the full DBR cycle, exemplars, and its strengths/limitations. A canonical AIEd example is the AI-Assisted [[collaborative-learning|Collaborative Learning]] model study ([[ai-assisted-collaborative-learning-model-dbr|Putra et al.]]), which ran a four-phase DBR cycle — needs analysis, model design, eight-week classroom implementation, and model refinement — iterating on a four-stage learning cycle (problem identification → AI-assisted collaborative inquiry → collaborative [[problem-solving]] → reflection and presentation). Other exemplars develop [[ai-literacy]] [[teacher-education|teacher training]] ([[genai-literacy-training-teacher-education-dbr-2026]]) and [[generative-ai|GenAI]] [[scaffolding]] for [[critical-thinking|critical thinking]] ([[critical-thinking-genai-scaffolding]]). - **Strengths:** high ecological validity and practical relevance; produces both usable artifacts and theory; responsive to the complexity of real classrooms and evolving AI tools; well-suited to developing a model and refining it based on authentic implementation evidence. - **Limitations:** weak internal validity (few/no control groups); findings are context-bound and hard to generalize; long timelines; difficult to isolate which design element caused an outcome — DBR demonstrates feasibility and improvement but cannot attribute learning gains to a specific mechanism. - **Exemplars:** [[ai-assisted-collaborative-learning-model-dbr]], [[genai-literacy-training-teacher-education-dbr-2026]], [[critical-thinking-genai-scaffolding]], [[human-centered-ai-teacher-educators-2026]]. DBR trades the causal control of [[rct|experiments]] for ecological authenticity and iterative refinement: it is the right tool for "how do we design this AI learning environment to work in practice?" questions, and its evidence is strongest as proof-of-concept and design guidance rather than causal efficacy. Reading DBR learning gains requires the same [[limitations-in-aied-research|caution]] as other designs — without an unassisted, controlled outcome measure, gains can reflect the same AI-inflated-performance confound documented under [[learning-gains|learning gains]]. ### Systematic reviews and meta-analyses Reviews synthesize the evidence base rather than running a new experiment. Systematic and scoping reviews apply a transparent protocol to search, screen, appraise, and synthesize a body of studies; meta-analyses additionally pool effect sizes across studies to produce a weighted summary estimate and test moderators. [[zerkouk-comprehensive-review-its-2025|A comprehensive ITS review]] and [[genai-higher-education-systematic-review-2026|a systematic review of GenAI in higher education]] exemplify the approach. - **Strengths:** efficient synthesis of a large, fragmented literature; meta-analysis yields pooled effect estimates and detects moderators; essential for evidence-based practice and identifying gaps. - **Limitations:** depend on the quality of included studies (garbage-in/garbage-out); publication bias; heterogeneous methods and outcome measures make synthesis hard; rapidly aging given the speed of AI change. - **Meta-research caveat (2026):** critiques of the AIED synthesis base show that many early meta-analyses are undermined by construct incoherence, unresolved heterogeneity, unaddressed dependence among effect sizes, and invalid publication-bias assessment — inflating headline AI effect sizes (see [[bartos-ai-learning-meta-meta-analysis-2026]], [[oneill-presumed-effective-meta-analysis-2026]], and [[weidlich-chatgpt-effect-search-cause-2025]]). Treat pooled AIED effect sizes as upper bounds. - **Exemplars:** [[zerkouk-comprehensive-review-its-2025]], [[genai-higher-education-systematic-review-2026]], [[chatgpt-critical-creative-thinking-review]], [[zerkouk-comprehensive-review-its-2025]], [[agentic-ai-education-scoping-review]]. See the dedicated [[meta-analysis-systematic-review]] concept page for a fuller treatment of systematic review and meta-analysis in AI in education — including their relationship to primary designs, PRISMA reporting, and their strengths and limitations. ### Computational and benchmark evaluation Computational evaluation assesses [[ai-technologies|AI systems]] directly — against benchmarks, ground-truth labels, or human judgments — rather than studying human learners. This includes [[benchmark|benchmarks]], [[llm]]-as-judge approaches. This is the closest method to [[ai-ed-evaluation]] (see the distinction below). - **Strengths:** fast, scalable, reproducible; enables head-to-head comparison of models and system versions; essential for system development and quality assurance. - **Limitations:** measures system output, not learning — high benchmark accuracy does not entail educational effectiveness; ground-truth and rubric quality are themselves contested; can miss pedagogical quality that humans perceive. [[rismanchian-ai-education-four-decades-aixed-2026|Rismanchian & Doroudi]] argue that LLMs' natural-language flexibility makes purely technical metrics insufficient, requiring human-inspired evaluation approaches — [[simulating-students|simulated students]], AI-[[teacher-role|teacher]] tests, and behavioral-science analyses previously reserved for human subjects — to judge learning-relevant quality, and that studying LLMs cautiously can generate insight into human learning. - **Exemplars:** [[teachbench-llm-teaching-evaluation]], [[jeon-isd-agent-bench-2026]], [[ground-truth-reliability-aied]], [[cong-confidence-asag-2026]], [[drawedumath-vlm-struggling-students-2026]]. - **Reporting standards for automated pipelines, and audited benchmarks.** Two 2026 papers extend methodological accountability beyond the study itself. PRISMA-LLM maps 888 review-automation papers and 14,726 annotations, finding that 38.0% of software or product papers reported no evaluation against 9.3% of LLM papers and that 52% of positive-only LLM evaluations left a high-bar concern unmet, and proposes reporting that identifies where in the review workflow automation acted ([[prisma-llm-ai-assisted-systematic-reviews-2026]]). An expert re-grading audit of six [[physics-education|physics]] benchmarks shows the same problem at the instrument level: of 250 audited rejections only 12 (4.80%) were genuine model errors, while 143 were item defects and 95 grader errors ([[frontier-models-physics-benchmark-audit-2026]]). Both argue that computational evaluations need an audited error budget before their results are read as findings about learners or models. ### Other designs: longitudinal, case, and simulation studies Beyond the major families, the knowledge base uses **longitudinal** designs that track learners over time ([[ai-lms-middle-school-longitudinal|a longitudinal LMS study]]), **case and in-the-wild** studies of authentic usage ([[ai-in-the-wild-college|large-scale analysis of real student interactions]]), and **simulation** studies in which LLMs stand in for students or patients ([[llm-student-simulation-teacher-insights|LLMs as simulated learners]], [[simulation]]). These trade breadth or control for realism and for access to phenomena that are otherwise hard to observe. ### Expert-consensus methods: the Delphi technique The Delphi method is a structured technique for establishing **expert consensus** on a question where the answer is not yet known empirically — most often used in the knowledge base to develop frameworks, competency lists, and definitions that practitioners and researchers can agree on. In a Delphi study, a panel of experts responds to successive rounds of questionnaires; after each round, an anonymized summary of the group's responses is fed back, and experts revise their answers until the group converges on agreement (typically defined by a pre-set threshold, e.g., 75%). It is a way to build construct validity and professional consensus through iterative, anonymized consultation rather than a single survey or vote. - **Strengths:** produces consensus from a diverse expert panel without in-person group pressures (anonymity reduces dominance effects); well-suited to defining constructs, competencies, and frameworks when no validated measure exists; iterative rounds let experts refine and converge; feasible where full experiments or large samples are impractical. - **Limitations:** consensus reflects expert judgment, not empirical evidence — it establishes agreement, not effect; results depend on panel [[writing-education|composition]] and the (subjective) consensus threshold; can be slow across multiple rounds; a single panel's judgment may not generalize. - **Exemplars:** [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|the SAIL framework study]] (three rounds, 17 experts, refining AI-literacy competency levels), [[hcap-human-centric-ai-pedagogy-framework-2026|the HCAP framework study]] (three rounds, 30 teachers, defining 25 AI-teacher competencies), [[ai-literacy-heptagon-2026|the AI Literacy Heptagon]] (which used expert input/consensus alongside a PRISMA-guided review), and. Delphi is often combined with other methods — for example, expert consensus can be used to validate a framework (as in SAIL and HCAP) that is then tested or implemented via design-based research or survey studies. It sits alongside qualitative and expert-judgment approaches and contributes to the [[educational-measurement|validity]] of framework-based instruments. ### Research vs. evaluation: connections and distinctions Research and evaluation are closely related but distinct. **Research** asks generalizable questions about how AI affects learning — "does scaffolding improve learning outcomes?" — and aims to build theory and evidence that transfers beyond the specific study. **Evaluation** (see [[ai-ed-evaluation]]) assesses whether a *specific* AI tool or system works — is accurate, reliable, pedagogically sound, and fit for purpose — against benchmarks, rubrics, or stakeholder-defined criteria. Research emphasizes internal validity and generalization; evaluation emphasizes system quality and local decision-making. The boundaries blur: benchmark studies are evaluation that can feed research, and evaluation instruments (rubrics, ground-truth sets, validity frameworks) depend on the [[educational-measurement]] and [[assessment-validity]] concerns that research clarifies. Conversely, research findings on what supports learning should inform how AI tools are [[ai-ed-evaluation|evaluated]]. The knowledge base treats them as complementary: computational and benchmark evaluation ([[benchmark]], [[ai-ed-evaluation]]) tells us whether an AI system is technically sound, while efficacy and survey research ([[rct]]) tells us whether it helps people learn. ### Choosing among methods Method choice follows the research question. Causal-effect questions favor experiments ([[rct]]); mechanism and perception questions favor surveys and qualitative work; system-quality questions favor computational evaluation ([[benchmark]], [[ai-ed-evaluation]]); synthesis questions favor reviews and meta-analyses; design questions favor DBR; and questions about what experts agree a construct, competency, or framework should contain favor expert-consensus methods like the Delphi technique. Given the field's heterogeneity and the speed of AI change, the knowledge base's corpus reflects a deliberate move toward triangulation — combining computational evaluation with efficacy, qualitative, and expert-consensus evidence to judge both whether a tool works and whether it helps learning. Equally important is reading any single study with awareness of the **cross-cutting limitations** that affect AIED research as a whole — methodological constraints, the fast pace of AI change versus slow publication, reproducibility and FAIR-practice gaps, reliance on proprietary tools, and weak or uncritical theory use. See [[limitations-in-aied-research]]. ## Contrasting the major research traditions The three major research traditions — [[quantitative-research|quantitative]], [[qualitative-research|qualitative]], and experimental — differ fundamentally in what they can claim, what they sacrifice, and when each is appropriate. Understanding these contrasts is essential for both designing and reading AI-in-education research. ### What each tradition establishes | Dimension | Quantitative / survey | Qualitative | Experimental | |---|---|---|---| | Core question | How much? How related? | What does it mean? How is it experienced? | Does X cause Y? | | Primary data | Numbers, scales, self-report | Words, observations, artifacts | Outcome measures across assigned conditions | | Inference target | Patterns, correlations, mediation | Meaning, mechanisms, categories | Causal effects | | Internal validity | Weak (correlational) | Weak (no control) | Strong (random assignment) | | External validity | Strong (large samples) | Limited (small, context-bound) | Moderate (controlled conditions) | | Ecological validity | Moderate | High | Lower (artificial conditions) | - **[[quantitative-research|Quantitative research]]** measures and models relationships among variables — surveys, SEM/PLS-SEM, measurement, longitudinal tracking. It provides breadth, precision, and generalizability but cannot establish causation from cross-sectional data and inherits [[educational-measurement|measurement]] limitations (including self-report bias). - **[[qualitative-research|Qualitative research]]** interprets meaning and experience — interviews, focus groups, thematic analysis, grounded theory, phenomenography, discourse analysis, observation/ethnography, case studies. It provides depth, mechanism, and theory-building (see [[theory-development-aied]]) but limited generalizability and weak causal support. - **Experimental and quasi-experimental designs** (see [[rct]]) estimate causal effects via random assignment or matched comparison — the gold standard for internal validity, at the cost of cost, speed, and ecological validity. ### The measurement and mixed-methods links Quantitative work depends on [[educational-measurement]] — reliable, valid instruments for the constructs being studied. Qualitative work reveals the mechanisms and meanings those instruments may miss. **Experimental** work estimates whether an intervention *causes* the outcomes the instruments measure. The three are complementary layers: instruments quantify constructs, experiments establish causality, and qualitative work explains the *how and why* behind the numbers. [[mixed-methods-research|Mixed-methods designs]] intentionally combine quantitative and qualitative strands so their strengths offset each other's weaknesses — quantitative breadth plus qualitative depth, with triangulation increasing confidence. ### Usability and HCI research A distinct methodological strand — [[usability-research|usability and HCI research]] — evaluates how users interact with an AI system: its usability, usefulness, learnability, and user experience, using think-aloud protocols, structured user studies, interviews, and observation. It is the closest to [[ai-ed-evaluation]] and answers a *prerequisite* question: even a pedagogically sound tool fails if it is unusable. Usability research shares data-collection methods with qualitative research but aims at evaluating an artifact rather than interpreting meaning. ### Benefits and limitations across traditions - **Quantitative/survey:** benefits — large samples, broad coverage, tests complex mediators, efficient. Limitations — no causation, self-report bias, convenience sampling, instruments may measure the wrong construct. - **Qualitative:** benefits — deep insight, surfaces unexpected phenomena and harms, essential for theory-building, centers under-represented voices. Limitations — limited generalizability, researcher dependence, small samples, weak causal support, hard to synthesize. - **Experimental:** benefits — strongest causal inference, clean outcome measurement, effect-size estimation. Limitations — costly/slow, artificial conditions, fast-changing AI dates results, underpowered small samples, ethical constraints. - **Mixed-methods:** benefits — triangulation, breadth + depth, explains unexpected results. Limitations — complex, resource-intensive, integration can be shallow, inherits each strand's weaknesses. - **Usability/HCI:** benefits — identifies adoption barriers, actionable design guidance, fast and cheap. Limitations — does not establish learning effects, small samples, self-report satisfaction can mislead. In practice, AI-in-education research rarely falls cleanly into one tradition. The strongest evidence triangulates: a computational or usability evaluation establishes that a system works, an experiment establishes that it causes learning, quantitative instruments measure the constructs, and qualitative work reveals the mechanisms and meanings — together answering both *whether* a tool helps learning and *how and why*. ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[ai-ed-evaluation]] - [[rct]] - [[benchmark]] - [[meta-analysis-systematic-review]] - [[educational-measurement]] - [[assessment-validity]] - [[simulation]] - [[ai-education]] - [[higher-ed]] - [[limitations-in-aied-research]] - [[learning-gains]] - [[theory-development-aied]] — Theory Development in AI in Education - [[qualitative-research]] — Qualitative Research - [[quantitative-research]] — Quantitative Research - [[mixed-methods-research]] — Mixed-Methods Research - [[design-based-research]] — Design-Based Research - [[usability-research]] — Usability Research - [[self-report-measures]] - [[learning-sciences]] ## Connected Articles - [[access-not-enough-ai-tutoring-2026]] — Access is Not Enough: Human Support Improves Engagement with AI Tutoring - [[genai-can-harm-teaching-rct-2026]] — Generative AI Can Harm Teaching - [[genai-over-reliance-learning-2026]] — From Enhancement to Over-Reliance: A Mixed-Method Study - [[acceptance-ai-english-tools-2026]] — Acceptance of AI-Assisted English Language Learning Tools - [[hazra-safetutors-pedagogical-safety-2026]] — AI Tutor Safety and Pedagogical Harms - [[zerkouk-comprehensive-review-its-2025]] — Comprehensive Review of Intelligent Tutoring Systems - [[ai-assisted-collaborative-learning-model-dbr]] — Design-Based Research for an AI-Assisted Collaborative Learning Model - [[teachbench-llm-teaching-evaluation]] — TeachBench: Evaluating LLM Teaching Ability - [[ground-truth-reliability-aied]] — Modernizing Ground Truth: Four Shifts Toward Reliability and Validity - [[llm-student-simulation-teacher-insights]] — Can LLMs Effectively Simulate Human Learners? - [[raise-framework-ai-education-reporting-2026]] — RAISE: 30 items in ten domains for transparent reporting of AI-in-education studies (Allison 2026) - [[ai-lms-middle-school-longitudinal]] — AI-Integrated Learning Management System: A Longitudinal Study - [[ai-in-the-wild-college]] — AI in the Wild: Large Scale Analysis of Authentic Interactions - [[same-ai-different-pathways]] — Same AI, Different Pathways: Unpacking Mechanisms - [[tep-aied-model-reporting-2026]] — The TEP-AIED model for reporting AI-in-education research with rigor (Hwang, Xie, Wah & Gasevic 2026) - [[t2i-competence-paradox-2026]] — The Competence Paradox: Text-to-Image GenAI in Art and Design - [[rismanchian-ai-education-four-decades-aixed-2026]] - [[weidlich-chatgpt-effect-search-cause-2025]] — ChatGPT in Education: An Effect in Search of a Cause - [[bartos-ai-learning-meta-meta-analysis-2026]] — Meta-meta-analysis of AI effect on learning - [[oneill-presumed-effective-meta-analysis-2026]] — Presumed Effective: flawed AIED meta-analysis audit --- ## [Qualitative Research](https://edtechdev.github.io/aied/concepts/qualitative-research/) > **Qualitative research** — the family of empirical methods that study how people *experience, interpret, and make meaning* of phenomena, typically through words, observations, and artifacts rather than numbers. In [[ai-education|AI in education]], qualitative methods reveal *how* students and teachers actually experience [[generative-ai|AI tools]] — the meanings, tensions, harms, and mechanisms that standardized measures miss. Because AI-in-education is fast-moving and its effects are often mediated by context, perception, and contested constructs like [[trust]] and [[agency]], qualitative work is essential alongside [[quantitative-research|quantitative]] designs (see [[research-methods-aied]]). ## Questions to Consider - A survey tells you 40% of students distrust an [[intelligent-tutoring|AI tutor]]; a focus group tells you *why* they distrust it. Which number feels more actionable to you, and what does the 'why' add that the percentage cannot? - Qualitative findings are usually not statistically generalizable, yet they're often 'conceptually generalizable' — mechanisms and dynamics that transfer elsewhere. Before you read, what does it mean for a finding to be generalizable in concept but not in statistics? - The page cautions that human–LLM coding *agreement* is not the same as coding *quality* when human consensus isn't ground truth. Have you ever treated 'two reviewers agreed' as proof something was correct? When is agreement a signal of truth, and when is it just shared error? - If AI can now assist qualitative coding at scale, does that threaten the interpretive depth that makes qualitative research valuable, or merely automate its drudgery? How would you decide which is happening in a given study? - Qualitative work centers under-represented voices — ethnic-minority students, skeptical nonusers — that large surveys often miss. Think of an AI-in-education claim you've heard. Whose experience of it is probably *not* captured by the headline number? - Choose one contested construct you care about — trust, agency, or harm. Before reading, sketch how you'd study it with words and observations rather than numbers, and note what you'd lose by doing so. ## Introduction Qualitative research is not a single method but a family organized by *what* they study and *how* evidence is gathered and analyzed. What unites them is an emphasis on meaning-making, context, and depth over breadth and causal control. Qualitative findings are typically **not** generalizable in the statistical sense, but they are often *conceptually generalizable* — revealing mechanisms, categories, and dynamics that transfer to other settings. In the knowledge base's corpus, qualitative work is prominent for studying AI acceptance, trust, harm, [[teacher-role|teaching]] practice, and learning processes. ## Major qualitative approaches ### Thematic analysis Thematic analysis identifies, codes, and interprets patterns ("themes") across qualitative data — typically interview or focus-group transcripts, open-ended survey responses, or documents. It is the most widely used approach in the knowledge base's qualitative studies. Interviews and open-ended responses are themselves [[self-report-measures|self-report data]], so they share the limits on what can be claimed about behavior — see that page for where self-report evidence is strong and where it breaks down. [[fouad-bentley-trust-utility-gap-physics-2026|A study of the trust–utility gap in physics]] uses thematic analysis of student interview data to surface [[discipline-specific-aied|domain-specific]] skepticism and adoption preferences; [[genai-teacher-feedback-comparison|a comparison of GenAI vs. teacher feedback]] analyzes student perceptions of usefulness and trustworthiness; and [[ai-adult-learning-guidelines-dis2026|guidelines for adult AI learning]] derive design principles from thematic coding of expert and learner input. [[ai-changing-teaching-workflows|How AI changes teaching workflows]] relies on thematic analysis of educator accounts. ### Grounded theory Grounded theory builds a theory *from the data* rather than testing an a priori framework, using iterative coding (open → axial → selective) until theoretical saturation. It is ideal for constructing new theory about emergent AI-in-education phenomena. [[liu-tool-tutor-crutch-programming-2026|Liu et al.]] develop a grounded theory of *tool, tutor, or crutch* — a three-mode typology of how students cognitively [[scaffolding|scaffold]] or offload onto AI in [[cs-education|programming education]] — directly theorizing [[cognitive-offloading]]. [[favero-critical-ai-tutors-empower-enslave-2025|A grounded-theory study of critical AI tutors]] examines whether such tutors empower or enslave learners. [[genai-feedback-design-multisite-experiment|Human-centered GenAI feedback design]] uses grounded analysis across a multisite study. See also [[theory-development-aied]] for how such grounded theories feed the field's theory building. ### Phenomenology and phenomenography Phenomenological approaches study the *lived experience* of a phenomenon — what it is like to learn with AI — while phenomenography studies the *qualitatively different ways* people experience and understand a phenomenon (producing "categories of description"). [[absent-cognitive-baseline-2026|The Absent Cognitive Baseline]] draws on students' lived experiences of self-assessment under AI to theorize a structural gap; [[metacognitively-discordant-completion-genai-2026|a phenomenological study]] captures the experience of completing work with AI while knowingly not understanding it; and [[genai-runaway-object-math-higher-ed|a socio-cultural, interpretive study]] analyzes how GenAI becomes a "runaway object" in [[math-education|mathematics]] academic practice. ### Discourse analysis Discourse analysis examines how language-in-use constructs meaning, identities, and power — analyzing classroom talk, written text, or interactional sequences. [[nspa-neuro-symbolic-pedagogical-alignment-2026|NSPA]] conducts long-horizon classroom *discourse analysis* (here computationally assisted) to mitigate dialect bias in understanding classroom interaction; [[scaffolding-critical-engagement-genai-minority-students|a study of ethnic-minority preparatory students]] analyzes collaborative *discourse* in [[prompt-engineering]] tasks. Discourse analysis bridges qualitative interpretation with computational methods when combined with [[educational-nlp]]. ### Observations and ethnography Observation studies watch behavior in context; ethnography extends this to sustained, immersive study of a setting, often with the researcher as participant-observer. [[trio-ethnography-llm-programming-education|A trio-ethnography]] of LLM-supported programming education traces how students' interpretations evolve; [[zha-ai-literacy-biology-case-study|a classroom case study]] observes [[ai-literacy]] integration in [[biology-education|biology]]. Observational methods capture *actual* behavior (what learners do with AI) rather than reported behavior — complementing the self-report surveys that dominate [[educational-measurement|quantitative measurement]] of attitudes. ### Case studies A case study is an in-depth investigation of a bounded case (a course, an institution, a single learner) using multiple data sources. [[drummond-genai-business-schools-framework-2026|A business-school case study]] generates a student-informed teaching and learning framework for GenAI; [[zha-ai-literacy-biology-case-study|a biology case study]] documents AI-literacy integration. Case studies trade breadth for depth and are strong for theory generation and transferable insight rather than generalization. [[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]] offer a qualitative case study of 21 visually impaired undergraduates across three Palestinian universities, using thematic analysis of semi-structured interviews to show how GenAI functions as an [[assistive-technology|assistive technology]] for [[inclusive-learning|inclusive learning]] — an example of case-study research surfacing mechanisms (personalized adaptation, teacher augmentation, educational parity) that quantitative measures miss. ### Interviews and focus groups Semi-structured **interviews** and **focus groups** are the primary data-collection instruments across all the above approaches. They elicit rich, contextual accounts. The knowledge base's qualitative corpus is built substantially on interviews (e.g., [[genai-expertise-pathways-sysadmin|expertise pathways]]) and focus groups (e.g., [[t2i-competence-paradox-2026|the text-to-image competence paradox]], [[ai-adult-learning-guidelines-dis2026]]). Quality depends on careful question design, sampling for variation, and rigorous analysis. ## How qualitative research appears in the knowledge base - **Mechanism and process.** Qualitative work reveals *why* AI helps or harms. [[same-ai-different-pathways]] uses qualitative strands to unpack mechanisms of AI-[[sociocultural-learning|mediated learning]] across contexts - **Trust, agency, and identity.** Contested, subjective constructs are often best studied qualitatively. [[t2i-competence-paradox-2026]] surfaces how art-and-design students negotiate ease, risk, and creative identity; [[genai-runaway-object-math-higher-ed]] captures AI's role in reshaping academic identity and practice. - **Equity and under-represented voices.** Qualitative work centers perspectives often excluded from large surveys — [[scaffolding-critical-engagement-genai-minority-students|ethnic-minority students]], [[becker-chatgpt-typology-physics-2026|skeptical nonusers]]. This connects to [[equity-in-ai-education]]. - **Typology and taxonomy building.** [[becker-chatgpt-typology-physics-2026|A qualitative typology of ChatGPT adoption]] distinguishes pragmatic users from skeptical nonusers — categories that inform later survey instrument design. ## AI and qualitative analysis A distinctive recent development is using [[llm|LLMs]] to assist qualitative coding. The knowledge base's evidence is cautionary: [[human-vs-llm-ordered-coding]] shows LLM and human coding diverge, with errors cascading through temporal analysis; [[agreement-not-quality-llm-coding-verification|Agreement Is Not Quality]] shows that human–LLM coding *agreement* is not the same as coding *quality* when human consensus is not ground truth. LLM-assisted coding can scale and accelerate qualitative analysis, but its outputs require verification against [[human-in-the-loop-ai|human judgment]] — an important intersection of qualitative research with [[educational-nlp]] and [[ai-ed-evaluation]]. **The fluent-output problem and warrantability.** [[chain-behind-claim-warrantability-2026|Holster (2026)]] sharpens why accuracy and disclosure are insufficient standards for AI-assisted qualitative work. Because LLMs reorganize corpora in minutes into fluent topics, quotations, and prevalence claims, they can conceal the analytic pathway that produced them — and people tend to rate easily processed, fluent output as more true. Holster proposes **warrantability** as a complement to accuracy and disclosure: an AI-assisted interpretation is warrantable when the pathway from source data to claim remains *inspectable, contestable, and revisable*. Its constructive machinery is **semantic lenses** (documented reorganizations of a corpus across levels of abstraction) and a claim-relative repertoire of **warrant artifacts** — source-linked topic tables, lens stacks, and evidence rivers — designed into tools so that fluent analysis also produces a retraceable pathway record. This extends the field's audit-trail tradition into the generative era and gives reviewers something concrete to contest beyond a final interpretation. - **Reliability is not accuracy in [[agentic-ai|multi-agent]] LLM coding.** A literature-informed pipeline in which two AI coders coded independently, debated and reconciled disagreements produced inter-coder agreement above Cohen's kappa 0.85 on every dataset and label while criterion F1 ranged from 0.31 to 0.89 (mean 0.68, SD 0.16) — high agreement between agents that were both wrong. Longer codebooks (t = -11.702) and more similar excerpts (t = -9.249) reduced initial accuracy, and discussion turns correlated positively with accuracy (t = 7.997) while correctly resolved conflicts and collaborating modes correlated negatively (t = -9.720 and -8.420), so convergence is a weak proxy for correctness. The practical implication is that AI-assisted coding needs checklist-level adjudication against a human reference rather than agent-to-agent agreement as its quality signal. ([[llm-qualitative-coding-consensus-2026]]) ## Strengths and limitations - **Strengths:** deep ecological and conceptual insight; surfaces unexpected phenomena, risks, and mechanisms; essential for theory-building (see [[theory-development-aied]]); captures meaning, context, and contested constructs; centers under-represented perspectives; strong for studying fast-moving phenomena where standardized measures lag. - **Limitations:** limited statistical generalizability; interpretive and researcher-dependent (reliability concerns); small samples; weaker support for causal claims; findings can be hard to synthesize across studies; time- and labor-intensive. Qualitative and quantitative methods are complements, not rivals — see [[research-methods-aied]] for how they contrast and triangulate, and [[mixed-methods-research|mixed methods]] for designs that combine them. ## Connected Concepts - [[research-methods-aied]] - [[mixed-methods-research]] - [[quantitative-research]] - [[theory-development-aied]] - [[educational-measurement]] - [[educational-nlp]] - [[ai-ed-evaluation]] - [[equity-in-ai-education]] - [[trust]] - [[agency]] - [[cognitive-offloading]] - [[self-report-measures]] ## Connected Articles - [[chain-behind-claim-warrantability-2026]] — warrantability standard for AI-assisted qualitative analysis - [[liu-tool-tutor-crutch-programming-2026]] — Tool, Tutor, or Crutch: a grounded theory of AI-assisted programming - [[trio-ethnography-llm-programming-education]] — A trio-ethnography of interpretation evolution in LLM-supported programming - [[absent-cognitive-baseline-2026]] — Theorizing a structural gap in AI-native students' self-assessment - [[t2i-competence-paradox-2026]] — The competence paradox in text-to-image GenAI use - [[fouad-bentley-trust-utility-gap-physics-2026]] — Trust–utility gap in physics education - [[genai-teacher-feedback-comparison]] — Comparing GenAI and teacher feedback: student perceptions - [[hazra-safetutors-pedagogical-safety-2026]] — AI tutor safety and pedagogical harms - [[zha-ai-literacy-biology-case-study]] — Case study of AI-literacy integration in a biology class - [[becker-chatgpt-typology-physics-2026]] — A qualitative typology of ChatGPT adoption in physics - [[scaffolding-critical-engagement-genai-minority-students]] — Collaborative discourse in prompt engineering among ethnic-minority students - [[human-vs-llm-ordered-coding]] — Comparing human and LLM ordered coding of qualitative data - [[agreement-not-quality-llm-coding-verification]] — Agreement is not quality in LLM qualitative coding - [[same-ai-different-pathways]] — Unpacking mechanisms of AI-mediated learning across contexts - [[drummond-genai-business-schools-framework-2026]] — Student-informed GenAI framework via case study - [[favero-critical-ai-tutors-empower-enslave-2025]] — Critical AI tutors: empower or enslave - [[genai-runaway-object-math-higher-ed]] — GenAI as a runaway object in higher-education mathematics - [[khlaif-assistive-genai-visually-impaired-2026]] — Assistive GenAI for visually impaired learners --- ## [Quantitative Research](https://edtechdev.github.io/aied/concepts/quantitative-research/) > **Quantitative research** — the family of empirical methods that collect and analyze *numerical* data to describe patterns, test relationships, and estimate causal effects. In [[ai-education|AI in education]], quantitative methods quantify whether and how AI tools affect [[learning-gains|learning outcomes]], [[student-engagement|engagement]], [[motivation]], and [[self-efficacy]], and model the psychological and behavioral mechanisms of AI use. They provide the breadth, precision, and causal inferential power that [[qualitative-research|qualitative methods]] trade away for depth and context. ## Questions to Consider - 'Students who use an AI tutor score higher' — before you read, is this a claim about causation or correlation, and what single piece of evidence would convert one into the other? - A cross-sectional survey can test a complex mediational model yet never establish causation. Why might two variables be correlated in a survey even when neither causes the other? Can you think of a way to be genuinely misled by such a correlation in education? - Quantitative instruments are only as good as what they measure — and the page notes instruments can measure the wrong construct. When you fill out a self-report survey about 'confidence' or 'engagement,' what could it actually be capturing instead, and how would you find out? - An RCT randomly assigns learners to conditions to estimate a causal effect. What makes random assignment powerful, and what practical and ethical problems arise when the 'treatment' is a possibly-helpful AI tool that some students are being denied? - Longitudinal designs track the same learners over time — essential for distinguishing AI-inflated performance from durable learning. Why would a single snapshot of high scores fail to reveal whether learning actually happened? - Quantitative work provides breadth and causal power; qualitative work provides depth and meaning. Before reading the pairing, where do you think numbers alone have most likely misled you about a learning claim, and what method would you add to check it? ## Introduction Quantitative research spans descriptive designs (measuring prevalence and patterns), correlational/observational designs (testing relationships among variables), and experimental and quasi-experimental designs (estimating causal effects). What unifies them is the systematic reduction of observations to numbers, analyzed with statistics, and the priority placed on **reliability, validity, and generalizability** — the core concerns of [[educational-measurement]]. ## Major quantitative approaches ### Survey and correlational research Cross-sectional surveys measure self-reported attitudes, perceptions, motivation, [[self-efficacy]], and technology acceptance, often modeled with regression or structural equation modeling (SEM/PLS-SEM) to test hypothesized relationships and mediators. These dominate the knowledge base's corpus, particularly for acceptance, motivation, and psychological-mechanism questions. [[acceptance-ai-english-tools-2026|Acceptance of AI-assisted English tools]] builds on the [[technology-acceptance-model|TAM]] with SEM; [[tian-genai-learning-adoption-pathways-2026|GenAI adoption pathways]] uses PLS-SEM, fsQCA, and importance-performance mapping; [[teacher-education-ai-literacy-sdt-2026|teacher AI literacy]] uses factor-validated surveys grounded in [[self-determination-theory]]. - **Strengths:** large samples; broad, low-cost coverage; tests complex mediational models; feasible for attitudes that are hard to observe. - **Limitations:** cross-sectional data cannot establish causation; self-report bias; convenience sampling limits generalizability; mediators inferred from covariance, not manipulation. See [[self-report-measures]] for the instrument-side treatment of these limits. ### Experimental and quasi-experimental research Experiments randomly assign learners to conditions (e.g., AI tutor vs. human tutor, or AI-scaffolded vs. unassisted) to estimate causal effects on outcomes. **Randomized controlled trials ([[rct]]s)** are the gold standard for internal validity. [[access-not-enough-ai-tutoring-2026|A randomized field study of human support plus AI tutoring]] and [[genai-can-harm-teaching-rct-2026|an RCT on generative AI in teaching]] use assignment to isolate causal effects. **Quasi-experimental** designs (pre/post, matched groups without randomization) are more feasible in intact classrooms but weaker on causal claims. - **Strengths:** strongest causal inference; clean outcome measurement; supports effect-size estimation and efficacy claims. - **Limitations:** costly and slow; artificial conditions reduce ecological validity; fast-changing AI tools date experiments quickly; small samples underpower detection of effects; ethical constraints on withholding helpful tools. ### Longitudinal research Longitudinal designs track the same learners over time, capturing change, growth, and durable learning that single-time-point measurement misses. [[ai-lms-middle-school-longitudinal|A longitudinal LMS study]] tracks students across a school year. Longitudinal designs are essential for distinguishing AI-inflated performance from [[genai-performance-vs-learning|durable learning]]. ### Computational and psychometric quantification Quantitative methods also include the direct measurement of constructs via instruments — the domain of [[educational-measurement]] and [[item-response-theory]]. The knowledge base's [[jin-glat-genai-literacy-assessment|GLAT]] is a 20-item quantitative instrument validated with IRT; [[educational-measurement|measurement instruments]] across [[ai-literacy|AI literacy]], acceptance, and self-efficacy provide the validated scales on which survey and experimental research depend. ## How quantitative research appears in the knowledge base - **Efficacy and causal claims.** RCTs and quasi-experiments test whether AI tools improve learning ([[access-not-enough-ai-tutoring-2026]], [[genai-can-harm-teaching-rct-2026]], [[adaptive-pretesting-retention]]). - **Mechanism modeling.** SEM/PLS-SEM tests mediators and moderators of AI adoption and learning ([[tian-genai-learning-adoption-pathways-2026]], [[acceptance-ai-english-tools-2026]], [[teacher-education-ai-literacy-sdt-2026]]). - **Measurement and scale development.** The knowledge base documents quantitative instrument development and validation ([[jin-glat-genai-literacy-assessment|GLAT]], [[educational-measurement]]). ## Strengths and limitations - **Strengths:** precision and statistical power; generalizability to defined populations; causal inference (with experimental designs); efficient large-sample coverage; cumulative and comparable across studies. - **Limitations:** captures what is measurable, often missing process, meaning, and context (see [[qualitative-research]]); self-report bias; instruments can measure the wrong construct (see [[educational-measurement|measurement issues]]); correlation without causation; can be artificial and slow relative to AI change. Quantitative and [[qualitative-research|qualitative]] methods are complements — quantitative work provides breadth and causal power, qualitative work provides depth and meaning. [[mixed-methods-research|Mixed-methods designs]] combine them. See [[research-methods-aied]] for the full method comparison and the contrasts among experimental, survey, qualitative, and other designs. ## Connected Concepts - [[research-methods-aied]] - [[qualitative-research]] - [[mixed-methods-research]] - [[educational-measurement]] - [[item-response-theory]] - [[rct]] - [[learning-gains]] - [[student-engagement]] - [[self-efficacy]] - [[technology-acceptance-model]] - [[self-report-measures]] ## Connected Articles - [[access-not-enough-ai-tutoring-2026]] — A randomized field study of human support plus AI tutoring - [[genai-can-harm-teaching-rct-2026]] — Generative AI can harm teaching: an RCT - [[acceptance-ai-english-tools-2026]] — Acceptance of AI-assisted English learning tools - [[tian-genai-learning-adoption-pathways-2026]] — GenAI adoption pathways (PLS-SEM, fsQCA) - [[teacher-education-ai-literacy-sdt-2026]] — Teacher AI literacy through self-determination theory - [[jin-glat-genai-literacy-assessment]] — GLAT: an IRT-validated GenAI literacy test - [[ai-lms-middle-school-longitudinal]] — A longitudinal AI-integrated LMS study - [[genai-over-reliance-learning-2026]] — From enhancement to over-reliance (mixed-method) - [[adaptive-pretesting-retention]] — Adaptive pretesting and retention --- ## [Mixed-Methods Research](https://edtechdev.github.io/aied/concepts/mixed-methods-research/) > **Mixed-methods research** — the design and practice of intentionally combining [[quantitative-research|quantitative]] and [[qualitative-research|qualitative]] strands within a single study so that their strengths complement and their weaknesses offset each other. In [[ai-education|AI in education]], mixed-methods designs are widely used because AI effects are simultaneously measurable ([[learning-gains|learning gains]], [[student-engagement|engagement]]) and meaning-laden (trust, [[agency]], identity) — and neither a survey nor an interview alone captures both. ## Questions to Consider - Suppose a survey shows students report high satisfaction with an AI tool, but test scores don't improve. Which number is 'right'? The page argues that combining quantitative and qualitative strands lets each explain the other — how might an interview explain what the numbers alone cannot? - A common view is that numbers are objective and interviews are anecdotal, so quantitative evidence is superior. The page instead treats the two as complementary — breadth and precision from numbers, depth and meaning from words. When does a number without context mislead more than an interview? - One design pattern (sequential explanatory) collects quantitative data first, then uses interviews to explain surprising results. Can you think of an educational outcome that was puzzling numerically until qualitative accounts revealed the mechanism? - Triangulation is the core value: when independent strands converge, confidence rises, and when they diverge, the discrepancy itself is informative. Have you ever had two sources of evidence about the same thing contradict each other — and what did you learn from the conflict? - Mixed methods can simultaneously capture effects (how much learning gain) and meaning (trust, identity, agency) — things no single survey or interview captures alone. If you were studying whether an AI tool helps students, what quantitative and qualitative questions would you pair, and in what order? ## Introduction Mixed-methods designs integrate the breadth, precision, and causal power of quantitative methods with the depth, context, and meaning of qualitative methods. The value proposition is **triangulation**: when independent strands converge on the same conclusion, confidence increases; when they diverge, the discrepancy itself is informative. ## Common design patterns - **Sequential explanatory (QUAN → qual).** Quantitative data are collected first, then qualitative data explain or contextualize surprising or significant quantitative findings. [[genai-over-reliance-learning-2026|A mixed-method study of GenAI and sustainable learning]] pairs three-wave surveys with educator interviews to explain the quantitative pattern of enhancement-to-over-reliance. - **Sequential exploratory (QUAL → quan).** Qualitative work builds theory, generates hypotheses, or informs instrument design that is then tested quantitatively. [[becker-chatgpt-typology-physics-2026|A qualitative typology of ChatGPT adoption]] yields categories that can inform later survey design. - **Convergent/parallel.** Quantitative and qualitative strands run simultaneously and are integrated in analysis. [[t2i-competence-paradox-2026|The competence-paradox study]] uses instructor focus groups, a student survey, and follow-up interviews in parallel; [[fouad-bentley-trust-utility-gap-physics-2026|the trust–utility gap study]] combines survey and interview evidence on physics students' AI adoption. ## How mixed methods appear in the knowledge base - **Explaining mechanisms.** [[same-ai-different-pathways]] combines strands to unpack the mechanisms of AI-mediated learning across discipline-institution contexts, where quantitative differences alone would be opaque. - **Complementing outcome data with experience.** [[hazra-safetutors-pedagogical-safety-2026|AI tutor safety]] pairs quantitative harm indicators with qualitative accounts of pedagogical harm; [[t2i-competence-paradox-2026]] pairs quantitative survey results with qualitative negotiation-of-identity accounts. - **Design and evaluation.** [[genai-feedback-design-multisite-experiment|A multisite experiment on GenAI feedback design]] combines experimental outcome measurement with qualitative feedback from learners, integrating quantitative effect estimation with design guidance. ## Strengths and limitations - **Strengths:** triangulation increases confidence; quantitative breadth plus qualitative depth; can explain unexpected results and bridge mechanism and magnitude; produces both effects and meaning; well-suited to complex, contextual AI-in-education phenomena. - **Limitations:** complex, resource-intensive, and methodologically demanding; integration can be shallow if strands are merely reported side-by-side rather than genuinely merged; still inherits the weaknesses of each strand (e.g., self-report bias in surveys, researcher dependence in interviews); requires proficiency in both methodological traditions. ## Relationship to the broader methods landscape Mixed-methods sits between the [[quantitative-research|quantitative]] and [[qualitative-research|qualitative]] traditions, drawing on the strengths of each while adding the design discipline of intentional integration. It is a methodological response to the recognition that AI-in-education phenomena are both effectful and meaning-laden — and that the field advances fastest when breadth and depth are combined (see [[research-methods-aied]]). It connects to [[educational-measurement]] (quantitative instruments), [[theory-development-aied]] (qualitative theory-building), and [[ai-ed-evaluation]] (integrating outcome and experience evidence). ## Connected Concepts - [[research-methods-aied]] - [[qualitative-research]] - [[quantitative-research]] - [[educational-measurement]] - [[learning-gains]] - [[theory-development-aied]] - [[ai-ed-evaluation]] - [[student-experience]] ## Connected Articles - [[genai-over-reliance-learning-2026]] — From enhancement to over-reliance (surveys + interviews) - [[t2i-competence-paradox-2026]] — The competence paradox (focus groups + survey + interviews) - [[fouad-bentley-trust-utility-gap-physics-2026]] — Trust–utility gap (survey + interviews) - [[same-ai-different-pathways]] — Mechanisms of AI-mediated learning across contexts - [[genai-feedback-design-multisite-experiment]] — Human-centered GenAI feedback design (multisite) - [[hazra-safetutors-pedagogical-safety-2026]] — AI tutor safety and pedagogical harms - [[becker-chatgpt-typology-physics-2026]] — A qualitative typology of ChatGPT adoption in physics --- ## [Design-Based Research](https://edtechdev.github.io/aied/concepts/design-based-research/) > **Design-based research (DBR)** — a methodological approach that iteratively designs, implements, and refines an educational intervention in authentic contexts, cycling between theory, design, and real-world practice to produce both a usable artifact and validated design principles. In [[ai-education|AI in education]], DBR is the method of choice for developing AI learning environments, [[pedagogy|pedagogical]] models, and teacher-training programs that must work in the messy reality of classrooms — trading the causal control of [[rct|experiments]] for ecological authenticity and iterative refinement. ## Questions to Consider - Design-based research iteratively designs, implements, and refines an intervention in real classrooms, cycling between theory and practice. How is that fundamentally different from a controlled experiment — and why might a classroom researcher prefer it? - DBR deliberately trades causal control for ecological authenticity. If you can't say for sure WHAT caused a learning gain, what is the knowledge still worth? - DBR produces two things at once: a usable artifact and validated design principles. Which of those two outputs matters more to you — a tool that works, or a principle you can apply elsewhere? - A strength of DBR is high practical relevance; a limitation is that findings are context-bound and hard to generalize. When would you trust a DBR finding enough to apply it in a very different setting? - Without an unassisted, controlled outcome measure, DBR learning gains can reflect the same 'AI-inflated performance' problem it studies. How would you design a DBR study so its reported gains actually mean learning, not assisted performance? ## Introduction DBR is a *design* tradition, not a data tradition: it does not fit neatly into the [[quantitative-research|quantitative]]/[[qualitative-research|qualitative]]/experimental contrast. It deliberately combines elements of all three — collecting both outcome and process data across iterative cycles — to answer *"how do we design this AI learning environment to work in practice?"* rather than *"does X cause Y?"*. It is closely related to [[learning-design]] (which specifies the design process) and to [[usability-research|usability evaluation]] (which feeds refinement), but is distinguished by its sustained, theory-driven, multi-cycle character and its dual goal of improving practice *and* generating theory. ## The DBR cycle DBR is not a single design but an iterative loop, typically comprising: 1. **Needs analysis** — understanding the problem, context, and learners in authentic settings. 2. **Design** — articulating an intervention and its underlying theoretical rationale. 3. **Implementation** — deploying the intervention in a real classroom or program. 4. **Refinement** — analyzing data, revising the design, and iterating (often over multiple cycles). 5. **Theory and artifact output** — producing both a usable intervention and generalizable design principles or a validated model. The canonical AIEd example is the [[ai-assisted-collaborative-learning-model-dbr|AI-Assisted Collaborative Learning (AACL) Model study]], which ran a four-phase DBR cycle — needs analysis, model design, an eight-week classroom implementation with Indonesian undergraduates, and model refinement — iterating on a four-stage learning cycle (problem identification → AI-assisted [[collaborative-learning|collaborative]] inquiry → collaborative [[problem-solving]] → reflection and presentation). ## How DBR appears in the knowledge base - **Developing learning models.** [[ai-assisted-collaborative-learning-model-dbr|The AACL Model study]] uses DBR to develop and evaluate an AI-assisted collaborative learning model targeting [[critical-thinking|critical thinking]] and problem-solving in [[higher-ed|higher education]]. - **Building AI-literacy teacher training.** [[genai-literacy-training-teacher-education-dbr-2026|Le et al.]] develop and evaluate a DBR [[generative-ai|GenAI]]-literacy training intervention for [[teacher-education]] students; [[human-centered-ai-teacher-educators-2026|Baran et al.]] use DBR across 2023–2025 to design professional learning for critical [[ai-literacy|AI literacy]] grounded in Human-Centered AI principles. - **Designing [[governance|institutional]] standards.** [[crompton-faculty-technology-integration-standards-2026|Crompton et al.]] use DBR across two iterative macro cycles and 114 participants to develop six faculty technology-integration standards. - **Iterative system implementation.** [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of|Rienties et al.]] describe six iterative DBR studies (18 months, 498 participants) implementing the Open University's AIDA AI assistant using an embedded-systems approach. - **[[scaffolding]] and intervention design.** [[critical-thinking-genai-scaffolding|GenAI critical-thinking scaffolding]] and other intervention-development studies use DBR to design and refine AI-based scaffolds. - **Youth- and expert-guided [[curriculum-design|curriculum design]].** [[science-integrated-ai-literacy-curriculum-dbr-2026|Moore et al. (2026)]] use a two-year DBR process with a youth and AI-expert advisory board to design a science-integrated [[ai-literacy|AI literacy]] / ML curriculum for high school, refining the design across cohorts and measuring ML-knowledge gains — an example of DBR as participatory co-design that integrates youth voice into the design cycle. - **Iterative rubric co-refinement as DBR.** [[yasar-llms-iterative-pedagogical-design-2026|Yaşar et al. (2026)]] exemplify DBR's iterative-refinement character applied to the assessment instrument itself: they treated the rubric as a revisable design artifact and cycled through co-refinement with an [[llm|LLM]] — clarifying performance descriptors and explicitly accepting implicit indicators of learning — which raised LLM–human agreement on 80 student design posters from 54.75% to 81.25% (Cronbach's Alpha 0.393 → 0.798). The study positions the rubric as a mediating interface between human pedagogical intent and machine inference, and its role-aware [[prompt-engineering|prompting]] (instructor, peer-reviewer, grant-reviewer) shows how design choices shape evaluative output — a DBR-style demonstration that the assessment instrument, not just the intervention, is a design object. - **Co-designing early-childhood AI literacy.** [[play-ai-pre-k-kindergarten-ai-literacy-2026|Lee (2026)]] uses DBR to co-design, pilot, and refine the Play With AI (PL-AI) curriculum across iterative cycles with two pre-K and two kindergarten teachers, drawing on teacher surveys, 32 hours of classroom video, field notes, and design-meeting transcripts. The study documents how DBR supports iterative refinement of developmentally appropriate AI literacy activities and yields transferable design principles ([[embodied-learning|embodied]] play, tangible coding, guided dialogue, teacher co-design). ## Strengths and limitations - **Strengths:** high ecological validity and practical relevance; produces both usable artifacts and theory; responsive to the complexity of real classrooms and evolving AI tools; well-suited to developing a model and refining it based on authentic implementation evidence; captures how an intervention actually works (or fails) in practice. - **Limitations:** weak internal validity (few/no control groups); findings are context-bound and hard to generalize; long timelines; difficult to isolate which design element caused an outcome — DBR demonstrates feasibility and improvement but cannot attribute learning gains to a specific mechanism. DBR trades the causal control of [[rct|experiments]] for ecological authenticity and iterative refinement: its evidence is strongest as proof-of-concept and design guidance rather than causal efficacy. Reading DBR learning gains requires the same [[limitations-in-aied-research|caution]] as other designs — without an unassisted, controlled outcome measure, gains can reflect the same AI-inflated-performance confound documented under [[learning-gains|learning gains]]. ## Relationship to other methods DBR is the signature method of the [[learning-sciences|learning sciences]] — the interdisciplinary field that studies learning and the design of learning environments, and that treats building an intervention and studying it as a single activity rather than two. It is complementary to, not a rival of, other research methods (see [[research-methods-aied]] for the full landscape). Where [[rct|experiments]] establish causality and [[quantitative-research|surveys]] establish breadth, DBR establishes *feasibility and design knowledge* — whether an intervention can be built to work in authentic practice and what design principles support it. It frequently pairs with [[usability-research|usability evaluation]] (to refine the interface) and [[qualitative-research|qualitative methods]] (to understand how learners experience the intervention). A mature DBR program typically culminates in an [[rct|efficacy trial]] or [[educational-measurement|measurement]] study that tests the developed intervention's causal effects at scale. ## Connected Concepts - [[research-methods-aied]] - [[rct]] - [[quantitative-research]] - [[qualitative-research]] - [[mixed-methods-research]] - [[usability-research]] - [[learning-design]] - [[educational-measurement]] - [[learning-gains]] - [[limitations-in-aied-research]] - [[theory-development-aied]] - [[learning-sciences]] ## Connected Articles - [[ai-assisted-collaborative-learning-model-dbr]] — DBR for an AI-Assisted Collaborative Learning Model - [[genai-literacy-training-teacher-education-dbr-2026]] — DBR for AI-literacy training in teacher education - [[crompton-faculty-technology-integration-standards-2026]] — DBR for faculty technology-integration standards - [[human-centered-ai-teacher-educators-2026]] — Human-centered AI for teacher educators (DBR) - [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of]] — Six DBR studies implementing AIDA at the Open University - [[critical-thinking-genai-scaffolding]] — DBR for GenAI critical-thinking scaffolding - [[science-integrated-ai-literacy-curriculum-dbr-2026]] — DBR for a science-integrated AI literacy curriculum (Moore et al. 2026) - [[play-ai-pre-k-kindergarten-ai-literacy-2026]] — Play With AI (PL-AI): play-centered AI literacy curriculum for pre-K and kindergarten (Lee 2026) - [[yasar-llms-iterative-pedagogical-design-2026]] — LLMs as agents of iterative pedagogical design --- ## [Usability Research](https://edtechdev.github.io/aied/concepts/usability-research/) > **Usability research** — the empirical study of how users interact with a software system, and of its usability, usefulness, and user experience (UX). Drawn from human–computer interaction (HCI), usability research evaluates whether an AI educational tool is usable, learnable, efficient, and satisfying — the qualities that determine whether learners actually adopt and benefit from it. It is distinct from, but complementary to, [[qualitative-research|qualitative inquiry]] into learning phenomena and [[quantitative-research|quantitative efficacy]]: usability research focuses on the *interaction between person and system*, not on learning outcomes per se. ## Questions to Consider - Can you recall a software tool — educational or otherwise — that was pedagogically sound or genuinely powerful, yet you stopped using it because it was confusing or frustrating? That failure is exactly what usability research tries to explain before it costs learning. - A common assumption is that a tool's educational effectiveness can be judged by whether learning gains improve. But the page argues a tool can be unusable and still seem to 'work' in a trial, or usable yet fail to teach. Why might a study showing learning gains still miss that the tool is a pain to use in practice? - Before reading the methods, how would you go about finding out whether an AI tutor is confusing, frustrating, or error-prone for learners? What would you actually do or observe — and what would a user's own report of satisfaction miss that careful observation would catch? - Think-aloud is a core method: users speak their thoughts while working, exposing confusion and mental models in real time. If you were a learner using an AI tool, what would you be able to articulate about your confusion that a simple 'did you like it?' survey would never capture? - The page notes self-reported satisfaction can diverge from objective performance — people can say they love a tool that secretly slows them down, or underrate one that actually helps. Where have you seen that gap between what people say and what their behavior shows? - Usability research tells you whether a tool is usable, not whether it teaches. If you're evaluating an AI learning tool, how would you combine usability evidence with evidence of learning — and what could a tool that passes both still fail to achieve? ## Introduction Usability and UX research answer questions like: Can students figure out how to use this AI tutor? Is the AI tool confusing, frustrating, or error-prone? Does it fit the workflow of teachers or learners? These questions are a prerequisite for — and sometimes the hidden cause of — the [[learning-gains|learning gains]] (or lack thereof) measured in efficacy studies. An AI tool that is pedagogically sound but unusable will fail in practice; usability evidence explains why. ## Core methods - **Think-aloud protocols.** Users verbalize their thoughts while performing tasks, revealing comprehension, confusion, and mental models in real time. [[code-anchor-multi-view-visualization|A study of multi-view code visualizations]] and [[learn-framework-responsible-genai-pbl-2026|the LEARN framework]] use think-aloud to understand how learners make sense of AI-assisted tools; [[feedback-futures-genai|feedback futures]] examines how learners process AI-generated feedback. - **User studies.** Structured task-based evaluations measure efficiency, error rates, satisfaction, and completion. [[rhaimi-productivemath-2025|ProductiveMath]] evaluates a generative-AI app's usability in supporting productive-failure teaching; [[supplynet-visual-exploratory-learning|SupplyNet]] runs a user study of a visual exploratory learning tool; [[llm-chatbots-cs-multiple-choice|LLM chatbots for CS multiple-choice]] assess interaction quality. - **Interviews and observation.** Qualitative usability interviews and observation capture user experience, preferences, and pain points. [[icub-humanoid-storytelling-llm-hri-2025|A usability study of a storytelling humanoid robot]] uses structured evaluation to ask whether parents would let the robot interact with a child; [[genai-architectural-design-studios|AI in design studios]] observes and interviews students using AI in authentic design work. - **Systematic usability evaluation.** Heuristic evaluation, cognitive walkthrough, and questionnaire-based UX measures (e.g., SUS) systematically assess usability against established criteria. ## How usability research appears in the knowledge base - **AI learning tool evaluation.** [[rhaimi-productivemath-2025|ProductiveMath]], [[supplynet-visual-exploratory-learning|SupplyNet]], and [[anvil-ai-educational-animations|educational animations]] are evaluated for usability and UX. - **Human–robot and [[conversational-ai|conversational AI]] interaction.** [[icub-humanoid-storytelling-llm-hri-2025|The humanoid storytelling study]] is an explicit usability study of [[llm]]-powered interaction; [[conversational-ai-agents-umbrella-review-2026|an umbrella review of conversational AI agents]] identifies usability and interaction quality as a recurring theme. - **Design and refinement.** Usability findings feed iterative design (see [[design-thinking]] and [[learning-design]]), improving tools before or alongside efficacy testing. ## Relationship to other research families Usability research shares data-collection methods with [[qualitative-research|qualitative research]] (interviews, observation, think-aloud) but differs in *aim*: qualitative research interprets meaning and experience to build understanding and theory, whereas usability research evaluates an artifact against usability/UX criteria. It also overlaps with [[ai-ed-evaluation]] (assessing whether a system works) and with [[educational-measurement|measurement]] (quantifying usability constructs). The knowledge base treats usability as a distinct but connected methodological strand — relevant to [[human-ai-collaboration]], [[student-experience]], and the design of effective AI learning tools. See [[research-methods-aied]] for how it fits the broader methods landscape. ## Strengths and limitations - **Strengths:** directly identifies usability barriers that block adoption and learning; produces actionable design guidance; complements efficacy and qualitative research by explaining *why* a tool works or fails in use; relatively fast and cheap compared to large experiments. - **Limitations:** usability findings do not establish learning effects (a usable tool can still fail to teach); small samples and task-specific settings limit generalizability; self-report satisfaction can diverge from objective performance; researcher and task-design dependence. ## Connected Concepts - [[research-methods-aied]] - [[qualitative-research]] - [[human-ai-collaboration]] - [[student-experience]] - [[ai-ed-evaluation]] - [[learning-design]] - [[design-thinking]] - [[intelligent-tutoring]] ## Connected Articles - [[icub-humanoid-storytelling-llm-hri-2025]] — A usability study of an LLM-powered storytelling humanoid - [[rhaimi-productivemath-2025]] — ProductiveMath: usability of a generative-AI app - [[supplynet-visual-exploratory-learning]] — SupplyNet user study - [[anvil-ai-educational-animations]] — Usability of AI-generated educational animations - [[code-anchor-multi-view-visualization]] — Think-aloud study of multi-view code visualizations - [[learn-framework-responsible-genai-pbl-2026]] — LEARN framework and think-aloud evaluation - [[feedback-futures-genai]] — How learners process AI-generated feedback - [[llm-chatbots-cs-multiple-choice]] — LLM chatbots for CS multiple-choice questions - [[conversational-ai-agents-umbrella-review-2026]] — Umbrella review of conversational AI agents --- ## [RCT](https://edtechdev.github.io/aied/concepts/rct/) > **Randomized controlled trial (RCT)** — a research design in which participants are randomly assigned to a treatment or control condition to estimate the causal effect of an intervention on an outcome. In [[ai-education|AI in education]], RCTs are the gold standard for establishing whether an AI tool or [[pedagogy|pedagogical]] approach *causes* [[learning-gains|learning gains]], engagement changes, or other outcomes, rather than merely correlating with them. ## Questions to Consider - If a school tells you 'students who used the AI tool scored higher,' why might that still fail to prove the tool caused the gain — even if the difference is large? - Randomization balances known *and unknown* confounders across groups. Before you read, what does random assignment accomplish that simply comparing two intact classrooms cannot, no matter how well-matched they look? - The page calls the RCT the gold standard but lists real costs: artificial settings, fast-changing AI that dates trials, underpowered small samples, and [[ethics|ethical]] constraints on withholding helpful tools. Which of these trade-offs do you think is most often ignored in education research headlines? - An RCT with 1,174 participants found [[generative-ai|GenAI]] closed about three-quarters of an education-based productivity gap. But a well-run RCT can still be conducted on a narrow task in a contrived setting. What should you check about the *outcome measure* before trusting the causal claim? - Consider the ethics problem directly: if you had genuine reason to believe an [[intelligent-tutoring|AI tutor]] helps students learn, is it defensible to randomly deny it to half a classroom for a semester? How would you design an ethically sound study that still isolates the cause? ## Introduction Randomization is what distinguishes an RCT from other designs: by randomly assigning learners to conditions, an RCT balances known and unknown confounders across groups, so any observed difference in outcomes can be attributed to the intervention with high internal validity. ### How RCTs appear in the research - **Micro-RCTs as a response to fast-moving technology:** [[ai-tutoring-micro-rct-gcse-science-2026|Harrison et al. (2026)]] argue that conventional large-scale trials cannot keep pace with tutoring platforms that change materially during a study, and use [[teacher-role|teacher]]-led micro-randomized controlled trials across English secondary schools (644 of 929 students completing post-testing, g = 0.33) to keep causal estimation repeatable. The trade-offs are stated in their own design: 30.7% attrition, [[curriculum-design|curriculum]]-aligned rather than independently standardized outcomes, and only four weeks of follow-up. - **Causal efficacy claims:** RCTs in AIED test whether an AI tutor, tool, or pedagogical treatment improves outcomes. [[generative-ai-education-productivity-gaps|A randomized experiment on generative AI]] with 1,174 participants found GenAI substantially narrows education-based productivity gaps, closing roughly three-quarters of the initial performance difference — a clear causal estimate of AI's effect. - **Comparison to the gold standard:** The [[research-methods-aied|research methods]] page situates RCTs as the strongest design for internal validity while noting their trade-offs — cost, artificial conditions, fast-changing AI, small underpowered samples, and ethical limits on withholding potentially helpful tools. ### Strengths and limitations - **Strengths:** strongest causal inference; clean outcome measurement; supports effect-size estimation; balances confounders through randomization. - **Limitations:** costly and slow; artificial settings can reduce ecological validity; AI tools change faster than trials can run; small samples often underpower detection of meaningful effects; ethical constraints on withholding potentially beneficial AI from a control group. For the fuller treatment of experimental design in AI in education — including when an RCT is appropriate versus quasi-experimental, survey, or computational designs — see [[research-methods-aied]]. ## Connected Concepts - [[research-methods-aied]] - [[ai-ed-evaluation]] - [[educational-measurement]] - [[generative-ai]] - [[higher-ed]] - [[ai-education]] ## Connected Articles - [[generative-ai-education-productivity-gaps]] — Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment - [[ai-changing-teaching-workflows]] — How AI is changing teaching workflows - [[genai-can-harm-teaching-rct-2026]] — Generative AI can harm teaching: an RCT - [[access-not-enough-ai-tutoring-2026]] — Access is not enough: human support improves engagement with AI tutoring - [[burneo-can-edtech-close-learning-gaps-2026]] — World Bank meta-analysis of 14 EdTech RCTs - [[ai-tutoring-micro-rct-gcse-science-2026]] — Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomized Trials of an AI Tutoring Platform in GCSE Science --- ## [Meta-Analysis and Systematic Review](https://edtechdev.github.io/aied/concepts/meta-analysis-systematic-review/) > **Meta-analysis and systematic review** — the family of evidence-synthesis methods researchers use to aggregate and appraise a body of studies, rather than run a single new experiment. A **systematic review** applies a transparent, reproducible protocol to search, screen, appraise, and synthesize the literature on a focused question; a **meta-analysis** goes further by statistically pooling effect sizes across eligible studies to produce a weighted summary estimate and to test moderators. In [[ai-education|AI in education]], these methods are central to establishing the evidence base for whether AI tools work, under what conditions, and for whom — and to exposing gaps, bias, and the field's methodological quality.([[genai-meta-analysis-programming-learning]])([[zerkouk-comprehensive-review-its-2025]]) ## Questions to Consider - Imagine you read ten studies on whether AI tutoring works — two show big gains, three show none, five show small positive effects. How would you decide what to conclude? That tension is exactly what systematic reviews and meta-analyses are built to resolve. - A systematic review and a meta-analysis are often treated as the same thing, but the page distinguishes them: a review synthesizes per a documented protocol, while a meta-analysis statistically pools effect sizes. When would pooling be inappropriate or impossible, even if a careful review exists? - Systematic reviews sit at the top of the evidence hierarchy partly because they compensate for small samples, heterogeneous designs, and conflicting results across studies. Where have you seen a single dramatic study shape opinion even though the pooled evidence was far more mixed? - Meta-analyses produce a weighted summary estimate — a single number like 'an average effect of 0.125 standard deviations.' What does a pooled average hide about the conditions, learners, or contexts where the effect differs — and why does that matter for whether you'd act on it? - Both methods commit to a transparent, reproducible protocol (often PRISMA) precisely because the choices of what to search and include can bias the result. How much would you trust a review that didn't disclose its search and screening decisions? ## Introduction Systematic reviews and meta-analyses sit at the top of the traditional evidence hierarchy precisely because they synthesize many individual studies, compensating for the small samples, heterogeneous designs, and conflicting results that characterize any fast-moving applied field. In AI in education, where new tools and studies appear constantly, reviews play the crucial role of taking stock: mapping what has been studied, aggregating what is known, and flagging where evidence is thin or methodologically weak. They differ from a narrative or integrative literature review, which provides [[qualitative-research|qualitative]] synthesis, in their commitment to a documented protocol and (for meta-analysis) statistical pooling.([[ai-literacy-heptagon-2026]]) ## Systematic review vs. meta-analysis | | Systematic review | Meta-analysis | |---|---|---| | **Core activity** | Search, screen, appraise, synthesize studies per a documented protocol | Statistically pool effect sizes across eligible studies | | **Output** | A narrative/thematic synthesis and evidence map, often with PRISMA flow | A pooled effect estimate with confidence intervals, plus moderator analysis | | **Statistical pooling** | Optional (many reviews are qualitative) | Required | | **When used** | Mapping a fragmented literature, answering "what has been studied and what does it show?" | When multiple comparable [[quantitative-research|quantitative]] studies exist, answering "how large is the effect overall?" | | **Strength** | Transparent, reproducible scope and appraisal | Increased power and precision; detects moderators and heterogeneity | Both follow **PRISMA** (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) as the reporting standard, which documents the search, screening, and [[inclusive-learning|inclusion]] process for [[explainable-ai|transparency]] and reproducibility. An integrative review may follow PRISMA principles for transparency while stopping short of statistical pooling.([[ai-collaborative-learning-systematic-review]])([[ai-literacy-heptagon-2026]]) ## Evidence-synthesis in AI in education ### What reviews accomplish Systematic reviews and meta-analyses in AI in education serve several distinct purposes: - **Establish the evidence base** — determining whether AI tools (tutoring, feedback, assessment, [[conversational-ai|chatbots]]) produce [[learning-gains|learning gains]], and how large those gains are. - **Map the field and its gaps** — a scoping review documents what has been studied, where the evidence is concentrated, and where it is missing (e.g., workplace settings, non-English work, failure cases).([[ai-vocational-education-training-review]]) - **Identify moderators and conditions** — meta-analysis tests whether effects differ by learner population, domain, AI system type, or study design, revealing for whom and under what conditions a tool works. - **Expose methodological quality** — reviews routinely find that the field relies on underpowered, pre-experimental, or quasi-experimental designs and immediate post-tests, tempering conclusions.([[ai-vocational-education-training-review]])([[zerkouk-comprehensive-review-its-2025]]) ### Examples from the knowledge base - **[[genai-meta-analysis-programming-learning|Meta-analysis of GenAI and programming]]** — pools evidence on the productivity-learning trade-off, finding significant productivity gains but no significant learning gain (g ≈ 0), illustrating meta-analysis's ability to separate short-term efficiency from durable learning.([[genai-meta-analysis-programming-learning]]) - **[[ai-vocational-education-training-review|Systematic review of AI in VET]]** — first systematic review of 26 studies, documenting the [[constructivist]]-in-name, behaviorist-in-practice gap and the absence of workplace studies.([[ai-vocational-education-training-review]]) - **[[genai-higher-education-systematic-review-2026|Systematic review of GenAI in higher education]]** — maps opportunities, challenges, and [[pedagogy|pedagogical]] innovations across a five-year window. - **[[zerkouk-comprehensive-review-its-2025|Comprehensive ITS review]]** — a systematic review of [[intelligent-tutoring|intelligent tutoring systems]] with a focus on methodological rigor. - **[[chatgpt-critical-creative-thinking-review|Systematic review of ChatGPT and critical/creative thinking]]** — synthesizes evidence on whether [[llm]] use supports or undermines [[critical-thinking|higher-order thinking]]. - **[[stanford-evidence-base-ai-k12-2026|Evidence base for AI in K-12]]** — reviews the strength of evidence for AI tutoring in schools. - **[[liu-ai-literacy-interventions-meta-analysis-2026|Meta-analysis of AI literacy interventions]]** — three-level meta-analysis of 59 studies (172 effects, 7,211 participants) estimating a large overall effect (g = 0.837) while showing that effectiveness varies by region and learning-outcome focus (knowledge-focused interventions outperformed those targeting skills, attitudes, or [[ethics]]). - **[[ai-literacy-heptagon-2026|AI Literacy Heptagon]]** — an integrative literature review following PRISMA principles, illustrating qualitative synthesis that stops short of meta-analysis.([[ai-literacy-heptagon-2026]]) - **[[genai-scenario-based-healthcare-education-2026|PRISMA 2020 review of GenAI in healthcare scenario learning]]** — Neto and colleagues (2026) systematically searched five databases (9 Nov 2025) for peer-reviewed GenAI studies across scenario-, case-, [[problem-based-learning|problem-]], and [[simulation]]-based healthcare education, screening 1,151 records down to 23 included studies appraised with the [[mixed-methods-research|Mixed Methods]] Appraisal Tool (MMAT). Their thematic synthesis surfaced six cross-cutting themes anchored on [[prompt-engineering|prompt design]] as instructional specification, and documented gaps in validation standardization, longitudinal/comparative designs, and efficiency quantification — a template for rigorous [[medical-education|domain-specific]] GenAI systematic review. - **[[li-language-educators-genai-review-2026|Systematic review of language educators and GenAI]]** — a PRISMA-aligned review of 23 SSCI-indexed empirical studies (December 2022–September 2024) on how language educators perceive, adopt, and learn to integrate GenAI, synthesized through the Aristotelian knowledge typology of episteme (theoretical understanding), techne (practical skill), and phronesis (practical wisdom). It documents cautious, selective adoption weighted toward behind-the-scenes preparation, persistent competency gaps across all three knowledge types, and only three structured PD interventions — illustrating how a review can expose both the evidence base and the field's methodological gaps (here, the scarcity of structured professional-development studies). - **[[riedmann-reinforcement-learning-education-review-2026|Systematic review of RL in education]]** — Riedmann, Schaper & Lugrin (2025) applied a PRISMA-standard protocol to synthesize 89 [[reinforcement-learning]] studies (2000–2024). The review is an instructive case study in synthesis methods and their limits: it maps a rapidly growing but methodologically uneven field — over half of studies (n = 54) reported no statistical testing — and runs effect-size analysis on only 15 papers suitable for pooling, reporting intermediate-to-large effects (Cohen's d) from live evaluations. It also performed publication-bias checks (funnel plot, Egger's test, PET-PEESE) that found no significant bias but had limited power (n = 6), and it stops at thematic-plus-limited-quantitative synthesis rather than a full meta-analysis precisely because heterogeneous evaluation protocols prevented wider pooling — a concrete illustration of the heterogeneity and garbage-in/garbage-out limitations described below. - **[[alsheikh-mapping-ai-integration-higher-education-2026|Mapping review of AI integration in higher education (FACETS + SAMR)]]** — a PRISMA 2020 review screening 959 records down to 22 intervention studies, illustrating how a coding framework (FACETS: Form, AI use case, Context, Education focus, Technology, [[samr-model|SAMR]]) plus an evaluative lens (SAMR) maps a fragmented literature and grades depth of transformation. Most included studies sat at SAMR Substitution/Augmentation, showing mapping reviews can reveal an integration field that is broad but shallow — an alternative to effect pooling when the aim is describing a landscape rather than estimating an effect. - **[[teacher-intervention-k12-ai-based-instruction-2026|Systematic review of teacher intervention in K-12 AI-based instruction]]** — Lee (2026) screened 1,565 records down to 29 studies with two independent raters at every stage (κ = 0.655 screening, κ = 0.647 eligibility) and an MMAT quality appraisal that removed one study. It is an instructive example of synthesis that deliberately stops short of pooling: because the independent effect of teacher intervention could not be separated from AI system design, instructional structure and classroom context in most of the included studies, the review reports conditional outcomes and an explanatory framework of process, strategy and effect rather than an effect size — the same limitation that separates a review from a meta-analysis. - **[[agarwal-ethical-values-norms-aied-2026|Systematic review of ethical values and norms in AIED]]** — Agarwal and colleagues (2026) screened 736 records across Web of Science, ERIC, IEEE CSDL, and ACM DL (plus backward snowballing) down to 25 included articles, consolidating the fragmented [[ethics|AIED ethics]] literature into six main ethical values (non-discrimination, data stewardship, [[human-in-the-loop-ai|human oversight]], goodwill, explicability, educational aptness) and mapping ethical norms onto a stakeholder-by-value matrix. The review illustrates how a systematic protocol can synthesize a conceptually fragmented, largely non-empirical literature (only three of 25 articles were methodology papers or original research) and turn it into an actionable framework — here, a foundation for [[governance]] and [[educational-policy-ai|policy]]. ### Interdisciplinary review of AI for dyslexia **[[dabaghi-ai-dyslexia-education-review-2026|Dabaghi, D'Urso & Sciarrone (2026)]]** present a PRISMA-guided, interdisciplinary systematic review (2018–2024, n=72) of AI and generative AI to support students with dyslexia in education. The review maps AI across detection, assistive support, and [[personalized-learning|personalized learning]], finding these strands evolve in parallel rather than in integration, driven more by technological opportunity than by consolidated educational theory. It documents that generative AI is under-utilized in this domain (GAI research, all from 2024, clusters into chatbots, teacher-training support, and exploratory studies) and that ML-based help-education tools fall into five areas (specific applications, [[student-engagement|engagement]], personalization, recommendation, generic support) while emphasizing technical performance over ecological validity. Open challenges include limited experimental validation, scalability and [[accessibility]] of diagnostic tools, ethics/privacy concerns with sensitive student data, limited teacher support, and language/cultural barriers. The review's own methodological limitations — interpretative classification bias, exclusion of non-English studies, heterogeneous evaluation protocols that prevent quantitative synthesis, and a rapidly evolving GAI evidence base — illustrate the systematic-review family's core tension: a transparent protocol can map a fragmented field, but heterogeneous evaluation prevents statistical pooling, so the review stops at thematic synthesis rather than meta-analysis. ## AI-era synthesis challenge: productivity vs. learning Reviews of generative-AI interventions face a distinctive challenge that the knowledge base's synthesis research highlights: **separating productivity gains from durable learning gains.** Because [[generative-ai|generative AI]] can inflate immediate task performance (homework, assisted practice) without producing learning, meta-analyses must be careful about which outcome they pool. [[genai-meta-analysis-programming-learning|The GenAI-and-programming meta-analysis]] found large productivity gains but no significant learning gain (g ≈ 0) — a clean illustration. [[stromberg-generative-ai-learning-penalty-secondary-2026|Large-scale field studies]] and [[generative-ai-reduced-study-time-math|unassisted-measure research]] show that the measured effect depends on whether outcomes are AI-assisted or proctored/unassisted. Reviews should therefore report assisted and unassisted outcomes separately, distinguish performance from [[learning-gains|learning]], and flag studies that measure only immediate AI-supported performance. A complementary caution emerges from the [[liu-ai-literacy-interventions-meta-analysis-2026|AI-literacy meta-analysis]]: which **outcome** is pooled also shapes the answer — knowledge-focused [[ai-literacy]] interventions showed larger effects than those targeting skills, attitudes, or ethics, so a review that pools only knowledge outcomes can overstate what AI-literacy instruction achieves overall. This connects to [[ai-ed-evaluation]] and [[summative-assessment]]. ## Strengths and limitations **Strengths:** - Efficient synthesis of a large, fragmented literature - Meta-analysis yields pooled effect estimates, increases statistical power, and detects moderators and heterogeneity - Systematic protocols improve transparency and reproducibility over narrative reviews - Essential for evidence-based practice and for identifying research gaps **Limitations:** - **Garbage-in/garbage-out** — the synthesis is only as good as the quality of included studies; weak primary designs yield weak pooled conclusions - **Publication bias** — null or negative results are under-published, inflating pooled effects - **Heterogeneity** — varied designs, outcome measures, and [[ai-technologies|AI systems]] make direct pooling hard and can undermine the meaning of a single effect size - **Rapid obsolescence** — the AI tool landscape changes quickly, so reviews can date fast - **Scope constraints** — single-database or English-only searches may miss relevant work.([[ai-collaborative-learning-systematic-review]])([[ai-vocational-education-training-review]]) **Meta-research warns the AIED synthesis base is currently weak.** A growing set of critiques documents that the field's headline AI-effect sizes — especially from early meta-analyses — are inflated by publication bias, construct incoherence, and methodological shortcuts. [[bartos-ai-learning-meta-meta-analysis-2026|Bartoš et al. (2026)]], meta-analyzing 1,840 effect sizes from 67 meta-analyses, estimate the publication-bias-adjusted AI effect at roughly one-third the reported magnitude (SMD ≈ 0.196), with extreme heterogeneity. [[oneill-presumed-effective-meta-analysis-2026|O'Neill (2026)]] audits 14 high-impact AIED meta-analyses and finds none had a coherent construct, valid publication-bias assessment, or resolved heterogeneity; twelve treated dependent effect sizes as independent. [[weidlich-chatgpt-effect-search-cause-2025|Weidlich et al. (2025)]] show most primary comparisons lack a well-defined treatment, control, and learning measure. This means readers should treat pooled AIED effect sizes as upper bounds until synthesis quality improves — see [[limitations-in-aied-research|Limitations in AIEd Research]] for the full analysis. - **Reporting standards have not kept pace with automation.** PRISMA-LLM analyses SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, and documents a growth rate near 4.7% per month alongside an accountability gap: since 2023, 38.0% of software and product papers reported no evaluation at all, against 9.3% of LLM papers, and 52% of positive-only LLM evaluations reported an unmet high-bar concern. Because LLM and software pipelines now participate in stages that can alter the evidence base, the framework requires disclosure of where in the review workflow automation operated, what was evaluated, and which limitations were checked — a direct extension of the transparency problem that [[limitations-in-aied-research|AIED review critiques]] have documented for meta-analyses of learning effects. ([[prisma-llm-ai-assisted-systematic-reviews-2026]]) ## Relationship to other methods Within the knowledge base's methodological landscape, meta-analysis and systematic review are the **synthesis** family, complementing primary designs: - **Primary studies** (experiments, surveys, qualitative work, [[design-based-research|design-based research]]) generate individual findings; reviews aggregate them. See [[research-methods-aied]]. - **Effect-size reporting** in primary studies (e.g., [[rct|RCTs]]) is what makes later meta-analysis possible — reviews depend on studies reporting comparable, extractable effect sizes. - **Evaluation** ([[ai-ed-evaluation]], [[benchmark]]) assesses individual systems; reviews assess the *literature* on systems and interventions. - **Educational measurement** ([[educational-measurement]], [[assessment-validity]]) concerns the quality of the outcome measures that reviews pool. ## Implications for researchers 1. **Report extractable effect sizes.** For a literature to be meta-analyzable, primary studies must report comparable effect sizes and adequate methods detail — a responsibility of every AIED study.([[research-methods-aied]]) 2. **Follow a transparent protocol.** PRISMA-guided search, screening, and appraisal make reviews reproducible and defensible. 3. **Interpret pooled effects cautiously.** Attend to heterogeneity, publication bias, and the quality of included studies before drawing strong conclusions. 4. **Use reviews to set the agenda.** Reviews' documented gaps (failure cases, workplace settings, non-English and non-indexed work, long-term outcomes) should guide where new primary research is needed.([[ai-vocational-education-training-review]]) ## GenAI in Healthcare Scenario Learning - **PRISMA 2020 review of GenAI in healthcare scenario learning.** Neto and colleagues (2026) systematically searched five databases (9 Nov 2025) for peer-reviewed GenAI studies across scenario-, case-, problem-, and [[simulation]]-based healthcare education, screening 1,151 records down to 23 included studies appraised with the [[mixed-methods-research|Mixed Methods]] Appraisal Tool (MMAT). Their thematic synthesis surfaced six cross-cutting themes anchored on [[prompt-engineering|prompt design]] as instructional specification, and documented gaps in validation standardization, longitudinal/comparative designs, and efficiency quantification — a template for rigorous [[medical-education|domain-specific]] GenAI systematic review. ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[research-methods-aied]] - [[rct]] - [[ai-ed-evaluation]] - [[benchmark]] - [[educational-measurement]] - [[assessment-validity]] - [[learning-gains]] - [[summative-assessment]] - [[ai-education]] - [[higher-ed]] - [[simulation]] ## Connected Articles - [[xia-ai-interdisciplinary-higher-education-review-2026]] — Systematic review of AI in interdisciplinary higher education (59 studies) - [[generative-ai-k12-teaching-learning-systematic-review-2026]] — Systematic review of generative AI in K-12 teaching and learning (Marzano 2026) - [[nguyen-genai-global-south-review-2026]] - [[espino-ai-business-education-review-2026]] - [[khalifeh-redefining-personalized-learning-ai-2026]] — Redefining personalized learning: systematic review - [[alrazeeni-transforming-nursing-education-ai-2026]] — AI in nursing education: systematic review - [[edurev-100741-tpack-genai-review]] — Systematic review of GenAI in student learning from a TPACK perspective - [[genai-meta-analysis-programming-learning]] — Meta-analysis of GenAI's effect on productivity and learning in programming - [[ai-vocational-education-training-review]] — First systematic review of AI in vocational education and training - [[ai-collaborative-learning-systematic-review]] — PRISMA systematic review of AI-powered collaborative learning - [[genai-higher-education-systematic-review-2026]] — Systematic review of GenAI in higher education - [[robot-assisted-language-learning-meta-analysis-2026]] — Meta-analysis of AI-enhanced embodied robot-assisted language learning - [[zerkouk-comprehensive-review-its-2025]] — Comprehensive systematic review of intelligent tutoring systems - [[chatgpt-critical-creative-thinking-review]] — Systematic review of ChatGPT and critical/creative thinking - [[stanford-evidence-base-ai-k12-2026]] — Evidence base for AI in K-12 - [[liu-ai-literacy-interventions-meta-analysis-2026]] — Meta-analysis of AI literacy intervention effects - [[ai-literacy-heptagon-2026]] — Integrative literature review of AI literacy dimensions (PRISMA-guided) - [[ai-metacognition-stem-review]] — Systematic review of AI and metacognition in STEM - [[llm-intervention-design-cs-review]] — Review informing LLM intervention design in CS - [[human-autonomy-agency-hri-review-2025]] — Review of human autonomy and agency in human-robot interaction - [[rail-ed-genai-literacy-teacher-education]] — Review of GenAI literacy in teacher education - [[student-llm-interaction-taxonomy-review-2026]] - [[zhao-genai-higher-order-thinking-meta-2026]] — GenAI and higher-order thinking meta-analysis - [[daniel-ai-sustainability-scoping-review-2026]] — Scoping review of AI for sustainability and sustainable AI (Daniel et al. 2026) - [[policy-deficit-ai-sel-2026]] — The Policy Deficit in AI × SEL Research - [[genai-scenario-based-healthcare-education-2026]] — Systematic review of GenAI in scenario-based healthcare education (Neto et al. 2026) - [[alsheikh-mapping-ai-integration-higher-education-2026]] — Mapping review classifying 22 AI-integration studies with FACETS + SAMR; most sit at Substitution/Augmentation - [[agarwal-ethical-values-norms-aied-2026]] — Ethical values and norms for AI in education - [[li-language-educators-genai-review-2026]] — Language educators' practices and development with GenAI - [[dabaghi-ai-dyslexia-education-review-2026]] — AI to help people with dyslexia in education - [[riedmann-reinforcement-learning-education-review-2026]] - [[weidlich-chatgpt-effect-search-cause-2025]] — ChatGPT in Education: An Effect in Search of a Cause (media-comparison critique) - [[bartos-ai-learning-meta-meta-analysis-2026]] — Meta-meta-analysis: bias-adjusted AI effects ~1/3 of reported size - [[oneill-presumed-effective-meta-analysis-2026]] — Presumed Effective: forensic audit of 14 AIED meta-analyses - [[teacher-intervention-k12-ai-based-instruction-2026]] — Teacher intervention in K-12 AI-based instruction: a systematic review --- ## [Latent Profile Analysis](https://edtechdev.github.io/aied/concepts/latent-profile-analysis/) > **Latent profile analysis (LPA)** — a person-centered method that sorts a sample into unobserved subgroups (profiles) when each case is described by several variables at once. It is the continuous-indicator member of the mixture-modeling family; its sibling **latent class analysis (LCA)** applies the same logic to categorical indicators. Both ask a different question from the [[quantitative-research|variable-centered]] models that dominate AI-in-education research: not "how much does X predict Y on average" but "how many different kinds of learner, [[teacher-role|teacher]], or manager hide inside that average." In this knowledge base the method shows that one AI tool lands very differently across subgroups — five ethical-awareness profiles among Ghanaian undergraduates, six readiness typologies among Ukrainian education managers, four [[generative-ai|ChatGPT]]-acceptance profiles among Taiwanese pre-service teachers. ## Questions to Consider - A study reports that students' average comfort with AI is 3.8 out of 5. What might that average conceal if one group is enthusiastic and another quietly resistant? - LPA models continuous scores; LCA models categories. If you code interview responses as "mentions [[cognitive-offloading|overreliance]]: yes/no," which do you need? - A five-profile solution reports entropy of 0.816; a four-profile rival scores 0.929 but describes the data less richly. Which would you publish? - Profiles are descriptive, not causal. If a "Resistant Skeptics" profile shows high ease of use but low intention to adopt, what does that support, and what does it not? ## Introduction Most AI-in-education evidence is variable-centered: it estimates average relationships, and averages assume homogeneity. Person-centered methods instead take the person as the unit of analysis and ask how many distinct configurations of attributes exist in the sample. LPA belongs to that family, and its payoff here is that it converts learner diversity from a rhetorical claim into a measurable finding: when close to a quarter of undergraduates sit in the two lowest ethical-awareness profiles despite a comfortable sample mean, differentiation stops being a design preference. ## What the method does, and when it is the right tool LPA assumes the sample is drawn from a mixture of subgroups, each with its own means and variances on the indicators. The researcher supplies the indicators and the number of groups; the algorithm estimates each case's probability of membership and assigns by highest probability, returning a mean profile per group, a size per profile, and a summary of how cleanly cases separate. Which member of the family you use follows the indicators, not the research question: - **LCA — categorical indicators.** Becker and colleagues converted coded response categories from 1,189 physics students' open answers into indicator variables and retained two classes: Pragmatic Users (70 percent) and Skeptical Non-Users (30 percent); because an unmentioned topic counts as "no endorsement," the authors flag zero-inflation in the analysis dataframe. - **LPA — continuous indicators** such as scale scores and construct means. Chen and colleagues profiled 128 Taiwanese pre-service teachers on five [[technology-acceptance-model|TAM]]/UTAUT2 constructs; Acquah and colleagues profiled 509 Ghanaian undergraduates on three ethical-awareness dimensions; Schweder and colleagues profiled 2,464 students on [[motivation|motivational]] need satisfaction. - **Latent (profile) transition analysis** extends the family over time, estimating profiles per wave and the probability of moving between them. Liang and colleagues tracked 2,086 students' AI learning motivation across a year; Wu followed 457 Japanese-language learners over three waves, the maladaptive profile shrinking from 26.48 to 17.74 percent. Ordinary clustering ([[machine-learning|k-means]] and hierarchical) pursues the same person-centered intent but partitions cases by geometric distance rather than estimating a probability model. Reach for LPA or LCA when the question concerns subgroups — whether "the learner" is a fiction at your sample's level of aggregation, and whether subgroups differ in shape as well as level. Do not reach for it when you need an average treatment effect: a five-profile solution inside a [[rct|randomized trial]] of an [[intelligent-tutoring|adaptive tutor]] predicted posttest differences (partial η² = .27) but found no condition × profile interaction (ps ≥ .198). ## How the corpus uses it Three uses recur. - **Establishing heterogeneity before designing for it.** Kremen and colleagues' survey of 395 Ukrainian education managers used person-centered LCA to show that "the manager" is a fiction: six typologies run from Competency-constrained (25.6 percent, willing but unskilled) to Barrier-free skeptics (highest readiness yet 74 percent AI distrust), and the authors read them as a mandate for differentiated training. - **Recovering subgroups an aggregate hides.** Acquah and colleagues retained five ethical-awareness profiles from Comprehensive Very High (26.1 percent) to Low Ethical Awareness (4.5 percent, beneficence 2.06), the two lowest together covering close to a quarter of the sample. Chen and colleagues found that Resistant Skeptics reported high perceived ease of use but very low behavioral intention — the corpus's clearest demonstration that the ease-of-use/intention paradox is invisible to a mean-level model. - **Profiling calibration rather than level.** The teacher AI-literacy study applied LPA to the agreement between [[self-report-measures|self-report]] and objective measures, yielding six profiles: overestimation, underestimation, alignment, and a low/low group concentrated among teachers without prior [[ai-literacy|AI literacy]] experience. Here profiles describe a pattern across instruments, not a score band. Profiles then serve as an independent variable: discipline shaped membership among pre-service teachers (Cramér's V = 0.532, STEM students concentrated in Technology Pioneers), and membership predicted later [[self-efficacy]], burnout, and [[anxiety-and-stress|AI anxiety]] elsewhere in the corpus. ## Choosing the number of profiles No single statistic selects the solution; the corpus treats retention as a judgment made from several criteria together. - **Information criteria (BIC, AIC).** Lower is usually better, but a monotonic decline signals a problem rather than a winner. In the Ukrainian manager study BIC fell monotonically across the two- to six-class range with no clear minimum, and the authors call their six-class solution exploratory on that evidence. - **Entropy.** A summary of classification certainty, closer to 1 meaning cleaner assignment. Chen and colleagues report 0.985 for four profiles; Acquah and colleagues report 0.816 for five profiles against 0.929 for four, and themselves suggest consolidating for a more stable grouping. - **[[explainable-ai|Interpretability]] and profile size.** The trust-in-AI profiling study kept three clusters although the Calinski–Harabasz index preferred two, because three were interpretable, and it notes that a silhouette coefficient of 0.288 signals weak or borderline separation. Chen and colleagues warn that their smallest profile (14.06 percent of 128 cases) may be unstable, and the Ghana study's smallest profile holds only 23 students, which its authors propose merging. - **Stability under resampling.** Bootstrap stability is the honest check on whether profiles would recur in a new sample: mean adjusted Rand index was 0.385 for the Ukrainian six-class solution but 0.989 across 100 random initializations in the trust-in-AI study — the same nominal design, very different evidential weight. - **Bootstrap likelihood ratio tests (BLRT)** and the Lo–Mendell–Rubin test are standard companions to BIC and entropy in the wider mixture-modeling literature, but the profile studies in this knowledge base do not report them. Where a page reports only BIC and entropy, treat the class count as provisional. Report the comparisons, not just the winner: a page that says "we retained five profiles" without the rival solution, the entropy values, and the smallest profile's size gives readers no way to judge the choice. ## Reading and reporting the results, and where they go wrong Read a profile by its shape as well as its level. In the Ghana study the profiles differ in the configuration of autonomy, beneficence, and fairness — the largest pairs strong autonomy endorsement with lower beneficence — so two profiles can sit at similar overall levels and still call for different teaching. Four cautions, each stated in the source pages: 1. **Profiles are descriptive.** They say who is in the sample, not why, and cross-sectional designs cannot show stability or movement (the pre-service teacher, Ghana, and physics studies). Only the longitudinal analyses — a year-long transition study and Wu's three waves — support claims about movement, and even there movement is association, not intervention effect. 2. **Membership is estimated, not observed.** Cases are assigned by highest posterior probability, so individuals near a boundary are classified with real uncertainty; low entropy and weak silhouette values mean the boundaries should be read as soft. 3. **Indicator choice defines the answer.** Profiles depend entirely on which variables enter the model, so heterogeneity the study never measures is heterogeneity it cannot find — the Ghana study collected only gender among backgrounds and therefore cannot say what predicts membership. 4. **Detected heterogeneity is not detected causation.** The tutor trial's profile × condition null is the reminder: profiling an outcome inside a trial is not a moderation test of the treatment. ## Implications for AI in education 1. **Measure heterogeneity before recommending an intervention.** Profiles in this corpus repeatedly place a quarter or more of a sample below the headline average; design for the profiles that exist rather than for the mean. 2. **Prefer person-centered methods when the deliverable is differentiation.** The pre-service teacher and manager studies both present that shift as their methodological contribution, because average predictors cannot reveal the configurations that justify differentiated support. 3. **Report retention evidence in full.** Publish BIC/AIC comparisons, entropy, profile sizes, and a stability check, and say plainly when a class count is exploratory. 4. **Treat profiles as diagnostics, not labels.** A profile is a research construct with an estimated boundary, so a tool that assigns individuals to named categories carries the caution of any [[educational-measurement|measurement]] instrument — and [[differential-effects-across-learner-groups|differential effects across learner groups]] warrant the scrutiny of any subgroup analysis. ## Connected Concepts - [[quantitative-research]] - [[mixed-methods-research]] - [[research-methods-aied]] - [[machine-learning]] - [[learning-analytics]] - [[student-modeling]] - [[educational-measurement]] - [[self-report-measures]] - [[differential-effects-across-learner-groups]] - [[technology-acceptance-model]] ## Connected Articles - [[ai-ethical-awareness-ghana-students-2026]] — Five ethical-awareness profiles; entropy 0.816 (five-class) vs. 0.929 (four-class); smallest profile n = 23 - [[ai-adoption-readiness-ukraine-education-managers-2026]] — Six manager typologies; BIC monotonic across 2–6 classes; bootstrap stability mean ARI 0.385 - [[chen-preservice-teachers-chatgpt-lpa-2026]] — Four ChatGPT-acceptance profiles (entropy 0.985); ease-of-use ≠ intention paradox - [[becker-chatgpt-typology-physics-2026]] — LCA on categorical indicators from 1,189 coded responses; zero-inflation caveat - [[ai-literacy-assessment-misalignment]] — LPA on self-report vs. objective agreement: six calibration profiles - [[wu-psychological-adaptation-ai-japanese-learning-2026]] — Three-wave latent profile transition analysis of psychological adaptation - [[liang-ai-learning-motivation-sdt-2026]] — Latent transition analysis of three motivation profiles over a year - [[trust-in-ai-psychological-profiles-ml-2026]] — K-means profiles; silhouette vs. Calinski–Harabasz disagreement; ARI 0.989 - [[saihi-ahmed-genai-adoption-personas-higher-ed-2026]] — Hierarchical and k-means clustering into four GenAI adoption personas - [[student-motivation-need-satisfaction-genai-sdt-2026]] — Person-centered LPA combined with variable-centered comparisons --- ## [Network Analysis](https://edtechdev.github.io/aied/concepts/network-analysis/) > **Network analysis** — the family of [[research-methods-aied|research methods]] that model entities (people, concepts, actions, or codes) as **nodes** connected by **edges** representing relationships or transitions, then analyze the structure and dynamics of the resulting network to reveal patterns invisible to frequency counts or pairwise comparisons. In AI-in-education research, network analysis is used to map interaction patterns between learners and AI tools, model how knowledge or discourse elements co-occur, and trace temporal sequences of behavior. It includes distinct variants — **Epistemic Network Analysis** (ENA, modeling the co-occurrence of codes/constructs), **Social Network Analysis** (SNA, modeling relationships between people), and **Transition Network Analysis** (TNA, modeling temporal sequences of states) — each of which operationalizes "learning as connection" in a different way.([[tracing-genai-literacy-interaction-patterns]])([[penny-transition-network-analysis-efl-writing-2026]])([[misiejuk-cognitive-offloading-prompting-2026]]) ## Questions to Consider - When you hear 'network analysis' in education, what images come to mind—friendship maps of students, links between ideas, or something else? How are those different from simply counting how often things occur? - Suppose you wanted to know whether students actually engage with an AI writing tool's feedback versus just getting answers. Why might a 'how many times did they click' metric miss the story that a sequence of actions (e.g., a revision loop vs. a chat loop) would reveal? - The page distinguishes Epistemic, Social, and Transition network analysis. Without knowing the details, can you guess which variant you'd use to study (a) how people collaborate, (b) which ideas co-occur in student reasoning, and (c) how learners move between states over time? - A researcher finds that high- and low-literacy learners use the same AI tool but produce very different network structures of reasoning. What does that tell you about evaluating AI tools with a single average score? - Network metrics like 'density' and 'centrality' describe whether interaction is random or organized around hubs. When would an organized network centered on one learner be a sign of good collaboration—and when a sign of a problem? ## Introduction Network analysis methods share a core premise: that the structure of connections — not just their presence or frequency — carries meaning. Rather than asking "how much of X occurred," they ask "how are elements connected, and what does that connectivity reveal about [[metacognition|cognition]], [[collaborative-learning|collaboration]], or learning processes?" This makes them especially valuable in AI-in-education, where researchers increasingly want to understand the *process* of learner–[[student-ai-interaction|AI interaction]] (how learners navigate [[feedback]], dialogue, and revision) rather than only the product (final scores, error rates). ## Variants used in the knowledge base's corpus - **Epistemic Network Analysis (ENA)** — the most common variant in the knowledge base (discussed in ~24 articles). ENA models the co-occurrence of codes or constructs within segments of discourse or activity, producing networks that show which ideas, skills, or epistemic actions tend to be connected in a given context. It is used to compare how different groups (e.g., high- vs. low-literacy learners, human vs. AI collaborators) structure their cognition.([[tracing-genai-literacy-interaction-patterns]])([[hao-human-ai-collaborative-problem-solving-cognition]]) - **Social Network Analysis (SNA)** — models relationships between people (learners, teachers, agents) to reveal collaboration structures, influence, centrality, and community. Useful for studying [[collaborative-learning|collaborative]] and peer learning.([[misiejuk-cognitive-offloading-prompting-2026]]) - **Transition Network Analysis (TNA)** — models temporal sequences of discrete states (e.g., learner actions in a tutoring session) as a directed network, quantifying the probability of moving between states. TNA is used to reveal behavioral loops, pathways, and uptake dynamics in learner–AI interaction.([[penny-transition-network-analysis-efl-writing-2026]]) These differ from a **[[knowledge-graph]]**, which is a data structure for representing and reasoning over facts (an ontology/triple store), not an analytical method for studying process or relationship structure. ## Network analysis in AI-in-education research Network methods are used across the knowledge base's evidence base to answer questions that aggregate metrics cannot: - **Open the "black box" of learner–AI interaction.** TNA reveals the *process* — the behavioral loops and pathways learners take when using AI tools (e.g., a "revision loop" vs. a "chat loop" in [[conversational-ai|chatbot]]-scaffolded [[writing-education|writing]]) rather than just final output.([[penny-transition-network-analysis-efl-writing-2026]]) - **Compare cognitive structuring across groups.** ENA shows how different groups connect constructs differently — e.g., how [[metacognition]] co-occurs with delegation vs. human reasoning in human–AI collaboration, revealing different collaboration modes.([[hao-human-ai-collaborative-problem-solving-cognition]]) - **Trace AI-literacy and interaction signatures.** ENA on interaction logs identifies distinct patterns of [[llm|LLM]] use (iterative strategic refinement vs. linear commands), distinguishing learner [[ai-literacy|proficiency]] and development.([[tracing-genai-literacy-interaction-patterns]]) - **Analyze discourse and framing.** ENA is applied to [[qualitative-research|qualitative]] and [[multimodal]] data (e.g., YouTube frames of ChatGPT in education) to reveal the structure of public or disciplinary discourse.([[youtube-frames-chatgpt-education]]) - **Complement self-report and product metrics.** Because network methods use observed behavioral data, they can expose discrepancies between what learners claim and what they actually do — a recurring finding in the knowledge base's feedback-uptake literature. ## Methodological considerations - **Coding is the foundation.** All network variants depend on reliably coding raw data (utterances, events, relationships) into discrete nodes/codes; automated LLM-based coding is increasingly used but requires human validation (e.g., Fleiss' κ of 0.70–0.71 in TNA studies).([[penny-transition-network-analysis-efl-writing-2026]]) - **Network-level metrics summarize structure.** Density, reciprocity, centralization, and in-/out-strength describe whether interaction is random or organized around "gravitational" hubs, and how reciprocal the exchange is. - **Statistical comparison is needed for group differences.** Chi-squared tests or permutation testing are used to establish that observed network differences (e.g., by proficiency) are not due to chance. - **Interpret with care.** Node granularity (e.g., a coarse "chat" node) can obscure intent; automated classification carries some ambiguity; and cross-sectional network structure does not establish causality. ## Implications for AI-in-education research 1. **Prefer process methods over product-only metrics.** To evaluate whether AI tools support learning, model how learners actually engage (uptake, dialogue, revision) with sequence/network methods rather than relying on final scores alone. 2. **Use ENA to compare cognitive structuring.** When asking how different learners or modes (human vs. AI) structure their reasoning, ENA provides a direct, visual comparison of co-occurrence networks — a technique well-suited to [[student-modeling|student modeling]] of how learners connect ideas. 3. **Validate automated coding.** With large log datasets, [[llm|LLM]]-based classification is powerful but must be checked against human coding (report inter-rater agreement) before interpreting network structure. 4. **Design for differentiation.** Network analysis often reveals that the *same* AI tool produces different interaction patterns across learner subgroups — informing adaptive design rather than one-size-fits-all evaluation. ## ENA Validation of Simulated Collaborative Dialogue - **ENA as validation for simulated dialogue.** Fang (2026) applies Epistemic Network Analysis to evaluate whether fine-tuned LLM agents reproduce the structure of real collaborative [[problem-solving]] dialogue. Comparing simulated adjacency vectors to the empirical network, he reports an ENA distance of 0.17 — within the 95th-percentile threshold of the null distribution, with a permutation p-value of 0.65 — demonstrating ENA's power as a [[quantitative-research|quantitative]] check on the fidelity of generative [[simulation|simulations]] of discourse, alongside other applications of ENA/SNA/TNA in education research. ## Connected Concepts - [[learning-analytics]] - [[knowledge-graph]] - [[meta-analysis-systematic-review]] - [[student-modeling]] - [[student-engagement]] - [[collaborative-learning]] - [[metacognition]] - [[ai-literacy]] - [[scaffolding]] - [[feedback]] ## Connected Articles - [[penny-transition-network-analysis-efl-writing-2026]] — TNA of learner-chatbot interactions in scaffolded EFL writing - [[tracing-genai-literacy-interaction-patterns]] — ENA of GenAI literacy interaction patterns - [[hao-human-ai-collaborative-problem-solving-cognition]] — ENA of human-AI collaborative problem solving - [[misiejuk-cognitive-offloading-prompting-2026]] — Cognitive offloading and prompting (SNA/network methods) - [[youtube-frames-chatgpt-education]] — ENA of YouTube frames of ChatGPT in education - [[agency-gap-ai-writing]] — The agency gap in AI-supported writing (ENA) - [[dai-chatbots-problem-posing-primary-2026]] — GenAI chatbots and problem posing in primary science - [[llm-agents-collaborative-problem-solving-simulation-2026]] — Fine-tuned participant-specific LLM agents reproducing collaborative problem solving dialogues (Fang 2026) --- ## [AI Ed Evaluation](https://edtechdev.github.io/aied/concepts/ai-ed-evaluation/) > **AI-ed evaluation** — the body of methods, benchmarks, and criteria used to assess whether [[ai-education|AI education]] tools ([[llm]]-based tutors, [[automated-assessment|automated graders]], feedback systems, agents) actually work — not just on headline accuracy, but on reliability, [[pedagogy|pedagogical]] quality, validity, and real learning impact. A recurring theme across the knowledge base's research is that evaluation must be domain-specific, reliability-aware, and anchored in human judgment and educational outcomes rather than single aggregate accuracy numbers. ## Questions to Consider - How would you decide whether an AI tutoring tool 'works'? What evidence — beyond a headline accuracy number — would convince you it actually improves learning? - A key finding is that reliability does not guarantee validity: a system can be highly consistent yet misjudge what good teaching looks like. Why might a stable, repeatable AI still be wrong? - Evaluation must be domain-specific — a benchmark that works for one subject can mislead for another. When you see a glowing benchmark result for an AI tool, what would you want to check about the context? - The 'ground truth' systems are judged against is often contested — what counts as a correct answer or grade varies across experts and disciplines. How does that uncertainty complicate trusting any evaluation? - Research shows text-based LLM evaluators privilege explicitly verbalized behaviors and under-weight implicit context — an 'explicit-cue bias.' What kinds of good teaching might a machine systematically miss because it only looks for the obvious? - Modern evaluation also weighs environmental and infrastructural cost — energy, hardware — not just output quality. Should [[sustainability]] factor into how you judge an AI tool's value? ## Introduction AI-ed evaluation spans several distinct objects of assessment. It can evaluate the **output** (is the AI's answer, grade, or feedback correct and reliable?), the **process** (does the tool support valid, defensible assessment and learning?), and the **agent** (does an AI tutor or agent teach effectively and behave appropriately?). Each requires different methods and raises different validity questions. ### How AI-ed evaluation appears in the research - **Output reliability and ground truth:** [[ground-truth-reliability-aied|Modernizing ground truth]] argues that reliability problems in AI-ed evaluation often trace back to the reference data itself — the "ground truth" labels systems are judged against — and proposes four shifts toward improving reliability and validity. [[calibrating-trustworthiness-llm-education-2026|Calibrating trustworthiness]] co-designs evaluation metrics and visualizations with stakeholders so that trust in an AI tool rests on demonstrated, interpretable evidence. - **Automated grading and scoring:** [[cong-confidence-asag-2026|LLM short-answer grading]], [[cong-confidence-asag-2026|confidence-aware ASAG]], [[cotal-formative-assessment-scoring-2026|CoTAL human-in-the-loop prompt engineering]], and [[llm-cognitive-diagnosis-handwritten-math|cognitive-diagnosis of handwritten math]] show that LLMs can grade and diagnose, but that reliability depends on [[human-in-the-loop-ai|human oversight]], domain-specific grounding, and confidence calibration rather than raw model size. - **Pedagogical quality and alignment:** [[machines-misread-pedagogical-quality|Why machines misread pedagogical quality]] documents human–machine misalignment in judging what makes instruction good, and [[tutoring-effectiveness-index|the Tutoring Effectiveness Index]] predicts tutor quality from teaching behavior. [[responsible-assessment-ai-era-stanford-2026|Responsible assessment in the AI era]] and [[authentic-products-authenticated-processes-2026|authenticated processes]] argue that evaluation must reach beyond correct answers to whether assessment remains authentic, valid, and defensible when AI can produce the "products" of learning. ### Why evaluation is hard in AI-ed AI-ed evaluation is difficult for several reasons. First, **reliability is not enough** — a system can agree with a rubric yet misjudge pedagogy, as [[machines-misread-pedagogical-quality|human–machine alignment research]] shows. Second, **ground truth is contested** — what counts as a "correct" answer, grade, or teaching move is itself a judgment that varies across disciplines and experts, per [[ground-truth-reliability-aied|ground-truth modernization]]. Third, **educational validity is multidimensional** — [[assessment-validity]], [[formative-assessment]], and [[authentic-assessment]] each impose different criteria that a single accuracy metric cannot capture. Finally, **the target keeps moving** — agentic AI and [[multimodal]] models demand evaluation frameworks ([[agentic-ai]], [[tool-invariant-framework-agentic-ai|tool-invariant assessment]]) rather than reuse of text-model benchmarks. Evaluation findings are also subject to the same cross-cutting limitations that affect all AIED research — they age as AI improves, depend on reproducibility and FAIR practices, and may rest on proprietary systems — so evaluation results should be read with the caveats in [[limitations-in-aied-research]]. A further, emerging dimension is **resource sustainability**: on-premise deployments increasingly report energy consumption and hardware requirements (e.g., VRAM, mWh per query) alongside accuracy — see [[shen-sustainable-ai-knowledge-base-cs-education-2026|sustainable on-premise knowledge-base assistants]] — so that a complete evaluation weighs environmental and infrastructural cost, not just output quality. **Reliability does not guarantee validity.** [[melo-llm-classroom-observation-teach-2026|Validation of LLM-based classroom observation]] shows that a model can be highly stable across repeated evaluations yet still misalign with expert judgment, and conversely that models aligning well with experts are often more variable — reliability and accuracy decouple, so a single-pass accuracy figure can overstate dependability. The same study documents an **explicit-cue bias**: text-based LLM evaluators privilege explicitly verbalized behaviors and under-weight implicit or contextual evidence (e.g., sustained student [[self-regulated-learning|self-regulation]] where a rubric allows high ratings on absence-tolerant criteria), producing systematic rather than random disagreement. This underscores that measurement reliability is a prerequisite for — not a proxy for — valid interpretation, and that evaluation must include repeated-measures stability checks alongside expert-anchored accuracy. **Aggregate accuracy hides who is served poorly.** [[drawedumath-vlm-struggling-students-2026|Evaluations of vision-language models on DrawEduMath]] show that overall accuracy obscures a systematic weakness: models underperform precisely on the student work that needs the most pedagogical help (erroneous, struggling-student work), so disaggregating evaluation by student proficiency and error status is necessary to avoid overstating capability and widening achievement gaps. - **Benchmark and grader errors are mistaken for model failure.** Expert re-grading of six widely used [[physics-education|physics]] benchmarks audited 250 rejected items and attributed 143 (57.20%) to benchmark defects and 95 (38.00%) to grader errors, leaving only 12 (4.80%) genuine model errors, so 95.20% of the measured gap was not attributable to the model. Repairing the items moved HLE-Physics mean@4 from 47.28% to 78.66% and CritPt mean@5 from 32.29% to a corrected 87.50%, converting an apparent frontier-model weakness into near-saturation. The audit argues that a reported score is a joint property of model, item bank and grader, and that expert adjudication should precede any capability claim drawn from a benchmark. ([[frontier-models-physics-benchmark-audit-2026]]) The velocity of the systems being evaluated is a further constraint. [[ai-tutoring-micro-rct-gcse-science-2026|Harrison et al. (2026)]] describe the temporal problem directly: by the time a large-scale trial has been designed, delivered, analyzed and published, the technology under study may have changed materially, which pushes practice toward weak observational or usage data at exactly the moment stronger evidence is needed. Their answer is not to accept weaker designs but to shorten the loop — practitioner-led micro-randomized trials that retain the causal contrast and repeat it as the platform evolves. ### Connections to related concepts AI-ed evaluation sits at the center of the knowledge base's methods and risks. It operationalizes [[assessment-validity]], [[educational-measurement]], and [[benchmark]] within [[assessment]] and [[automated-assessment]]. Its call for human oversight connects to [[human-in-the-loop-ai]] and [[teacher-role]], while its focus on reliability connects to [[hallucination-risk]], [[automated-assessment|Confidence Aware AI Assessment]], and [[trust-calibration]]. The distinction between evaluating performance and evaluating learning links to [[genai-performance-vs-learning|performance vs. learning]] and to [[student-modeling]]; and evaluation of pedagogical agents connects to [[intelligent-tutoring]], [[pedagogical-llm-training]], and [[pedagogical-safety]]. ### Evaluating learning gains A central object of AI-ed evaluation is the **learning gain** — the measurable improvement in knowledge or skill an AI tool produces (see [[learning-gains]]). Evaluating gains rigorously requires choosing the right outcome measure, because [[genai-performance-vs-learning|performance and learning diverge]]: AI can inflate immediate, AI-assisted task performance while leaving durable, unassisted learning unchanged or reduced (see [[generative-ai-reduced-study-time-math]], [[stromberg-generative-ai-learning-penalty-secondary-2026]]). Effective gain evaluation therefore: - **Uses unassisted, AI-resistant outcome measures.** [[generative-ai-guardrails-harm-learning|Guardrail evidence]] and [[summative-assessment|summative-assessment research]] show that proctored, closed-book, unassisted measures — not AI-assisted homework or take-home work — reveal genuine [[learning-gains|learning gains]]. - **Distinguishes assisted performance from durable learning.** [[genai-meta-analysis-programming-learning|Meta-analysis]] shows AI can produce large productivity gains with no significant learning gain (g ≈ 0), so evaluations must report both. - **Pairs pre/post measures with validity checks.** [[assessment-validity]] and [[educational-measurement]] ground gain measurement; [[genai-educational-outcomes-meta-analysis|meta-analytic review]] pools effect sizes across studies to establish the field's gain evidence. - **Disaggregates by learner and context.** Because [[learning-gains]] vary by population, domain, and AI configuration, evaluation should report gains for different student subgroups (e.g., by prior proficiency, as [[drawedumath-vlm-struggling-students-2026|VLM evaluations]] reveal for error status) rather than a single aggregate, and should connect gain findings to [[meta-analysis-systematic-review]] to situate them in the wider evidence base. Context-conditioned benchmarks are needed: [[zhang-tutormoments-2026|Zhang et al. (2026)]] argue that prior tutoring benchmarks (MathTutorBench, MRBench, LearnLM) reward one side of the assistance dilemma or give underspecified guidance. TutorMoments instead replays teacher-identified pedagogical decision points, evaluating whether a tutor's help is appropriate to the specific learning moment — [[scaffolding]] vs. rigor. - **The metric you choose can reverse your conclusions.** [[zhang-platform-scores-miss-ai-teaching-agents-2026|Zhang et al. (2026)]], evaluating AI teaching agents in [[medical-education|medical education]], found that an [[edtech-platform|educational platform]]'s undisclosed aggregate scores ranked agents nearly opposite to a transparent, expert-validated 8-dimension teaching-quality rubric (medical knowledge accuracy, pedagogical guidance, knowledge coverage, role-play quality, adaptive difficulty, medical safety, engagement, feedback). Platform scores index student performance; the rubric indexes agent teaching behavior -- choosing the wrong metric determines which agents get adopted or refined. They also found LLM-as-evaluator leniency differs by model (some too lenient to discriminate), so automated scoring needs human calibration and is most trustworthy on cognitive-process dimensions. ## Connected Concepts - [[interpreting-and-applying-aied-research]] - [[assessment-validity]] — Validity of interpretation in AI-ed evaluation - [[educational-measurement]] — Measurement theory for assessing learning - [[psychometrically-aware-ai]] — Applying psychometrics to AI-based assessment - [[benchmark]] — Standardized benchmarks for evaluating AI systems - [[automated-assessment]] — AI-based grading and scoring systems - [[learning-analytics]] — Data-driven analysis of learning behavior - [[research-methods-aied]] — Research methods for AI in education - [[learning-gains]] — Measuring learning gains from AI tools - [[formative-assessment]] — Ongoing assessment to guide instruction - [[summative-assessment]] — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams) - [[authentic-assessment]] — Assessment of real-world, transferable performance - [[human-in-the-loop-ai]] — Human oversight of AI evaluation - [[trust-calibration]] — Calibrating trust in AI systems - [[hallucination-risk]] — Risk of fabricated content in AI outputs - [[intelligent-tutoring]] — Evaluating AI tutoring systems - [[agentic-ai]] — Evaluating autonomous AI agent behavior ## Connected Articles - [[zhang-platform-scores-miss-ai-teaching-agents-2026]] — What platform scores miss: multidimensional evaluation of AI teaching agents - [[assessment-latent-structure-human-llm-2026]] — Do assessment instruments measure the same thing for humans and LLMs? (Strugatski et al. 2026) - [[assessing-quality-ai-generated-exams-field-2025]] — Assessing the quality of AI-generated exams: a large-scale field study - [[nspa-neuro-symbolic-pedagogical-alignment-2026]] — Neuro-symbolic pedagogical alignment (NSPA) - [[yasir-llm-tutoring-agents-2026]] — Three-way classification benchmark of LLM tutoring agents (Yasir et al. 2026) - [[drawedumath-vlm-struggling-students-2026]] — Evaluating VLMs on DrawEduMath: error content hardest (Lucy et al. 2026) - [[cdpk-pedagogy-benchmark-llms]] — Benchmarking LLM pedagogical knowledge (CDPK + SEND) - [[melo-llm-classroom-observation-teach-2026]] — LLM classroom observation validation: reliability vs accuracy (Melo et al. 2026) - [[shen-sustainable-ai-knowledge-base-cs-education-2026]] — On-premise OER AI knowledge-base assistants: multi-dimensional evaluation - [[ground-truth-reliability-aied]] — Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity - [[calibrating-trustworthiness-llm-education-2026]] — Calibrating Trustworthiness: Co-Designing Metrics and Visualizations - [[teachbench-llm-teaching-evaluation]] — TeachBench: Evaluating LLM Teaching Ability - [[machines-misread-pedagogical-quality]] — Why Machines Misread Pedagogical Quality: Human-Machine Alignment - [[cong-confidence-asag-2026]] — Automatic Short Answer Grading With LLMs - [[cotal-formative-assessment-scoring-2026]] — CoTAL: Human-in-the-Loop Prompt Engineering for Formative Assessment - [[llm-cognitive-diagnosis-handwritten-math]] — Benchmarking LLMs for Diagnosing Students' Cognitive Skills - [[tutoring-effectiveness-index]] — The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality - [[jeon-isd-agent-bench-2026]] — ISD Agent Benchmark - [[tool-invariant-framework-agentic-ai]] — A Tool-Invariant Framework for Teaching and Assessing Computational Methods - [[valid-student-simulation-llm-2026]] — Toward Valid Student Simulation With Large Language Models - [[llm-difficulty-calibration-programming-exams-2026]] — From Evaluated Models to Evaluation Aids - [[socratic-tests-conversational-assessment]] — The Theoretical Foundation of Socratic Tests - [[responsible-assessment-ai-era-stanford-2026]] — Responsible Assessment in the AI Era - [[authentic-products-authenticated-processes-2026]] — From Authentic Products to Authenticated Processes - [[zerkouk-comprehensive-review-its-2025]] — Comprehensive Review of Intelligent Tutoring Systems - [[genai-educational-outcomes-meta-analysis]] — Meta-analysis of generative AI educational outcomes - [[zhang-tutormoments-2026]] — When Help is Unhelpful: evaluating AI tutors for productive struggle - [[elbench-education-llm-benchmark-2026]] — ELBench: education LLM benchmark - [[teaching-monster-pck-benchmark-2026]] — Teaching Monster: PCK benchmark - [[ai-grading-handwritten-physics-2026]] — AI grading of handwritten physics assessments (Olympiad) - [[distilling-self-explaining-lm-learning-analytics-2026]] — Distilling self-explaining LM for learning analytics - [[burneo-can-edtech-close-learning-gaps-2026]] — Meta-analytic evaluation of adaptive + AI EdTech - [[xiong-ai-educational-measurement-review-2026]] — AI's role across scoring, psychometrics, assessment - [[liu-ai-literacy-interventions-meta-analysis-2026]] — Meta-analytic evaluation of AI literacy outcomes - [[ai-tutoring-micro-rct-gcse-science-2026]] — Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomized Trials of an AI Tutoring Platform in GCSE Science - [[proiqa-math-item-quality-assessment-2026]] — ProIQA: Process-Based Math Item Quality Assessment - [[durable-skills-measurement-ai-teammates-2026]] — Toward Scalable Measurement of Durable Skills - [[pivot-generative-video-tutors-stem-2026]] — From Content Generation to Learning Support: Pedagogy-Guided Generative Video Tutors for STEM Learning --- ## [Benchmark](https://edtechdev.github.io/aied/concepts/benchmark/) > **Benchmark** — standardized test suites and evaluation frameworks used to measure AI model performance on educational tasks. Benchmarks enable reproducible comparison across models and approaches, and are essential for evaluating the reliability, fairness, and [[pedagogy|pedagogical]] quality of AI in education systems. ## Questions to Consider - A benchmark is a standardized test suite for measuring AI model performance on educational tasks. Before reading, what do you think most AI benchmarks actually test — and why might that be a different thing from what a good tutor needs? - This page highlights a Pedagogy Benchmark that tests pedagogical knowledge — [[teacher-role|teaching]] strategies, assessment methods, special-education pedagogy — rather than content knowledge. Why might an AI that knows a subject brilliantly still fail at teaching it, and why would a benchmark that ignores pedagogy miss that? - One lesson here is [[research-methods-aied|methodological]]: how you validate a benchmark changes the results dramatically, with naive validation reporting far higher performance than rigorous trial-independent methods. How might a model developer or vendor be tempted to design validation to look good, and how would you spot that? - Benchmark performance often doesn't transfer to real-world utility. Can you think of a scenario where an AI 'wins' a benchmark yet fails in an actual classroom — and what does that gap tell you about relying on benchmark scores alone? - The page notes that benchmark design can itself encode or amplify bias. If a benchmark is made of certain tasks, in certain languages, from certain populations, whose learning does it end up measuring — and whose does it ignore? ## Introduction Benchmarks serve as the evidentiary foundation of [[ai-education|AI in education research]]. They provide standardized datasets, tasks, and metrics that allow researchers to compare models, track progress, and identify failure modes. In the knowledge base's research, benchmarks appear across multiple domains: - **[[cstutorbench-slm-tutors|CSTutorBench]]** evaluates small language models for CS tutoring tasks. - **[[anvil-ai-educational-animations|ANVIL]]** benchmarks AI-generated educational animations against human-created alternatives. - **[[teaching-feedback-classification-benchmark|Teaching feedback benchmarks]]** assess cross-language transfer of [[ai-feedback-quality|feedback quality]] classification. - **[[cdpk-pedagogy-benchmark-llms|The Pedagogy Benchmark (CDPK + SEND)]]** tests pedagogical knowledge — teaching strategies, [[assessment|assessment methods]], and [[special-education|special-education pedagogy]] — rather than content knowledge, and reports a cost-vs-accuracy "value frontier" across 97 models (most general benchmarks test content knowledge; pedagogy is a distinct, education-critical dimension). - **[[jeon-isd-agent-bench-2026|ISD-Agent-Bench]]** benchmarks [[llm]]-based [[learning-design|instructional-design]] agents across 25,795 instructional-design scenarios, showing that hybrid agents grounded in classical ISD frameworks (ADDIE, Dick & Carey, Rapid Prototyping) outperform pure theory or pure technique — a benchmark result with direct implications for [[agentic-ai|agentic AI]] design in education. ### Why benchmarks matter in AIED Benchmarks connect to [[ai-ed-evaluation]] and [[assessment-validity]] — without rigorous benchmarks, claims about [[intelligent-tutoring|AI tutoring]] effectiveness are unverifiable. They also intersect with [[bias-mitigation]], as benchmark design can encode or amplify biases. The tension between benchmark performance and real-world utility is explored across multiple articles, connecting to [[transfer-of-learning]] concerns in [[generative-ai]] applications. - **When the metric, not the system, is the failure:** [[algorag-rag-theoretical-cs-education-2026|AlgoRAG]] scored BLEU-4 = 0.0000 on all 179 theoretical [[cs-education|computer science]] exam questions while a six-criterion pedagogical rubric gave 0.7620, because logically equivalent proofs routinely differ in notation, variable names and proof strategy. A zero score says more about n-gram overlap than about answer quality, which is the general case against treating surface metrics as the headline number for formal-domain AIED systems. - **Construct-level counterfactual benchmarks.** CFES-P24 expresses multimedia-learning principles as deterministic, reversible slide transformations to audit whether MLLMs respond to specific instructional-design constructs rather than producing plausible holistic ratings. A frozen pilot showed construct recognition (operation, principle, repair, evidence localization) at 8/8 while comparative judgment (direction 6/8) and severity calibration (0/8) failed — arguing for layered scorecards over composite scores.([[cfes-p24-multimodal-slide-auditing-2026]]) - **Trial-independent evaluation in physiological benchmarks.** [[eeg-familiarity-automated-assessment-2026|Nanayakkara & Halloluwa (2026)]] benchmark 15 ML/DL models for EEG-based familiarity prediction and show that the choice of validation scheme changes headline results dramatically: standard stratified cross-validation allows temporal leakage and reports up to 0.9853 F1, while trial-independent Group K-Fold validation drops the peak to 0.6038 F1. The lesson — temporal/leakage-aware evaluation is essential for credible educational benchmarks — extends beyond EEG to any benchmark using sequential or time-structured data. - **Synthetic benchmarks for AI tutoring.** Open, reproducible datasets for evaluating AI tutoring remain scarce. ASTRA (Adaptive Socially-intelligent Team Reasoning Agents) is a multi-agent tutoring prototype and benchmark framework for studying collaborative programming with socially differentiated agents, supporting alone-tutor, pair-tutor, and pair-multiagent configurations (N=540; 360 sessions; 1,440 episodes) with a trace-ready schema for reproducible analysis of interaction, participation balance, and verification. - **Auditing benchmarks is now a research contribution in its own right.** Three 2026 artifacts push benchmark work past leaderboard aggregation. EduFair-Bench holds a simulated student fixed and varies demographic attributes, turning a tutoring benchmark into a fairness audit with turn-level pedagogical metrics ([[edufair-bench-pedagogical-fairness-llm-tutors-2026]]). GeoVAD-Bench diagnoses intermediate visual constructions — perception, auxiliary quality, utilization — rather than final correctness on 600 [[math-education|geometry]] problems ([[geovad-bench-visual-chain-of-thought-geometry-2026]]). Expert re-grading of six [[physics-education|physics]] benchmarks quantified the error such scores carry: 57.20% of audited rejections were item defects, 38.00% grader errors and only 4.80% true model failures ([[frontier-models-physics-benchmark-audit-2026]]). Together they argue that a benchmark score should always be read with its own audited error budget, which is the same discipline [[assessment-validity]] asks of classroom instruments. - **Decoupled annotation and question generation as a construction paradigm.** Most benchmarks build task-specific question–answer pairs per item or image, which makes extending to new tasks expensive, makes data hard to reuse across tasks, and leaves limited control over question form and complexity. MUSE inverts the order: annotate each artwork once into a reusable structured representation of its visual and semantic content, then instantiate 12 tasks from predefined generation rules, so one image yields a multi-view evaluation instance with difficulty and format treated as explicit design variables rather than by-products ([[muse-vlm-artistic-image-benchmark-2026]]). Its correlation evidence is a second argument for the design — the 12 tasks measure related but non-redundant capabilities (Jigsaw Puzzle correlates weakly with most others, ρ = 0.25 to −0.10), and on six external benchmarks general multimodal scores transfer unevenly to artistic educational imagery (BLINK Jigsaw vs. MUSE Jigsaw ρ = −0.20), which is the construct-coverage case against reading any single aggregate score as a proxy for educationally relevant capability ([[muse-vlm-artistic-image-benchmark-2026]]). ## Connected Concepts - [[ai-ed-evaluation]] - [[bias-mitigation]] - [[human-in-the-loop-ai]] - [[formative-assessment]] - [[knowledge-tracing]] - [[generative-ai]] - [[automated-essay-scoring]] ## Connected Articles - [[omniphys-multimodal-physics-benchmark-2026]] - [[assessment-latent-structure-human-llm-2026]] — Do assessment instruments measure the same thing for humans and LLMs? (Strugatski et al. 2026) - [[cdpk-pedagogy-benchmark-llms]] — The Pedagogy Benchmark: LLM pedagogical knowledge (CDPK + SEND) - [[jeon-isd-agent-bench-2026]] — ISD-Agent-Bench: benchmarking LLM-based instructional-design agents - [[shen-sustainable-ai-knowledge-base-cs-education-2026]] — On-premise OER AI knowledge-base assistants: multi-dimensional benchmark - [[authentic-products-authenticated-processes-2026]] — From authentic products to authenticated processes: authentic assessment in AI-rich higher education - [[llm-cognitive-diagnosis-handwritten-math]] — Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work - [[educlaw-bench-pedagogical-llm-agents-2026]] — EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners - [[responsible-assessment-ai-era-stanford-2026]] — Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference - [[anvil-ai-educational-animations]] — ANVIL: Analogies and Videos for Lecturers - [[icle-plus-plus-essay-scoring]] — ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring - [[elbench-education-llm-benchmark-2026]] - [[teaching-monster-pck-benchmark-2026]] - [[cfes-p24-multimodal-slide-auditing-2026]] — CFES-P24: Benchmarking Multimodal LLMs for Slide Auditing - [[diagramir-educational-math-diagram-evaluation]] — DiagramIR: benchmark for evaluating generated math diagrams - [[eeg-familiarity-automated-assessment-2026]] — Automating Learner Assessment: EEG-Based Familiarity Prediction - [[distilling-self-explaining-lm-learning-analytics-2026]] — Distilling self-explaining LM for learning analytics - [[astra-multi-agent-tutoring-benchmark-2026]] — ASTRA synthetic benchmark for multi-agent tutoring and participation-balanced collaboration - [[algorag-rag-theoretical-cs-education-2026]] — AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory - [[muse-vlm-artistic-image-benchmark-2026]] — MUSE: annotation-first, task-generative benchmark construction, and dimension-level non-redundancy across 12 artistic-imagery tasks (Zhu et al. 2026) - [[mental-health-literacy-students-llms-2026]] — Mental Health Literacy Across Psychology Students and Large Language Models --- ## [Administrators](https://edtechdev.github.io/aied/concepts/administrator/) > **Administrators** — the institutional, leadership, and decision-making view of AI adoption, strategy, and governance in education. Administrators and institutional leaders shape whether and how AI is adopted — through policy, funding, infrastructure, and the strategic framing of AI's role — and must weigh competing concerns about learning, [[equity-in-ai-education|equity]], risk, and organizational capacity. ## Questions to Consider - AI adoption is often framed as a classroom decision, but administrators set the conditions — policy, funding, infrastructure. In your institution, who actually decides whether and how AI gets used, and who is left out of that decision? - Research finds a recurring gap between high-level policy ambition and day-to-day implementation. Where have you seen a well-intentioned AI policy fail to reach the classroom, and why do you think it fell short? - Administrators must balance pedagogical opportunity against risk, equity, and resourcing. If you had to choose between funding one new AI initiative and shoring up data governance and faculty training, which would you pick and why? - Procurement decisions made at the institutional level can enable or constrain what teachers and students can actually do. What questions would you want answered before your institution signs a contract for an AI edtech platform? - Institutions translate AI capability into acceptable-use frameworks and data standards. How do you weigh the promise of innovation against the privacy and equity implications of the data these tools collect? - Administrator choices about AI well-being tools directly shape student experience. What evidence would you want to see before deploying an AI tool meant to support — rather than monitor — students? ## Introduction AI adoption in education is not purely a classroom decision; it is also an institutional one. Administrators — provosts, deans, CIOs, and institutional leaders — set the conditions under which faculty and students use AI, balancing pedagogical opportunity against governance, resourcing, and risk. ### How the administrator perspective appears in the research - **Policy and institutional decision-making:** [[ai-uk-higher-education-policy-2026|UK higher-education AI policy research]] finds that AI integration is accelerating but fragmented, with a gap between high-level policy ambition and institutional implementation — a recurring theme for administrators navigating strategy without clear operational guidance. - **Well-being and student experience:** [[ai-campus-wellbeing-tools|AI campus well-being tools]] examine how institutions deploy AI for student support, linking administrator choices to [[student-experience]] outcomes. - **Governance and regulation:** Administrator decisions interact with [[educational-policy-ai]], [[governance]], and [[regulation]] — institutions translate AI capability into acceptable-use frameworks, assessment rules, and data-governance standards (see [[genai-policies-higher-ed-computing|institutional GenAI policy analysis]]). ### Connections The administrator perspective connects to [[educational-policy-ai]] (policy formation), [[governance]] and [[regulation]] (governance), [[edtech-platform]] (procurement and infrastructure), [[learning-analytics]] (institutional data), and [[higher-ed]]. Administrator decisions enable or constrain the [[educational-development]], [[teacher-role]], and [[ai-literacy]] work covered elsewhere in the knowledge base. ## Connected Concepts - [[educational-policy-ai]] - [[governance]] - [[regulation]] - [[higher-ed]] - [[edtech-platform]] - [[learning-analytics]] - [[privacy]] - [[llm]] - [[generative-ai]] - [[student-experience]] - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) - [[educational-development]] - [[teacher-role]] - [[ai-literacy]] ## Connected Articles - [[sposato-ai-educational-leadership-taxonomy-2025]] — AI in educational leadership: comprehensive taxonomy - [[baroudi-anticipatory-governance-ai-higher-ed-2026]] — Anticipatory governance and leadership for AI - [[alrahmi-org-drivers-ai-adoption-he-2026]] - [[ai-uk-higher-education-policy-2026]] — AI in UK higher-education policy and institutional decision-making - [[ai-campus-wellbeing-tools]] — AI-driven tools for campus well-being - [[genai-policies-higher-ed-computing]] — Institutional GenAI policy in computing - [[institutional-change-framework-ai]] — Institutional change framework for AI --- ## [Educational Technology Developers](https://edtechdev.github.io/aied/concepts/educational-technology-developers/) > **Educational Technology Developers** — the people and organizations that build educational technology: product designers, software developers, learning engineers, learning-analytics designers, and the edtech companies, university labs and [[open-source]] projects they work in. In AI in education this is the role that turns a model capability into something a teacher or learner can actually use, and it carries decisions no later stage can undo: what evidence a design claim rests on, how far a [[learning-analytics|analytics]] pipeline or a [[intelligent-tutoring|tutoring]] system is grounded in the institution's own licensed material, whether teachers and learners are included in design, which [[learning-design|instructional design]] assumptions are baked into the defaults, and what happens to the product after the funding stops. Across the knowledge base's system reports and deployment studies, the recurring lesson is that the deployment context, not the model, is usually the binding constraint. ## Questions to Consider - If the meta-analyses claiming that "AI improves learning" rest on invalid methodology, as the audit in [[oneill-presumed-effective-meta-analysis-2026]] found, what evidence is a product roadmap actually allowed to build on? - Should an algorithm's explanations be written in the teacher's curricular language even when that costs more design effort than exposing feature importances — and who pays for that effort? - When a tool is co-designed with students, whose verdict decides: measured learning gains, or the 96% who said they wanted it kept? - Is on-premise, open-licensed deployment a technical choice or a governance one — and should transparency requirements become a condition of purchase? - What does a developer owe an institution when the grant ends: a maintained product, a forkable repository, or a candid statement that the system was never a validated intervention? ## Introduction The developer sits one level below the platform. [[edtech-platform]] describes the deployed system and the stakeholder it becomes once it is in a school or university; this page is about the people who decide what that system does. The distinction matters because platform-level findings — low take-up, equity skew, procurement friction — are usually consequences of design choices made earlier, by someone who never met the learners. The role is also distinct from its neighbours. [[learning-design]] and [[curriculum-design]] design a course for a known cohort; a technology developer designs a product that many courses, taught by people they have never met, will use — which is why defaults, configurability and documentation carry pedagogical weight. [[educational-development]] supports the teaching staff of an institution from inside it; developers sit outside or alongside, supplying the tools those staff are then asked to adopt. And [[design-based-research]] is the evidence standard such developers are increasingly asked to meet: iterative, contextual, and reported with its own limitations. ### Who builds educational AI **Research labs building public infrastructure.** [[oatutor-open-source-adaptive-tutor-2023|OATutor]] was built at UC Berkeley as the first fully open-source adaptive tutoring system on [[intelligent-tutoring|ITS]] principles: an MIT-licensed codebase with a Creative Commons algebra library, [[knowledge-tracing|Bayesian Knowledge Tracing]] mastery estimation, A/B testing infrastructure and LTI support. Its reason for existing is a design decision — proprietary platforms had confined [[adaptive-learning]] research to closed systems — and its authoring route is another: 16 creators produced a College Algebra course in six months after 2.27 hours of training. **Model builders.** [[learnlm-improving-gemini-learning]] reframes improving a model for learning as [[prompt-engineering|pedagogical instruction following]]: behavior is set per application through system instructions rather than one fixed definition of [[pedagogy]], and expert reviewers preferred it by +31% over GPT-4o and +13% over base Gemini. The practical point is that pedagogy is too context-dependent to define globally; the useful capability is adherence to the instructions a developer writes, measured by conversation-level scenarios rather than single-turn [[benchmark|benchmarks]]. **Architects of knowledge models.** [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] argues [[personalized-learning]] needs more than a static ontology, proposing systems of mapped ontologies plus rules and analytics in place of the classic four-model ITS architecture, and a reuse framework of eight metadata classes that aims to cut the cost of each new build. **Infrastructure and measurement engineers.** [[a4l-analytics-pipeline]] describes a modular, domain-agnostic pipeline for learner interaction data validated across three educational AI assistants, where methods built for one domain extended to another — reusable [[learning-analytics]] infrastructure rather than a one-course dashboard. [[stanbkt-bayesian-knowledge-tracing]] shows the complementary case: a Bayesian reimplementation produced *identical* prediction to the established point-estimate tool (AUC 0.711), differing only in cost and in credible intervals that make a condition comparison interpretable. **Builders inside institutions.** [[moodle-ai-tutoring-deep-learning]] embeds LLM tutoring in an existing LMS rather than shipping a standalone tool, lowering the adoption threshold the ITS literature names as a reason systems fail in practice. [[savvy-student-attention-video-learning]] turns multimodal attention signals into an interface teachers can read before releasing a video. [[instructional-agents-multi-agent-course-gen|Instructional Agents]] automates ADDIE's first three phases with role-specialized agents, and its ablation is a design lesson: the single-agent baseline scored worst, Full Co-Pilot beat Autonomous by 0.5–0.9 points, and no quality difference between backends made the cheapest the default. ### What a design claim can rest on **The evidence base is weaker than it looks.** [[oneill-presumed-effective-meta-analysis-2026]] audited 14 meta-analyses claiming that AI improves education and found none justified its claims: all but two defined the treatment as a tool rather than a pedagogical intervention, 61% of 59 vetted primary studies had validity problems, heterogeneity was high wherever reported, moderator analyses were underpowered, and publication bias was never validly assessed. A retracted meta-analysis was still cited as authoritative by 60% of sampled later papers. For a developer, "AI improves learning" is a product-category claim, not a design input. **Report uncertainty and full cost, not just accuracy.** For [[stanbkt-bayesian-knowledge-tracing|StanBKT]], Bayesian inference buys nothing in prediction and everything in being able to say which effects were credible. [[shen-sustainable-ai-knowledge-base-cs-education-2026]] reports retrieval ablations, quantization-aware fine-tuning, VRAM, energy per query and hallucination measured against retrieved open resources — with the authors' own caveat that the system is not a validated tutor. That is the [[ai-ed-evaluation|evaluation]] discipline that makes a deployment claim checkable. ### Co-design with teachers and learners **Explanations must speak the teacher's language.** [[xai-teachers-trust-edtech-recommendations-2026]] ran a within-subject experiment with 41 chemistry teachers on an ML recommendation tool: understandability, [[trust]] and acceptance correlated positively, and domain-driven explanations in curricular language produced significantly higher understandability, learned [[trust-calibration|trust]] and acceptance than feature-importance explanations. Trust was also dynamic — several teachers said only classroom experience would settle it — and acceptance depended on pedagogical alignment and workload reduction. **Co-design at institutional scale.** [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of|AIDA]] at the Open University was built through six design-based studies over 18 months with 498 students and 20 staff. About 20% were initially skeptical; after hands-on use 96% wanted it kept, and an exploratory RCT found twice the usage time but no significant differences on learning process data. The enabling factors were organizational — senior sponsorship, cross-unit collaboration, data-informed iteration — with gaps in systems-thinking capacity. ### Procurement, openness and after the funding stops [[shen-sustainable-ai-knowledge-base-cs-education-2026]] supplies the inputs a procurement decision needs: a 12 GB VRAM hardware floor, an accuracy ceiling for a 7B-class model, per-query energy, and an ordering of choices — retrieval first (without it the model scored 52.3%, below a TF-IDF baseline), then fine-tuning, then quantization-aware compression; open licensing is the precondition for serving a corpus locally. [[reclaiming-epistemic-agency-co-agency-2026]] frames the same decision as governance: transparency requirements turn purchasing into epistemological governance, contestability and provenance become conditions, and districts with the least capacity face the highest bar. [[credential-cognitive-stewardship-ai-assessment]] adds that vendor governance appeared in only 29% of 30 audited policy packages, which specified what AI may do far more readily than what evidence of learning remained. [[genai-mindtool-generative-learning]] poses the design question — does the product encourage learning *with* the tool or offload cognitive work — and [[vocabulary-difficulty-prediction]] shows the trade in miniature: the top-scoring black-box model (r > 0.91) was less explainable than the interpretable one (r > 0.77). ## Connected Concepts - [[edtech-platform]] - [[learning-design]] - [[curriculum-design]] - [[design-based-research]] - [[educational-development]] - [[open-source]] - [[learning-analytics]] - [[intelligent-tutoring]] - [[human-in-the-loop-ai]] - [[human-ai-collaboration]] - [[teacher-ai-competency]] - [[technology-acceptance-model]] - [[universal-design-for-learning]] - [[assessment-validity]] - [[ai-ed-evaluation]] - [[governance]] - [[educational-policy-ai]] - [[sustainability]] - [[privacy]] ## Connected Articles - [[oatutor-open-source-adaptive-tutor-2023]] - [[moodle-ai-tutoring-deep-learning]] - [[learnlm-improving-gemini-learning]] - [[savvy-student-attention-video-learning]] - [[xai-teachers-trust-edtech-recommendations-2026]] - [[oneill-presumed-effective-meta-analysis-2026]] - [[a4l-analytics-pipeline]] - [[stanbkt-bayesian-knowledge-tracing]] - [[instructional-agents-multi-agent-course-gen]] - [[ontology-layered-hybrid-knowledge-model-personalized-elearning-2026]] - [[credential-cognitive-stewardship-ai-assessment]] - [[reclaiming-epistemic-agency-co-agency-2026]] - [[genai-mindtool-generative-learning]] - [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of]] - [[vocabulary-difficulty-prediction]] - [[shen-sustainable-ai-knowledge-base-cs-education-2026]] --- ## [Learners](https://edtechdev.github.io/aied/concepts/learners/) > **Synthesis:** Learners are the primary audience of [[ai-education|AI in education]] — the students in [[k-12]], [[higher-ed]] and [[adult-learning]] whose work, understanding, and sense of self AI now shapes. This page is the umbrella for the knowledge base's learner-side coverage: what learners experience ([[student-experience]]), who they are becoming ([[learner-identity]]), what they still choose ([[agency]]), how they actually interact with the tool ([[student-ai-interaction]]), whether their effort and [[self-regulated-learning|self-regulation]] hold up under it ([[cognitive-offloading]]), and how AI systems model them ([[student-modeling]]). The recurring finding across that research is that the same tool helps and harms different learners differently — gains concentrate where [[prior-knowledge|prior knowledge]], [[ai-literacy|AI literacy]] and verification habits are already present, and reversal concentrates where AI substitutes for the thinking the task was meant to build. ## Questions to Consider - Which learners in your own context would a new AI tool reach first, and which would it leave behind — and what makes you think so? - The research often reports that students *believe* AI helped them while unassisted measures show no gain, or a loss. If a learner's own account is unreliable evidence, what would you accept as evidence that learning happened? - Learners are described here both as people who use AI and as objects AI models ([[student-modeling|learner models]], [[knowledge-tracing|knowledge tracing]], [[simulating-students|simulated students]]). Where should a model of a learner inform teaching, and where should it stop being trusted? - [[self-regulated-learning|Self-regulation]] and [[prior-knowledge|prior knowledge]] decide whether AI support becomes learning or substitution. Is that a learner deficit to remediate, a design problem to solve, or an assessment problem to fix? - If a group of learners does poorly with an AI tool, the tool's designers are usually the last to find out. What would a routine for hearing from learners before and after deployment actually look like in your institution? - [[agency|Learner agency]] and [[learner-identity|identity]] are at stake alongside achievement. Which of those would you refuse to trade for measurable [[learning-gains|score gains]] — and would your assessment design make that refusal visible? ## Introduction Learners appear in this knowledge base in two very different guises, and blurring them causes most of the confusion in the field. In the first, learners are **people who use AI**: they ask questions, accept or resist answers, lose or keep their footing, and report experiences of support, guilt, anxiety, and dependency. In the second, learners are **objects that AI systems model**: a skill estimate in [[knowledge-tracing|knowledge tracing]], a latent state in [[cognitive-diagnosis|cognitive diagnosis]], or a [[simulating-students|simulated student]] standing in for a real one. The evidence about the first comes from surveys, interviews, log analysis, and experiments; the evidence about the second comes from the machinery of measurement itself — and a model of a learner is a claim, not a learner. This page is the entry point for both. It collects the learner-side concepts the knowledge base covers, explains how they relate, and links the studies behind them. [[stakeholders]] covers the other side of the same field — [[teacher-role|teachers]], [[administrator|administrators]], designers and policymakers — and the two pages are meant to be read together. ## Who counts as a learner The knowledge base treats learners across the full span of formal education: [[k-12]] pupils, [[higher-ed|university students]], [[vocational-education|vocational]] and [[professional-training|professional]] learners, and [[adult-learning|adult]] and [[lifelong-learning|lifelong]] learners, including [[special-education|learners with disabilities]], [[neurodiversity|neurodivergent]] learners, and [[multilingual-learning|multilingual]] learners. Two conventions matter. First, "learners" and "students" are not interchangeable in a useful way: *students* names an institutional role, *learners* names an activity, and a person can be one without the other (an employee in [[professional-training|workplace training]] is a learner but not a student). Second, learners are not a homogeneous group, and the heterogeneity is exactly what the research keeps surfacing — [[prior-knowledge|prior knowledge]], [[self-regulated-learning|self-regulation]], language, access, and disability status all change whether an AI tool helps. ## What learners experience [[student-experience]] is one of the most researched dimensions of AI in education, and its central lesson is that effects are mixed rather than uniform. [[student-perceptions-ai-study-productivity-2026|Survey work on study productivity]] finds learners reporting genuine efficiency gains — 92.3% said AI improved their understanding — alongside a gap the headline numbers hide: only 38.5% said it reduced their overall study time, and half reported sometimes relying on AI instead of trying to learn independently. [[uneven-impact-generative-ai-student-learning-2026|Analyses of reliance patterns]] go further: students with different levels of [[ai-literacy|AI literacy]] and evaluation skill end up in qualitatively different relationships with the tool, so the same course policy produces different outcomes for different learners — support for some, substitution for others. [[genai-student-experiences-uk-he-survey-2026|Students describe the pull of least effort]] in their own words, which the [[academic-integrity]] and [[misconceptions]] pages treat as a design and policy problem rather than a moral one. Affect runs alongside the cognitive story: [[anxiety-and-stress]] and [[well-being]] document anxiety about being outpaced, about being accused of misconduct, and about the value of the degree being pursued. Learners' mental models of AI — what they think it is and what they think it is for — are upstream of whether they use it well, which is why [[misconceptions]] and [[framing-ai-use-for-students]] appear throughout the learner-side research. Learner accounts also complicate the integrity story that frames so much learner-side policy. [[mulisa-students-genai-integrity-perspectives-2026|Interviews with 27 undergraduates at an Ethiopian university]] find near-universal GenAI use alongside a genuinely divided [[ethics|ethical]] reading — most crediting the tools with raising their achievement, a minority calling coursework use misconduct, and almost all reporting an uneven playing field in which AI users score above diligent independent workers, one describing the effect as killing their sense of diligence. The student side of misconduct procedure is thinner in the literature than the student side of use, but [[munoz-misconduct-allegation-evidence-2026|case-file analysis of 1,162 GenAI allegations]] shows what learners face when the institutional response arrives: the evidence most often cited is the weakest-rated kind, no minimum evidentiary threshold governs whether a case progresses, and students whose cases rest on thin evidence are pushed toward appeals. ## Identity, agency, and authorship Learner-side research is not only about outcomes. [[learner-identity]] asks who a learner is becoming in relation to a discipline and to AI, and the evidence runs in both directions: well-designed use can scaffold disciplinary belonging, while outsourcing can erode the sense that the work is one's own. [[agency]] asks what the learner still controls. The knowledge base treats both as genuinely at stake, not as soft adjuncts to achievement: a learner who produces correct output with an AI and no longer recognizes the reasoning behind it has lost something the grade does not record. The authorship question is where learners themselves struggle most: [[mulisa-students-genai-integrity-perspectives-2026|students interviewed about GenAI and integrity]] claimed originality because no other author existed — "If it is not my original idea, then whose?" — while others concluded the work did not represent them, and one reasoned toward recognizing the tool as a co-author it could not be, given that AI is not a person. ## Interacting with AI: what learners actually do The [[student-ai-interaction]] page collects what learners ask of AI, how their prompts and dialogues evolve, and why interaction quality predicts [[learning-gains|learning outcomes]] better than access does. [[student-llm-interaction-taxonomy-review-2026|A rapid scoping review of 46 categorizations drawn from 33 studies]] finds that evidence base conceptually fragmented — studies differ in data source, category scheme and unit of analysis, so "quality interaction" is not yet a comparable construct across them, and the review's call is for a convergent taxonomy of learning-oriented [[llm]] use rather than a claim that one exists. [[student-ai-conversations-cognitive-engagement-2026|Studies of chat content]] find students discipline themselves in characteristic ways, with engagement ranging from probing and testing claims to accepting the first plausible answer. [[help-seeking]] supplies the older frame: asking for help is a skill, and asking the wrong helper in the wrong way is a known failure mode that AI does not remove. ## Effort, self-regulation, and the performance–learning gap This is where the learner-side evidence is most consequential, because it separates what learners *can do with AI* from what they *can do without it*. [[cognitive-offloading]] collects the over-reliance evidence; [[genai-performance-vs-learning]] states the central methodological point that assisted performance and unassisted capability must be measured separately; [[layer-sensitive-cognitive-offloading-writing-2026|Layer-sensitive studies of writing]] separate surface, structural, idea and reasoning offloading, and find the highest supported performance in the least bounded condition alongside the lowest independent performance eight weeks later; and [[shaw-nave-cognitive-surrender-2026|Shaw and Nave's cognitive-surrender account]] names the disposition that makes delegation habitual rather than strategic. [[metacognitively-discordant-completion-genai-2026|Metacognitive discord]] documents the uncomfortable middle case — learners who notice they do not understand and submit anyway — and [[verification-quality-reliance-calibration-genai-2026|verification research]] shows that "checking" is itself a graded skill, not a binary habit. On the design side, [[desirable-difficulties|productive difficulty]] and [[reducing-ai-misuse]] collect the interventions that restore the effort the task was meant to require. ## Learners as models The learner-side concepts with the longest technical lineage are the ones that represent the learner to the system. [[student-modeling]] covers the family: [[knowledge-tracing]] estimating skill acquisition over time, [[cognitive-diagnosis]] locating specific misconceptions, and the adaptive and [[personalized-learning|personalized]] systems that consume those estimates. [[simulating-students]] and [[simulating-students-llm-review-2026|its review]] treat [[simulation]] as a way to test [[intelligent-tutoring|tutors]] and generate data when real learners are unavailable — an explicitly provisional stand-in, not a substitute. Two cautions run through this literature: model estimates are inferences from behavior that are sensitive to how items and interfaces are built, and [[demographic-signals-llm-student-assessment-2026|studies of demographic signals]] show that assessment systems can pick up proxies for learner identity (language, background) that were never intended to be part of the construct. [[self-report-measures]] covers the mirror-image problem on the research side: what learners say about their own learning often diverges from what they can do. ## Equity across learners Because benefits follow prior advantage, learner-side work is inseparable from [[equity-in-ai-education]]. [[digital-divide]] and access to paid model tiers shape who gets the strongest tools; [[bias-mitigation|fairness]] concerns govern how learned models treat different groups; [[inclusive-learning]], [[accessibility]], [[special-education]] and [[neurodiversity]] cover learners whose needs the default design ignores. The practical lesson from this research is that "AI helps students" is not a finding — the finding is always *which* students, under *what* conditions, with *what* prior knowledge and access. Two groups carry a distinctive share of that risk. [[wright-transcription-not-generation-2026|Wright (2026)]] argues that prohibitions written around "[[generative-ai|generative AI]]" rather than around function capture transcription tools that convert the format of work a learner has already authored, so the resulting false positives fall hardest on disabled learners who rely on voice-to-text and OCR, including where AI-powered OCR has replaced discontinued assistive software; [[harerimana-remote-proctoring-nursing-scoping-2026|a scoping review of remote proctoring]] makes the parallel point about assessment conditions, finding that connectivity, data cost and device failure decide who can be assessed at all — an equity finding rather than a technical one, and one concentrated in low- and middle-income settings. ## Where this page sits Read with [[stakeholders]] for the people around the learner, and with [[pedagogy]] and [[learning-design]] for what teachers do with these findings. [[assessment]] determines which learner capabilities are ever made visible; [[limitations-in-aied-research]] explains why so much of the learner-side evidence is short-term, self-reported and conducted on convenience samples; and [[misconceptions]] is the usual entry point for learners themselves. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[student-experience]] — How learners perceive, interact with, and are affected by AI - [[learner-identity]] — Who learners are becoming in relation to a discipline and to AI - [[agency]] — What learners still control and choose - [[student-ai-interaction]] — What learners actually ask AI and how dialogues evolve - [[cognitive-offloading]] — Over-reliance and the substitution of AI for thinking - [[self-regulated-learning]] — Planning, monitoring, and adjusting one's own learning - [[help-seeking]] — Asking for help well, and the failure modes when learners do not - [[metacognition]] — Knowing what one does and does not understand - [[motivation]] — Why learners persist or stop - [[self-efficacy]] — Learners' confidence in their own capability - [[student-engagement]] — Behavioral, emotional, and cognitive engagement - [[prior-knowledge]] — The background knowledge that decides whether support becomes learning - [[student-modeling]] — Representing the learner inside the system - [[knowledge-tracing]] — Estimating skill acquisition over time - [[simulating-students]] — LLM-simulated learners as provisional stand-ins - [[ai-literacy]] — The capability that decides whether learners use AI well - [[equity-in-ai-education]] — Who benefits and who is left out - [[well-being]] — Anxiety, stress, and the affective cost of AI-mediated study - [[misconceptions]] — The mental models learners bring to AI - [[stakeholders]] — The other side of the same field: teachers, leaders, designers - [[assessment]] — Which learner capabilities are ever made visible ## Connected Articles - [[uneven-impact-generative-ai-student-learning-2026]] — Reliance patterns and evaluation literacy split student outcomes - [[student-perceptions-ai-study-productivity-2026]] — Learners report efficiency gains alongside dependency concerns - [[student-llm-interaction-taxonomy-review-2026]] — A taxonomy of learning-oriented student-LLM interaction - [[student-ai-conversations-cognitive-engagement-2026]] — Discipline-specific patterns in student-AI chat - [[layer-sensitive-cognitive-offloading-writing-2026]] — Assisted performance gains without independent capability - [[genai-performance-vs-learning]] — Why assisted performance and unassisted learning must be measured apart - [[shaw-nave-cognitive-surrender-2026]] — Cognitive surrender as a disposition, not an accident - [[metacognitively-discordant-completion-genai-2026]] — Learners who notice they do not understand and submit anyway - [[verification-quality-reliance-calibration-genai-2026]] — Verification quality and reliance calibration - [[simulating-students-llm-review-2026]] — Simulated students: architecture, mechanisms, and limits - [[stanbkt-bayesian-knowledge-tracing]] — Parameter estimation in Bayesian knowledge tracing - [[demographic-signals-llm-student-assessment-2026]] — Implicit and explicit demographic signals in LLM-based assessment - [[ai-literacy-learning-engagement-psych-capital-2026]] — AI literacy, engagement, and psychological capital - [[genai-student-experiences-uk-he-survey-2026]] — Students describe the pull of least effort - [[mulisa-students-genai-integrity-perspectives-2026]] — Students on whether GenAI is a cheating tool or a learning partner - [[munoz-misconduct-allegation-evidence-2026]] — What misconduct allegation files actually contain as evidence - [[wright-transcription-not-generation-2026]] — Over-inclusive AI rules and the students they catch - [[sharma-judgment-visible-genai-assessment-2026]] — Integrity as evaluative judgment rather than compliance - [[harerimana-remote-proctoring-nursing-scoping-2026]] — Remote proctoring's emotional and equity costs for students --- ## [Librarians](https://edtechdev.github.io/aied/concepts/librarians/) > **Librarians** — the library and information professionals who teach information literacy, staff research consultations, build collections and discovery systems, and increasingly sit as the human counterpart to AI search and recommendation tools. In the knowledge base's research they appear as the people who handle the credibility question a model cannot settle — whether a patent, a technical standard, or an industry report deserves to be relied on — and as partners in course and policy co-design. Their place in [[ai-education|AI in education]] is therefore double: as educators who teach source evaluation and search strategy, and as the [[human-in-the-loop-ai|human layer]] inside systems built around AI retrieval. ## Questions to Consider - AI can now answer in seconds the factual reference question a librarian used to field. What part of that work was ever *the answer* — and what part was the judgment about whether a source deserved belief? - In the embedded-librarian study, source-evaluation questions dominated consultations (38.5%) in the very cases where the AI was least confident. Is that a design failure to fix, or evidence that credibility judgment resists automation? - [[genai-academic-search-workshop|The CHIIR 2026 workshop]] reported a trust gap — students [[trust-calibration|over-trust]] GenAI while faculty distrust it. Who closes a gap like that: a librarian, an instructor, or an interface that shows its sources and its confidence? - [[ithaka-sr-ai-skills-college-graduates-2026|Ithaka S+R]] found that institutions lack both a consensus on what AI skills are and a framework for assessing them. Is information-literacy instruction the natural home for that work — and does saying so ask libraries to carry a load their staffing does not match? ## Introduction Librarians are the professionals who sit between a learner's question and the scholarly record: reference and liaison librarians, subject specialists, teaching librarians who deliver information-literacy instruction, and the staff who manage collections and repositories. In [[ai-education|AI in education]] they are named far less often than [[teacher-role|teachers]] and [[administrator|administrators]], yet the corpus gives them a specific, testable role. They are the human answer to the credibility problem that [[generative-ai]] creates, the instructors who teach [[evaluative-judgment]] and search strategy rather than merely supplying sources, and — in the strongest case in the knowledge base — a designed component of an AI system rather than a service beside it. This page treats the role as distinct from its neighbors. [[educational-technology-developers]] build the retrieval and recommendation systems; librarians decide what those systems should not be trusted to conclude. [[academic-integrity]] frames AI use as authorship and disclosure; the librarian's framing — attribution, provenance, access — is adjacent but more concrete. ## Information literacy as instruction, not just assistance The knowledge base's most detailed evidence on this role comes from [[ai-assisted-seminar-learning-information-literacy-2026|Huang's embedded-librarian seminar platform]], built for engineering research teams in a quasi-experiment with 60 students in three programs over eight weeks (30 on the integrated platform, 30 in conventional library instruction). The integrated group gained 0.78 points on a five-point ACRL-based information-literacy instrument against 0.25 for controls, with the largest within-group gains in search skills (+0.87) and source evaluation (+0.83), followed by information synthesis (+0.73) and ethical-usage awareness (+0.70). Notably, performance-based items alone still produced large effects, so the gains are not purely confidence. What was taught maps onto library work: search-strategy formation, source appraisal, citation, and repository navigation — the four categories that later dominated the consultation logs. The wider corpus adds a staffing and delivery question. [[genai-academic-search-workshop|The CHIIR 2026 workshop report]] describes one librarian's 15-week digital literacy [[curriculum-design|curriculum]] built on a survey of 2,076 students and 101 librarians, and records the presenters' judgment that single-session workshops leave students with a misunderstood foundation; the same report counts one librarian's instruction reach at roughly 1,700 first-year students. ## Where the AI was weakest, the librarian carried the load The seminar platform split labor deliberately: algorithms handled retrieval at scale, humans handled evaluation under uncertainty — and the logs show the split working. Across 312 consultations (10.4 per participant), source evaluation and quality accounted for 38.5%, search-strategy formation 29.2%, citation management 18.6%, and repository navigation 13.7%, and 73.4% were answered within six hours. The authors connect this directly to model performance: the [[recommender-systems-and-learning-paths|recommendation engine]] reached precision 0.68 and recall 0.61, up from a Boolean keyword baseline (0.52 and 0.48) but still insufficient for credibility judgments on patents and standards, and natural-language query handling answered 72.8% of engineering-specific queries correctly. Satisfaction followed the same order — librarian assistance highest at 4.3 out of 5.0, ahead of recommendations (4.1), querying (3.9), seminars (3.8), learning paths (3.6), and [[knowledge-graph|knowledge graph]] (3.4) — and interviews recorded 21 of 30 students citing trust in the librarians' presence. Use was heterogeneous — 11 of 30 leaned on AI, 8 on librarians, and 11 used both. The tension cuts both ways. The cases AI handled worst were credibility questions, and the platform's own authors conclude that keeping a human expert at the credibility decision point is the practical lesson. But the study is single-site and non-randomized, its three components were never isolated, and the authors call the results "merely preliminary," noting that persistence after librarian support is withdrawn was never tested. Whether the human layer was the cause of the gains or the safety net around a 0.68-precision recommender remains open. ## Co-design, partnership, and who pays for verification Librarians also appear as co-designers rather than service providers. [[maybee-disruptive-partnerships-sap-2025|Maybee, LeGrand, and Fundator]] document Partners for Algorithmic Literacy at Purdue — a six-week student–faculty learning community facilitated by academic librarians, where undergraduates and faculty co-produce AI course [[educational-policy-ai|policies]] and AI-integrated [[group-work|group projects]]. The authors frame it as an alternative to deficit-based narratives about student [[ai-misuse-learning-harm|misuse of AI]], and the SPIRaL program's student reflections show information literacy being relearned: participants initially equated it with source evaluation, then moved toward reading it as a layered scholarly practice tied to [[agency]]. That shift — from checking sources to judging knowledge production — is the one the librarian is positioned to produce. Two more strands set the limits of the role. [[pearls-epistemic-verification-2026|The PEARLS framework]] names Access, Legitimacy, and Source among its six verification dimensions and notes that verification consumes database access, disciplinary language, and sometimes paid tools — so learners with fewer resources bear a larger burden to prove responsible use, and librarians are named among the parties co-design requires. [[aarc-ai-research-competency-2026|The AARC framework]] makes the instructional goal explicit: verify, cite, and reflect as recurring commitments, assessed through the research process — including detecting fabrication and bias — rather than the product. [[hingle-collaborative-ai-literacy-2025|Hingle and Johri]]'s review shows the same designs libraries already use transfer: group work's benefits for information literacy generalize to [[ai-literacy]] learning, including with [[agentic-ai|AI agents]] as partners. ## Implications for AI in education - **Keep a human at the credibility decision point.** Precision around 0.68 left roughly a third of surfaced material unmatched, and source evaluation was the most-requested consultation type — a design cue for any AI search or recommendation layer. - **Treat information literacy as instruction.** The gains that held up on performance-based items were search strategy and source appraisal, which teaching librarians already teach and which attributing and verifying AI output also requires ([[academic-integrity]], [[critical-thinking]]). - **Design for heterogeneous pathways.** With roughly equal groups preferring AI, librarians, or both, routing consultations by subject and workload absorbed 312 of them at 10.4 per participant; a single mandated route will under-serve some learners. - **Make confidence and sources visible.** The workshop's search-as-learning thread asks which cognitive processes should not be offloaded and calls for interfaces that preserve decision points — a [[scaffolding|scaffold]] librarians can specify from their ordinary working standard of source visibility. - **Budget the human layer, not just the model.** The consultations were answered in hours, at scale, because staff were routed deliberately; assuming learners will handle credibility judgment alone transfers that cost to the people least equipped to pay it. ## Connected Concepts - [[ai-literacy]] - [[evaluative-judgment]] - [[critical-thinking]] - [[academic-integrity]] - [[human-in-the-loop-ai]] - [[human-ai-collaboration]] - [[recommender-systems-and-learning-paths]] - [[knowledge-graph]] - [[scaffolding]] - [[inquiry-based-learning]] - [[research-methods-aied]] - [[higher-ed]] - [[learners]] ## Connected Articles - [[ai-assisted-seminar-learning-information-literacy-2026]] — Embedded-librarian seminar platform; 312 consultations led by source evaluation - [[genai-academic-search-workshop]] — CHIIR 2026 workshop on GenAI and academic search; librarian trust gap - [[maybee-disruptive-partnerships-sap-2025]] — Academic librarians co-designing AI literacy with students as partners - [[aarc-ai-research-competency-2026]] — Verify, cite, reflect as teachable research commitments - [[pearls-epistemic-verification-2026]] — Six-dimension verification protocol; access and legitimacy as gates - [[hingle-collaborative-ai-literacy-2025]] — Collaborative learning for AI literacy; information literacy benefits transfer - [[ithaka-sr-ai-skills-college-graduates-2026]] — Instructors rank attribution and responsible use highest, but teach few of 26 skills - [[bird-multimodal-educational-literature-2026]] — Text-complexity tool built for non-technical users such as English teachers and librarians - [[cognitive-offloading-llm-synthesis-writing]] — What learners offload when AI synthesizes sources for them - [[citation-errors-hallucinations-computing-education-2026]] — Fabricated and erroneous citations as a verification problem --- ## [Parents and Families](https://edtechdev.github.io/aied/concepts/parents-and-families/) > **Parents and families** — the household as a stakeholder in [[ai-education|AI in education]]: the caregivers who tutor their children at home, choose and pay for tools, supervise or fail to notice what those tools do, and receive whatever their school communicates about AI. The page collects the evidence on AI-mediated parent–child [[intelligent-tutoring|tutoring]] and [[conversational-ai|conversational]] companions used in the home, on the handoffs between home and school, on parents' own [[ai-literacy]], and on the [[equity-in-ai-education|equity]] consequences of home devices, connectivity and cost. Its scope is the family as an actor — what families do with AI, what reaches them from school, and what the research does and does not establish about the results. ## Questions to Consider - [[paratutor-parent-child-tutoring|Luo et al. (2026)]] found that generic [[llm]] assistance reduced the parent's tutoring role, while a role-separated interface preserved it. If you were designing a tool for a parent and child to use together, what would you protect the parent from, and what would you hand them? - In [[stromberg-generative-ai-learning-penalty-secondary-2026|Strömberg, Lei and Wu's (2026)]] data, homework scores rose 18% while [[summative-assessment|closed-book exam]] scores fell 20% within six months. What does a parent actually see happening at home, and what would they need to be told to notice the difference? - [[k12-teachers-ai-companion-literacy-2026|Xiao et al. (2026)]] found teachers naming parents as primarily responsible for children's relationships with AI companions while describing those same parents as unaware or overstretched. Is that a reasonable assignment of responsibility, and what would it take to make it real? - [[virtual-tutoring-computer-assisted-learning-takeup-2026|Oreopoulos, Dong and Low (2026)]] nearly doubled first-session attendance by removing family-facing paperwork. In your own school or program, which parts of the home-facing process exist for the family's benefit and which exist for the institution's? - [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] treat family–school coordination as an untested hypothesis rather than an established practice. If it has never been tested, on what evidence does the advice your school currently gives to parents rest? - [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|MacCallum, Parsons and Mohaghegh (2026)]] describe second- and third-level digital divides that persist after devices have been supplied. What would closing the skills and outcomes gaps look like in a household, as opposed to a classroom? ## Introduction A large share of children's time with AI happens outside the classroom, in a setting the research literature mostly treats as background. Work on tutoring and learning outcomes is usually school-based, while work on family guidance is usually advisory rather than tested, so families are simultaneously the target of a great deal of guidance about [[ethics|responsible AI]] use and the subject of very little evidence about whether that guidance works. This page treats the household as an actor in its own right. It covers AI-mediated parent–child tutoring, the companion and reading tools children meet at home, what schools do and do not tell families about AI use, parents' own [[ai-literacy]] and what is known about their views, and the access and cost conditions under which families use these tools. Where the evidence is thin, the page says so rather than filling the gap. ## Parents as co-educators at home The clearest design evidence on home tutoring comes from [[paratutor-parent-child-tutoring|Luo et al. (2026)]], who built ParaTutor for the Chinese home-tutoring context after a [[formative-assessment|formative]] study of how parents help with math word problems. That formative work identified recurring failures: parents struggle to understand the structure of the problem, often lack the content knowledge to support it, and run into communication difficulties that break shared understanding. The design response was role separation — parents received tutoring guidance through [[agentic-ai|multi-agent]] chatbots while children received visual grounding for the problem itself, with phase-gated support to stop dyads jumping to answers. In an evaluation with 23 parent–child dyads (children aged 10–12) across four conditions, generic [[llm]] assistance tended to displace the parent's role, whereas the role-separated interface kept parents delivering and adapting support while children stayed in the reasoning. The authors' implication is context-specific: in [[math-education]], where model accuracy is limited, children should not be interacting with the model independently. [[virtual-tutoring-computer-assisted-learning-takeup-2026|Oreopoulos, Dong and Low (2026)]] treat the household as a procedural actor, and their results are as much about family behavior as about tutoring. Their [[rct|randomized trial]] with the Toronto District School Board placed human virtual tutors over Khan Academy practice for struggling Grade 4–8 students. In Year 1, only about 45% of [[teacher-role|teacher]]-nominated students attended even one session, with families lost at each handoff: invitation, interest, account creation on a [[edtech-platform|third-party platform]], matching, showing up. In Year 2, inviting families directly through their own teacher for a fixed weekly time and handling enrollment and scheduling on the family's behalf raised first-session take-up to 83%. Weekly attendance still hovered near two-fifths, mostly intermittent absence rather than dropout; practice rose about 17 minutes a week, and the intent-to-treat effect on the topic test was 0.055 SD, with report-card marks up about 0.08 SD. Removing friction changed who showed up, not how much anyone sustained. ## Companions and reading tools in the home Families more often meet AI as a companion app than as a tutoring system. [[liao-role-adaptive-ai-companion-book-talk-2026|Liao (2026)]] compared a fixed "student peer" companion (Whisper with GPT-3.5) against experienced homeroom teachers in book-talk sessions with 19 elementary students in Taoyuan, Taiwan — 12 in Grade 4 and seven in Grade 5, four sessions each. Students spent significantly more time with the AI, yet contributed a markedly lower proportion of words and sentences; for Grade 5 the share was less than a third of what they produced with the teacher, a pattern the paper calls conversational dominance. The companion was effective at eliciting factual recall and significantly weaker than the teacher at [[prompt-engineering|prompting]] emotional and future-oriented reflection — an affective ceiling. Liao's response is a role-adaptive framework in which one companion occupies Student Peer, Teacher Assistant and Parent Advisor roles, the last extending discussion into the home. The diagnosis behind it is a support vacuum: most companion designs stay focused on the student–AI dyad and give teachers and parents no defined place. Two early-childhood papers bear on what reaches the home. [[creative-project-approach-ai-early-childhood-2025|Yang, Li and Lee (2025)]] argue that physical, [[embodied-learning|embodied]] agents are developmentally preferable to screens for young children, but warn that generative social robots can fabricate content that preoperational children accept as truth, lack developmental calibration in their [[feedback]], and carry cost and access barriers that widen the [[digital-divide]]. [[ai-play-framework-early-childhood-2026|Malallah et al. (2026)]] take a different route with the same age group: their unplugged, [[game-based-learning|play]]-based AI-Play framework was implemented through a family-centered Hour of Code event, with parent surveys and child reflection sheets used to gauge engagement and [[usability-research|usability]] for at-home replication. They report high engagement and emerging understanding that AI learns from examples. ## Home–school communication and transparency What schools tell families about AI, and what families tell schools, is largely unstudied; the clearest findings concern the absence of a channel. In scenario-based interviews with 33 US [[k-12]] teachers, [[k12-teachers-ai-companion-literacy-2026|Xiao et al. (2026)]] found that decisions to intervene turned on jurisdiction rather than perceived harm. Their visibility test drew the boundary at the school door: use in class or on an assigned tool sat inside the teacher's role, while use at home entered it only through observable effects such as withdrawal, falling grades or isolation — absent those, teachers described watching and informal documentation rather than action. The Parent scenario was named most concerning by 10 of 33 teachers, 17 said a teacher should get involved and 11 answered "it depends". Parents were named first as the party primarily responsible for children's relationships with AI companions, then described as unaware or overstretched. The authors call this a jurisdictional vacancy and warn that without an established curricular claim, the adults with the most consistent access to children default to surveillance and referral. Counselors, parents and students were not interviewed. [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] make the coordination problem explicit. Reviewing the guidance literature, they observe that it pushes risk management onto families and schools while treating their coordination as established, and they formalize three competing models — additive, synergistic and compensatory — noting that the synergistic version embedded in most guidance has no supporting evidence. Their proposed test is a factorial trial contrasting family-only, school-only, coordinated and usual-practice guidance. On the family side they carry over a [[self-determination-theory]] distinction between dependent offloading, where core thinking is delegated, and autonomous offloading, where AI [[scaffolding|scaffolds]] while the learner keeps epistemic [[agency]]; immediate performance is identical in both cases, which is why maladaptive use is hard for a parent to see. The same invisibility appears in achievement data. [[stromberg-generative-ai-learning-penalty-secondary-2026|Strömberg, Lei and Wu (2026)]] followed 26,811 Chinese secondary students in grades 7–12 over 30 months and found that [[generative-ai|generative AI]] adoption raised homework scores by 18% and cut completion time by 30% while lowering closed-book exam scores by 20% within six months, with entrance-exam drops of 18% and 24% emerging only after about two years. About 81% of users showed a homework-outsourcing pattern, while AI users who spent as much homework time as non-users scored similarly on exams. The authors' recommendation is addressed to the adults around the child: monitor inputs such as homework time and effort rather than outputs such as homework scores. They caution that closer parental monitoring is itself one of the unobservables distinguishing the non-outsourcing group. ## Parents' own AI literacy and advocacy Direct evidence on what parents know, fear and ask for is thin, and parents mostly appear in the literature as reported by others. The parents in [[all-girls-genai-makerspace-gender-equity-2026|Liu et al.'s (2026)]] critical case study of an all-girls makerspace were interviewed alongside girls and practitioners, and reported that girls often "got scared" and spoke less in mixed-gender settings. In [[k12-teachers-ai-companion-literacy-2026|Xiao et al. (2026)]], the companion literacy teachers sketched was shared work across counselors, parents, platforms and policymakers, sequenced as an age-graded spiral, with parents named first and then described as already carrying more than they could. Parents appear in the AI-Play evaluation through surveys of activity usability rather than through accounts of their own understanding. What a parent would need in order to take part in any of this is defined, in effect, by AI literacy frameworks built for schools. [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|MacCallum, Parsons and Mohaghegh (2026)]] designed SAIL to be age-agnostic, treating its Level 1 competencies — understanding and exploring AI — as essential for all ages, which places adults in the same progression as children rather than as its instructors. Families also occupy a formal role in research itself: [[raise-framework-ai-education-reporting-2026|Allison (2026)]] asks authors to report participant context and socio-demographic characteristics, [[accessibility]] and cultural fit, and ethical review, consent and data [[governance]] — the obligations through which a family's consent, data and context become visible at all. ## Access and equity across families Cost and connectivity shape what families can use before any question of pedagogy arises. In the makerspace study, [[all-girls-genai-makerspace-gender-equity-2026|Liu et al. (2026)]] document both a cost barrier and a design barrier: the free image generator Playground dropped its free functions for a paid "pro" tier the under-resourced nonprofit could not absorb, and the tools accepted only English prompts, disadvantaging girls without English confidence. [[creative-project-approach-ai-early-childhood-2025|Yang, Li and Lee (2025)]] flag cost and access as widening the [[digital-divide]]. The take-up result in [[virtual-tutoring-computer-assisted-learning-takeup-2026|Oreopoulos, Dong and Low (2026)]] is an equity result as well as a design one: the families lost in Year 1 were lost to administrative steps, and the intervention that recovered them was built around the family's own teacher and a fixed weekly slot. [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] name the measurement conditions that make blanket family advice unsafe: instruments for AI literacy, dependence and [[cognitive-offloading|overreliance]] were largely developed in single-country or specialized university samples, and most studies are cross-sectional and [[self-report-measures|self-reported]]. SAIL's [[equity-in-ai-education|equity]] argument applies in the household register as well: supplying a device addresses only the first of the three divides, and the skills and outcomes gaps persist regardless of who owns the hardware. ## What the evidence does and does not show The strongest family-facing results in this knowledge base concern behavior rather than learning. Reducing household-facing friction raised first-session attendance from about 45% to 83% in one trial, and take-up remained the binding constraint even then. The same trial produced small and imprecise outcome gains, while the largest [[learning-gains|achievement]] study reports large negative effects concentrated among students who outsourced homework, a pattern the authors partly attribute to parental monitoring they could not measure directly. Neither finding isolates what parents do from what schools or tools do. What is missing is more conspicuous. Parents' own accounts are rare: the teachers in [[k12-teachers-ai-companion-literacy-2026|Xiao et al. (2026)]] excluded parents from their sample by design, and the parents who do appear are interviewed as one voice among girls, practitioners or teachers. There is no direct evidence here about how families learn a school's [[educational-policy-ai|AI policy]], what they are told about monitoring or disclosure, or how they respond to it, and family–school coordination has been theorized rather than tested. Home-tutoring designs such as ParaTutor have been evaluated with 23 dyads in one national context, and companion comparisons rest on 19 students in one school. **Relationship to other pages.** [[stakeholders]] is the short overview of the whole stakeholder set, with families listed as an audience not yet covered in depth; this page is that depth. [[early-childhood-elementary-ai-education]] covers the developmental and [[pedagogy|pedagogical]] domain for young learners across school and home; this page follows the family across ages, from preschool play activities to secondary homework. [[ai-use-disclosure]] covers whether learners tell others that they used AI; this page covers the adjacent question of whether that information reaches the adults at home, and what families are told about school AI use. ## Connected Concepts - [[stakeholders]] — the umbrella page for audiences and actors in AI in education - [[early-childhood-elementary-ai-education]] — the developmental domain where home and family use is most studied - [[ai-use-disclosure]] — disclosure practices and the question of what reaches families - [[ai-literacy]] — parents' own competencies and the school-built frameworks that define them - [[equity-in-ai-education]] — how benefits and burdens distribute across households - [[digital-divide]] — access, skills and outcomes layers as they apply at home - [[conversational-ai]] — the dominant form AI takes in family settings - [[self-determination-theory]] — autonomy support as the frame for family guidance - [[cognitive-offloading]] — dependent versus autonomous offloading, and why parents cannot see the difference - [[intelligent-tutoring]] — the model behind AI-mediated parent–child tutoring ## Connected Articles - [[paratutor-parent-child-tutoring]] — role-separated LLM support for 23 parent–child dyads in Chinese home math tutoring - [[liao-role-adaptive-ai-companion-book-talk-2026]] — fixed peer-role companion versus teacher with 19 elementary students in Taiwan - [[family-school-autonomy-support-genai-2026]] — review treating family–school coordination as an untested hypothesis - [[raise-framework-ai-education-reporting-2026]] — reporting obligations covering consent, participant context and equity of access - [[k12-teachers-ai-companion-literacy-2026]] — 33 teachers on jurisdiction, parents and children's relationships with AI companions - [[ai-play-framework-early-childhood-2026]] — unplugged AI literacy tested through a family-centered Hour of Code event - [[all-girls-genai-makerspace-gender-equity-2026]] — parents among interviewees; tool cost and language barriers in informal settings - [[creative-project-approach-ai-early-childhood-2025]] — physical agents for young children, with cost, access and privacy caveats - [[demir-akar-ai-media-literacy-children-2026]] — school program on data privacy, safe communication and media ethics in 36 fourth graders - [[virtual-tutoring-computer-assisted-learning-takeup-2026]] — take-up rose from 45% to 83% when family-facing steps were removed - [[stromberg-generative-ai-learning-penalty-secondary-2026]] — homework outsourcing, exam declines and the case for monitoring inputs - [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl]] — age-agnostic AI literacy levels and the three-level digital divide --- ## [Stakeholders](https://edtechdev.github.io/aied/concepts/stakeholders/) > **Stakeholders** — the range of human stakeholders involved in, affected by, and responsible for [[ai-education|AI in education]], and the umbrella concept for the knowledge base's coverage of who the actors are. [[ai-education|AI in education]] is a multi-stakeholder field: learners who use AI, [[teacher-role|teachers]] and [[educational-development|faculty]] who integrate it, [[administrator|administrators]] who govern it, instructional designers who build learning experiences around it, and policymakers who regulate it. Each audience has distinct needs, competencies, roles, and perspectives, and the knowledge base treats them as the human context in which AI tools are designed, deployed, and evaluated. ## Questions to Consider - The page argues the same AI system looks different from every vantage point — a tool a student experiences as support may look to a teacher like an integrity risk and to an administrator like a governance decision. Which role do you most identify with, and what do you think you're prone to miss from the others? - Before reading on, try to list everyone in your institution who is touched by an AI-in-education decision — beyond just students and teachers. Who did you forget, and what would each of them care about most? - If learners, teachers, administrators, instructional designers, and policymakers each have distinct needs and competencies, who should have the final say over how an AI tool is deployed — and why? - A student's 'personalized support' can simultaneously be a teacher's 'integrity risk.' How would you design a conversation or process that gives each stakeholder's concern genuine weight rather than letting the loudest voice win? ## Introduction [[ai-education|AI in education]] is fundamentally about people — the learners and educators whose work it transforms, and the leaders and designers who decide how it is used. Understanding the distinct stakeholders is essential because the same AI system looks very different from different vantage points: a tool a student experiences as personalized support may appear to a teacher as an integrity risk, to an administrator as a procurement and governance decision, and to a designer as a pedagogical choice. The knowledge base organizes coverage of these audiences across several concept pages. ## The stakeholder landscape - **Learners (students).** The primary audience — students in [[k-12]], [[higher-ed]], and [[adult-learning]]. The knowledge base covers learners through [[student-experience]], [[student-engagement]], [[misconceptions]], [[student-modeling]], [[well-being]], and [[agency]]. Learners' AI [[ai-literacy]], self-regulation ([[self-regulated-learning]]), and risk of [[cognitive-offloading|over-reliance]] are central concerns. - **Teachers and faculty.** Educators who integrate AI into instruction. Covered by [[teacher-role]], [[teacher-ai-competency]], [[teacher-education]], [[educational-development]], and [[tpack]]. Teachers face the dual challenge of using AI in their own teaching and teaching students to use it responsibly (see [[pedagogy|pedagogies and teaching strategies]]). - **Instructional designers and learning technologists.** The professionals who design courses, curricula, and learning experiences around AI. Related to [[learning-design]] (the discipline) and [[curriculum-design]], though the *people/role* of instructional designer is not yet a dedicated page — it is grouped here. - **Administrators and institutional leaders.** Provosts, deans, CIOs, and leaders who set policy, allocate resources, and govern adoption. Covered by [[administrator]], and connected to [[educational-policy-ai]], [[governance]], and [[regulation]]. - **Policymakers and regulators.** Government and institutional bodies that set the legal and regulatory framework. Related to [[educational-policy-ai]], [[regulation]], and [[governance]]. - **[[parents-and-families|Parents and families]].** Present in the [[research-methods-aied|research]] (e.g., monitoring [[student-ai-interaction|student AI use]], attitudes toward AI) but not yet a dedicated page — grouped here as a stakeholder. ## How stakeholders appear in the research - **Role-specific competency frameworks.** [[teacher-ai-competency]] and [[tpack]] define what teachers need to use AI effectively; [[ai-literacy]] defines what all audiences (especially students) need. - **Differential impacts by role.** Research examines how AI affects different audiences differently — [[student-experience]] studies student outcomes, [[teacher-role]] studies pedagogical integration, [[administrator]] studies institutional strategy, and [[educational-development]] studies professional learning. - **Multi-stakeholder governance.** [[governance]] and [[educational-policy-ai]] research emphasizes aligning national, institutional, and classroom stakeholders — policymakers set expectations, administrators implement, teachers adapt, and students experience the result. - **Equity across audiences.** [[equity-in-ai-education]] examines how AI's benefits and harms distribute across learners and institutions, connecting stakeholders to [[bias-mitigation|fairness]] and access. ## Identity across audiences A common thread across these stakeholders is **identity** — the sense of who one is and is becoming in relation to AI and to the domain. The knowledge base treats identity as distributed across audiences rather than belonging to any single group. - **Learner identity** — the evolving disciplinary, professional, creative, and academic identity of students ([[learner-identity]]). It is distinct from, but causally connected to, [[agency]]: agency is the [[situated-learning|situated]] capacity to act, while identity is the durable sense of self that accumulates from agentic acts and is threatened by authorship loss and competence doubt under AI. - **Teacher identity** — the professional self-understanding of educators ([[teacher-role]]), reshaped by AI as a question of purpose and role rather than skills alone (see [[laidlaw-genai-identity-crisis-faculty-2026|GenAI as identity crisis]]). - **Designers and leaders** — professional identity also shapes how instructional designers, [[administrator|administrators]], and policymakers orient to AI, though the knowledge base's explicit identity coverage concentrates on learners and teachers. Identity is the human anchor of the stakeholder landscape: it is what AI must support (not erode) for each audience, and it is the construct that connects otherwise separate role pages — [[learner-identity]], [[agency]], and [[teacher-role]] in particular. Where agency concerns *control* and identity concerns *self*, AI design must preserve both: control over one's learning and a robust, authorial sense of who one is in the domain. ## Implications for AI in education - **Design for the full stakeholder set:** effective [[ai-education|AI in education]] must serve learners, support teachers, inform administrators, and align with policy — not just optimize one audience. - **Build role-specific competencies:** teachers, students, designers, and leaders each need tailored AI literacy and support (see [[ai-literacy]], [[teacher-ai-competency]], [[educational-development]]). - **Align across levels:** the knowledge base's governance research shows AI succeeds when institutional leadership, teacher practice, and student experience are aligned rather than fragmented. - **Consider parents and the broader community:** families are stakeholders in AI adoption whose role and concerns deserve explicit attention. ## Connected Concepts - [[learners]] — Learners: the umbrella for the learner-side concepts - [[teacher-role]] - [[learner-identity]] - [[agency]] - [[teacher-ai-competency]] - [[teacher-education]] - [[educational-development]] - [[tpack]] - [[student-experience]] - [[student-engagement]] - [[misconceptions]] - [[administrator]] - [[learning-design]] - [[curriculum-design]] - [[educational-policy-ai]] - [[governance]] - [[ai-literacy]] - [[equity-in-ai-education]] - [[higher-ed]] - [[k-12]] - [[adult-learning]] - [[parents-and-families]] ## Connected Articles - [[genai-student-experiences-uk-he-survey-2026]] — Student experiences of GenAI in UK higher education - [[ai-uk-higher-education-policy-2026]] — AI in UK higher-education policy (students and institutions) - [[ai-campus-wellbeing-tools]] — AI-driven tools for campus well-being - [[ai-tpack-mathematics-teacher-education-2026]] — AI-TPACK readiness in mathematics teacher education - [[genai-policies-higher-ed-computing]] — Institutional GenAI policy in computing - [[ethical-ai-higher-ed-game-theory]] — Coordination game framework for ethical AI use in higher education - [[student-rationalization-ai-writing]] — Student rationalization of AI use in academic writing - [[ai-changing-teaching-workflows]] — How AI is changing teaching workflows - [[beyond-hype-stakeholder-perceptions-genai-2026]] — Stakeholder perceptions of GenAI in higher ed (Humble & Mozelius 2026) --- ## [Educational AI Policy](https://edtechdev.github.io/aied/concepts/educational-policy-ai/) > **Educational AI policy** — the formal and informal rules governing AI use in educational institutions, from national legislation to classroom guidelines. Policy research in the knowledge base spans institutional governance, [[curriculum-design|curriculum]] mandates, and teacher preparation requirements. ## Questions to Consider - Institutional AI policies are often written as if telling people the rules changes what they do. Research suggests these policies can lag far behind actual AI use. Why do you think a policy 'on paper' so often fails to match what happens in classrooms? - If homework outsourcing causes learning losses that go largely unnoticed — because no single teacher connects a student's decline to AI use — what kind of policy would even be possible? What would monitoring look like? - Assessment policy decisions like oral exams, proctored tests, and closed-book formats are themselves responses to AI-enabled cheating. Do you think policing the 'output' (catching AI use) is a better strategy than redesigning assessment so the work itself is harder to outsource? - Educational AI policy spans national legislation, government guidance, institutional rules, and classroom-level choices. Which of these levels do you think actually shapes student and teacher behavior the most — and why? - Whose voices tend to shape AI policy — and whose are missing? Who should be at the table when an institution decides what AI use is allowed and how it's governed? ## Introduction ## Policy levels - **Institutional policy:** [[genai-policies-higher-ed-computing|Institutional policy analysis]] compares how universities develop AI policies. [[institutional-change-framework-ai|Institutional change frameworks]] provide models for policy development. [[qian-governing-genai-higher-ed-policy-2026|Qian (2026)]] coded the principal AI pages of the 50 US universities ranked most innovative in 2025 and found that only 6 framed them as a "policy" while 44 published guidance, guidelines, principles or resource hubs, and that 33 of 50 positioned instructors as the primary translators of institutional expectations into course rules — institutional policy is being deliberately written as adaptable guidance for course-level interpretation rather than fixed rule. - **Government policy:** and [[ai-lifelong-learning-policy|lifelong learning policy]] examine regulatory approaches at national and regional levels. A critical, infra-level view of AI in government policy-making is [[perrotta-zero-shot-governance-2026|Perrotta (2026)]], whose analysis of the UK Redbox civil-service [[llm]] prototype shows how general-purpose AI enters the professional toolkit of policy through "zero-shot governance" — domain-agnostic foundation models intervening in decisions, wrapped in thin [[discipline-specific-aied|domain-specific]] [[scaffolding|scaffolds]]. - **[[k-12]] policy:** [[stanford-evidence-base-ai-k12-2026|Stanford's evidence reviews]] of 818 papers inform K-12 AI policy. - **Assessment policy:** [[ai-assessment-scale-reform|Assessment reform policies]] and [[authentic-assessment]] frameworks represent policy-level responses to [[academic-integrity|AI-enabled cheating]]. The choice of summative assessment format — oral, proctored, closed-book — is itself an assessment-policy decision (see [[summative-assessment]]). ### Policy maturity gap The knowledge base documents that institutional AI policies [[genai-policies-higher-ed-computing|lag behind actual AI use]]. [[educational-development]] programs, [[teacher-ai-competency]] frameworks, and [[regulation]] all require coherent policy foundations. Large-scale field evidence [[stromberg-generative-ai-learning-penalty-secondary-2026|(Strömberg, Lei, & Wu 2026)]] shows that the learning losses from homework outsourcing go largely unnoticed because individual subject teachers and students rarely connect the decline to AI use — a gap that evidence-informed policy (e.g., weighting closed-book assessment, informing students of long-run costs, monitoring inputs rather than outputs) can address. **Advisory documents outnumber binding policy.** A census of every accredited program in one professional field shows that the maturity gap is one of form as well as timing. [[institutional-ai-policy-health-informatics-2026|Eldredge et al. (2026)]] collected AI policy and guidance documents from all 48 CAHIIM-accredited health informatics and health information management master's programs in the United States and found that 40 (83%) had at least one publicly available AI-related document while 8 had none, but that the documents were mostly guidance rather than enforceable rules: 21 guidelines (53%) against 7 formal policies (18%). Their content centered on academic integrity and acceptable use, with privacy, intellectual property, and regulatory concepts (HIPAA, FERPA, research compliance) appearing far less often, and topic modeling returned the same student-conduct emphasis. Neither document type nor intended audience varied by delivery mode (Fisher's exact P = .85 and P = .71). The authors read this as evidence that academic program policy is a distinct activity from curriculum design and workforce competency development, and argue that accrediting bodies could reduce the resulting variation by providing AI policy frameworks that integrate academic integrity, data ethics, and equitable access. **A sector making policy without an evidence base.** The maturity gap also appears as an evidence problem in professional programs. [[gutowski-hurley-genai-policy-legal-education-2025|Gutowski and Hurley (2025)]] scored the institutional policies of ABA-approved US law schools on five dimensions — prohibitiveness, permissiveness, educational integration, transparency and accountability, and depth — and found most schools taking generally prohibitive positions while reserving discretion to individual instructors and committing to review the rules as the technology moves. They attribute the caution to time pressure rather than evidence: the ABA's 2024 survey drew responses from only about 15% of accredited schools, and the authors report no consensus on whether AI use must be disclosed or how it should be cited. Their conclusion is procedural — build policy with faculty rather than for them, train proactively, and treat policy generation as a recurring event rather than a one-time act — which makes [[legal-education]] one of the clearest cases of the policy-versus-governance distinction this page draws. **Instructors regulate; institutions mostly do not.** The maturity gap has a second face, visible from the teaching side. [[watson-rainie-ai-challenge-faculty-survey-2026|Watson and Rainie (2026)]] surveyed 1,057 US college and university faculty and found that rules have been written far faster below the institution than at it: 87% of respondents created their own assignment-level policies for [[student-ai-interaction|student AI use]], against 35% who said their department has written guidelines and 48% who said their institution has. The structural response behind those documents is thin — 55% reported a task force or oversight group, 37% new AI-focused classes, 17% an AI major or minor, 16% new academic leadership offices, and only 13% adoption of [[ai-literacy|AI literacy]] as a general education outcome — while 68% said their schools had not prepared faculty to teach with the tools. Because students meet the resulting patchwork as a single institutional regime, the finding is direct evidence for this page's distinction between a policy document and the governance that makes it consistent. (The authors describe the sample as non-scientific and not generalizable, so it is best read as the sector's expressed concerns.) **Task-level regulation is the emerging pattern.** A large-scale longitudinal study of 31,000+ course syllabi (2021–2025) at a large public research university ([[chirikov-regulate-ai-syllabi-2026|Chirikov 2026]]) shows how instructors actually regulate AI in practice: explicit AI regulation grew from near zero to 55% of courses by Fall 2025, but the direction shifted from restrictive toward permissive, and instructors increasingly **differentiated by task type** — restricting AI for drafting/reasoning (displacement-risk tasks) while permitting it for editing/proofreading and study support (augmentation tasks). Framing also shifted from academic integrity (63%→49% of syllabi) toward learning impact (1%→29%). This task-based pattern — built on the labor-economics mechanisms of task displacement, augmentation, and reinstatement — offers a more granular alternative to blanket adoption-or-ban policies and is a direct empirical anchor for the policy-vs-governance distinction above. **Instructor-level policy design as an alignment exercise.** Task-level regulation has a design counterpart below the syllabus. [[mccorkle-aligned-genai-course-policy-2025|McCorkle's (2025)]] design case derives a course's allowed and unallowed GenAI uses from "what, specifically, am I assessing?" — inventorying every task in a project, mapping each task to a learning objective, pairing each with an emerging GenAI workforce competency ([[career-development-and-readiness]]), and then deciding which concern takes priority: the need to assess student performance or the value of building the competency. The resulting policy differs task by task inside a single project — brainstorming topics and curating images permitted, composing learning objectives and designing slides not — with the [[framing-ai-use-for-students|rationale written directly to students]]. The case also documents why blanket prohibitions fail in practice: students who do not see themselves as dishonest read a prohibition as not applying to them, and the policy becomes inequitable ([[equity-in-ai-education]]) when expectations are left implicit. It complements the top-down policy levels above by showing how a policy's *content* can be generated from [[assessment]] alignment at the course and assignment level. **The boundary–evidence gap in assessment policy.** A 30-university audit of public [[generative-ai|GenAI]] assessment guidance ([[credential-cognitive-stewardship-ai-assessment|Yao 2026]]) finds that institutional policies are better at *classifying* AI use than at explaining what evidence of learning remains valid under each class: the mean delegation-boundary score (2.47/4) exceeded the mean evidence-standard score (1.89/4), safeguards were sparse (2.75 of 8), and guidance was clearest for final-output substitution. The framework of *cognitive stewardship* argues that policies must make the certification logic visible — what learners may delegate, what they must still demonstrate, and how institutions protect fair evidence — rather than merely monitor AI use. **Over-inclusive prohibition language and misapplied rules.** A rule's scope is a policy decision with consequences beyond compliance. [[wright-transcription-not-generation-2026|Wright (2026)]] argues that the blanket prohibitions written across [[higher-ed|higher education]] since 2023 bar "generative AI" without the technical precision to separate content generation from AI-powered format conversion — optical character recognition, handwritten text recognition and speech-to-text are recognition [[ai-technologies|technologies]] that infer what a student already wrote rather than producing new content. On this reading, sanctioning transcription-only use is best characterized as policy misapplication rather than misconduct, and the cost of imprecision falls unevenly: [[equity-in-ai-education|equity]]-exposed and disabled students who rely on those tools carry disproportionate false-positive risk, an exposure that [[legal-issues-and-risks|legal issues and risks]] treats as a reasonable-adjustment question rather than an integrity one. **A purpose-based boundary instead of a tool list.** [[li-genai-assessment-language-equity-2026|Li (2026)]] supplies the policy content that over-inclusive prohibition language lacks: a support-versus-substitution line defined by function and the assessment construct rather than by naming tools, on the argument that tool lists go obsolete and encourage compliance by interface. Permitted support is a surface intervention that adds no ideas, no sources and no material re-ordering of analysis; substitution is generating arguments or counterarguments, applying disciplinary rules to facts, restructuring an analytical sequence, or generating citations for inclusion. The framework is delivered as three instruments — working definitions with shared quick tests so two markers characterize the same conduct alike, calibrated [[ai-use-disclosure|disclosure]] templates ranging from nothing for embedded low-risk functions to a one-line statement for external language editing and roughly 80–120 words where limited ideation is expressly authorized, and a decision rubric that names weak proxies (polished language or sudden fluency, non-native phrasing, single detection scores) beside the evidence types that can carry weight. It also states the conditions under which scaled permissive regimes such as the AI Assessment Scale remain workable: use expressly authorized for the specific task, the construct re-specified so students and markers know what is assessed, and disclosure kept feasible. **Staged automation with mandatory human oversight.** Uruguay's [[human-in-the-loop-ai-scoring-national-assessment-2026|*Acredita EB*]] national lower-secondary accreditation test offers a concrete model of oversight-preserving automation in assessment policy: multiple-choice sections and existing procedures stay unchanged, AI is used only for the writing section inside a mandatory calibration stage on 50–100 texts, and two pruning rules decide which results require expert review — candidates who cannot pass on the other two sections need no writing review, and candidates Proficient in all three sections pass with only 0.2%/0.6% observed residual risk. The authors estimate the design could cut written responses needing full human scoring by at least 50% while keeping the guiding principle that no final decision about test outcomes is made without appropriate [[human-in-the-loop-ai|human oversight]]. This complements the boundary–evidence work above by showing a policy that specifies not only where AI is permitted but which human evidence and review remain mandatory. **The adversarial side of [[automated-assessment|automated grading]].** A policy that routes grading through an AI tool inherits an attack surface. [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] red-teamed an institutional grading workflow by hiding five indirect prompt injections inside the files of a synthetic submission that Microsoft Copilot (GPT-5.2) had graded as fail on all six baseline runs. Two strategies raised the grade with no visible warning to the user — reported attack success rates of 100% (9 of 9 iterations) and 94% (17 of 18) — and the tool both silently disabled a chat after blocking the simplest attack and, in one iteration, announced it would ignore embedded instructions before raising the grade on the next six runs. A grade obtained through a hidden instruction carries no [[assessment-validity|validity]] claim, and the same technique could degrade a submission with no durable trace, so the paper's asks are policy-level rather than technical: sector-level guidance, professional development, restricted AI use with human review for high-stakes work, and standardized, domain-agnostic testing of prompt-injection resilience so the attack surface is measured rather than assumed. It is the adversarial counterpart to the oversight designs above: where those specify which human review is mandatory, this specifies a class of manipulation that review exists to catch but has no reliable signal for. **A value/norm matrix as a policy foundation.** [[agarwal-ethical-values-norms-aied-2026|Agarwal et al. (2026)]], a [[meta-analysis-systematic-review|systematic review]] of 25 articles, consolidate AIED ethics into six main [[ethics|ethical]] values (non-discrimination, data stewardship, [[human-in-the-loop-ai|human oversight]], goodwill, explicability, educational aptness) and map the ethical norms extracted from the literature onto a stakeholder-by-value matrix. The review positions the matrix as a foundation for building detailed ethical frameworks and regulation for AIED, giving educational institutions, developers, and regulators concrete norms to implement specific values. It finds goodwill norms aimed at regulators are far more numerous (nine) than for any other stakeholder set, signaling regulators' role in ensuring AIED benefits learners through policy and legislation — a concrete, value-anchored starting point for the policy-vs-governance machinery this page describes. **Claims that travel further than their qualifications.** A pilot announcement is a policy instrument too, and its reporting standards are part of the policy. [[el-salvador-ai-tutoring-selection-claim-2026|Restrepo Morales et al. (2026)]] analyze the September 2026 El Salvador episode in which an AI-tutoring pilot in 171 public schools was reported as reaching results comparable to Germany and Sweden: against the German PISA 2022 [[math-education|mathematics]] average, the top 5.5% of the national achievement distribution would have matched the benchmark with no learning gain at all, and the pilot's assessment covered 7.0 students per school against 25.4 in the national survey the same year, so the published evidence cannot exclude selection as the explanation. Their proposal is a six-item reporting standard — the participating schools' baseline, the sampling protocol (eligible students, assessed students, identification rule, participation rate), disaggregated scores with standard errors, the operational date of each reform component, overlap with the representative national sample, and the qualification stated in the same document, post or paragraph as the claim — motivated by the disclosure arithmetic of the announcement itself: the World Bank post carrying the claim recorded about 441,000 views against about 11,000 for the post carrying the qualification four places later in the same thread. The policy lesson is that where a system uses school-level results as evidence about itself, the format of disclosure decides which statement becomes the public fact. ### Policy vs. governance Policy and governance are closely related but distinct, and keeping them apart matters for understanding the knowledge base's research. - **Policy is the *content* — what is decided.** A policy is a formal rule, principle, or statement: what AI use is allowed, prohibited, or required; what must be disclosed; what assessment formats are permitted. It exists as a documented artifact (legislation, institutional guideline, syllabus statement) and answers *"what are the rules?"* - **Governance is the *machinery* — how it is decided, implemented, and enforced.** Governance encompasses the institutional structures, norms, and accountability mechanisms that produce, carry out, and monitor policy: who sets the rules, how they are communicated and resourced, how compliance is enforced and challenged, and how they are revised as AI evolves. It answers *"who decides, and how do the rules actually take effect?"* - **They are interdependent.** Policy without governance is unenforced — a written rule no one owns, monitors, or updates. Governance without policy lacks direction — structures that administer nothing in particular. The knowledge base's research repeatedly shows that the two must be built together: a policy that only *classifies* AI use without governance to specify evidence, safeguards, and revision processes remains weak in practice ([[credential-cognitive-stewardship-ai-assessment|the cognitive-stewardship audit]]), and governance that merely monitors without clear policy risks surveillance without [[bias-mitigation|fairness]] (see [[governance|AI governance]]). The practical test that separates them: a policy can be read on paper, but governance is observed in whether the rule is implemented, enforced, and adapted. This is why [[governance]] extends [[regulation]] and policy into institutions, and why the knowledge base treats assessment-format choices ([[summative-assessment]]) as *policy* decisions that only become effective through *governance* structures like review boards, declaration frameworks, and appeal routes. **Classroom policy as the interpretive frontier of governance.** The machinery of governance does not stop at the institutional document; the teachers who read and adapt it are the last link in the chain, and [[nash-preservice-teachers-classroom-ai-policies-2026|Nash & Burriss (2026)]] show what that link looks like in practice. Their [[teacher-education|preservice teachers]] expected to work in districts with AI policies yet still had to author classroom-specific rules for their own students — an essential adaptation skill — and the authors argue that teachers can decline to adopt AI for specific reading and writing tasks without ignoring it, with principled refusal distinguishable from uncritical rejection. Keeping that distinction viable is itself a governance task: districts and preparation programs need the guidance, [[educational-development|professional development]], and policy infrastructure that let teachers refuse particular uses without being framed as behind the times. **Governance indicators for the authentication problem.** [[coates-governing-academic-integrity-indicators-2025|Coates, Croucher and Calderon (2025)]] supply the measurement layer this policy-versus-governance distinction implies, treating [[academic-integrity]] in the GenAI era as a governance problem before a detection problem and judging contemporary academic governance resilient but "not well positioned or poised" to protect the authentication of student [[assessment]]. Their framework holds 130 governance questions under eight dimensions — Designing, Developing, Training, Implementing, Analyzing, Reporting, Evaluating and Improving — and the items are questions about institutional machinery rather than psychometric scales: whether the institution's top-most board or council receives updates on assessment processes and outcomes, whether key performance indicators cover assessment quality, what percentage of students are known individually by the teachers who assess them, whether extreme low or high marks are cross-checked, and whether there is a simple route for referring contract cheating cases. The accompanying reform program targets governance architectures, the people who hold governance roles, and the technologies and resources supporting assessment, and the authors argue that little institutional development pays out without external affordance from regulation, [[benchmark|benchmarking]] and cross-institutional competition — the same regulatory pressure the [[regulation]] page treats as the binding constraint on governance reform. ### The policy deficit in AI × SEL research A [[meta-analysis-systematic-review|systematic review]] of 65 papers at the intersection of [[ai-education|AI]] and [[social-emotional-learning|social-emotional learning]] ([[policy-deficit-ai-sel-2026|Tran, Liu & Nguyen 2026]]) documents a "policy deficit": nearly three-quarters of studies state no policy implications, and those that do often lack actor-oriented specificity. The review finds policy [[student-engagement|engagement]] correlates with publication venue, reflecting academic incentives that reward technical novelty over [[governance]] and [[regulation]]. It proposes a "WH-question" framework (Who, What, Why, When/Where, How) and a shift from "implication-as-afterthought" to "implication-as-methodology" — treating policy articulation as a design constraint of research rather than a post-hoc add-on. ## Connected Concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[regulation]] - [[governance]] - [[educational-development]] - [[equity-in-ai-education]] - [[higher-ed]] - [[k-12]] - [[academic-integrity]] - [[ethics]] - [[teacher-ai-competency]] - [[framing-ai-use-for-students]] - [[summative-assessment]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) ## Connected Articles - [[nash-preservice-teachers-classroom-ai-policies-2026]] — Preservice English teachers' classroom AI policies: what they permitted, limited, and banned (Nash & Burriss 2026) - [[human-in-the-loop-ai-scoring-national-assessment-2026]] — A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment - [[mccorkle-aligned-genai-course-policy-2025]] — Deriving allowed and unallowed GenAI uses task by task from what is assessed (McCorkle 2025) - [[chirikov-regulate-ai-syllabi-2026]] — How instructors regulate AI across 31,000 course syllabi (Chirikov 2026) - [[chirikov-ai-grade-inflation-2026]] — AI task displacement as a mechanism of grade inflation (Chirikov 2026) - [[ai-adaptation-gap-higher-education-2026]] — The AI Adaptation Gap in Higher Education - [[crompton-governing-genai-higher-ed-delphi-2026]] — Global Delphi on governing generative AI in higher education - [[qian-governing-genai-higher-ed-policy-2026]] — Guidance over policy: instructor-set syllabus rules and a four-unit support ecosystem across 50 innovative US universities (Qian 2026) - [[gutowski-hurley-genai-policy-legal-education-2025]] — Five-factor comparison of law school GenAI policies in a sector making policy without an evidence base (Gutowski & Hurley 2025) - [[wright-transcription-not-generation-2026]] — Transcription is not generation: over-inclusive AI rules, format conversion and disability accommodation (Wright 2026) - [[baroudi-anticipatory-governance-ai-higher-ed-2026]] — Anticipatory governance for AI in higher education (scoping review) - [[institutional-ai-policy-health-informatics-2026]] — AI policy documents across all 48 accredited health informatics master's programs: mostly guidance, centered on academic integrity (Eldredge et al. 2026) - [[credential-cognitive-stewardship-ai-assessment]] — Cognitive stewardship for AI-mediated assessment (30-university policy audit) - [[adarkwah-genai-unesco-policy-2026]] - [[genai-policies-higher-ed-computing]] - [[institutional-change-framework-ai]] - [[ai-assessment-scale-reform]] - [[stanford-evidence-base-ai-k12-2026]] - [[ai-uk-higher-education-policy-2026]] - [[ssaho-ai-academic-integrity-review-2025]] — Call for explicit, co-developed AI-use policies - [[young-people-learning-generative-ai-rapid-review-2026]] — Move beyond adoption-or-ban; staged, developmentally responsive guidance - [[policy-deficit-ai-sel-2026]] — The Policy Deficit in AI × SEL Research - [[agarwal-ethical-values-norms-aied-2026]] — Ethical values and norms for AI in education - [[perrotta-zero-shot-governance-2026]] — Zero-shot governance: general-purpose AI in policy (Perrotta 2026) - [[el-salvador-ai-tutoring-selection-claim-2026]] — Bounding the learning claim of El Salvador's AI tutoring pilot (Restrepo Morales et al. 2026) - [[watson-rainie-ai-challenge-faculty-survey-2026]] — 1,057 US faculty: 87% write their own assignment rules against thin institutional and departmental guidelines (Watson & Rainie 2026) - [[li-genai-assessment-language-equity-2026]] — A purpose-based support–substitution boundary with calibrated disclosure and decision rubrics (Li 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Prompt injection in AI-mediated grading: grades changed undetected, and the policy-level response (Humble 2026) - [[coates-governing-academic-integrity-indicators-2025]] — 130 governance indicators for authenticating assessment, and the external pressure reform needs (Coates, Croucher & Calderon 2025) - [[physics-faculty-learning-community-ai-2026]] — A Workshop Series for Effective Use of AI in Uncertain Times: Building a Physics Faculty Learning Community --- ## [AI Governance](https://edtechdev.github.io/aied/concepts/governance/) > **AI governance** — the frameworks, policies, institutional structures, and norms that guide the responsible design, deployment, and use of [[ai-education|artificial intelligence in education]]. Governance spans formal institutional mechanisms (AI steering groups, policies on academic integrity and acceptable use, ethical review) and informal norms (faculty guidelines, professional development, cultures of [[reducing-ai-misuse|responsible AI use]]). In the AI era, effective governance is a prerequisite for ethical, [[equity-in-ai-education|equitable]], and sustainable adoption of GenAI — it determines whether AI is integrated transparently, with accountability, or adopted reactively in ways that deepen inequities. ## Questions to Consider - If your institution has an AI policy on paper but no clear process for enforcing, updating, or communicating it, is that really governance? What separates a rule that exists from a rule that is actually governed? - Students in one study said their university had no clear AI policy and that this shaped how they used AI. How might policy ambiguity influence what students decide is 'acceptable use' — and should institutions be worried about leaving that negotiation to individuals? - The page distinguishes detection-based governance (policing AI use) from design-based governance (redesigning assessment so AI use is expected and declared). Which approach does your own assessment practice lean toward, and what are the trade-offs of each? - Governance is described as operating at national, institutional, and classroom levels. Think of one AI rule in your context: who set it, how is it communicated and enforced, and how well do those levels align? - One framework reframes AI governance as a collective-action problem of sustaining shared expertise — the 'cognitive commons.' How does protecting a profession's collective knowledge pool change how you think about AI governance versus just regulating a tool? - The page warns that mandatory AI-use declarations fail when they feel punitive or ambiguous. When have you seen a compliance rule backfire because people didn't understand or trust it — and what did that teach you about governance? ## Introduction AI governance in education is increasingly urgent because [[generative-ai|generative AI]] introduces new epistemic, ethical, and organizational challenges: it destabilizes assumptions about knowledge production, [[agency|learner agency]], [[assessment]] validity, and the [[teacher-role|role of educators]] as epistemic authorities. Governance addresses questions of [[academic-integrity|academic integrity]] (what counts as acceptable AI use), [[privacy|data privacy]] and security, [[bias-mitigation|algorithmic bias]] and fairness, [[explainable-ai|transparency]] and accountability, and the alignment of AI adoption with institutional mission and values. A recurring finding across the knowledge base's [[research-methods-aied|research]] is that **institutional governance is often lagging** — many institutions lack clear, unified AI policies, leaving students and faculty to negotiate acceptable use on their own. - **Zero-shot governance as a structural condition of platformisation.** [[perrotta-zero-shot-governance-2026|Perrotta (2026)]] develops the concept of **zero-shot governance** — domain-agnostic [[generative-ai|generative AI]] intervening in policy decisions — through a code-level analysis of **Redbox**, a discontinued UK civil-service prototype built on off-the-shelf [[llm|LLMs]]. Reading its architecture (an invisible system prompt + a thin Python wrapper over a [[rag]] retrieve→format→generate pipeline and a provider-agnostic cloud stack), the article argues that the general-purpose nature of foundation models is a *structural feature* of platformisation that can be mitigated but never ruled out: [[agentic-ai|agentic AI]] does not interrupt the monopolistic, rentier logic of platforms, and the same probabilistic mechanism that generates novelty also produces [[hallucination-risk|hallucinations]]. For governance, the implication is that oversight of general-purpose AI must treat aberrant output as an irreducible, only-mitigable risk rather than a fixable bug — a caution that applies equally to [[educational-policy-ai|education policy]] reasoning. - **[[baroudi-anticipatory-governance-ai-higher-ed-2026|Baroudi]]** [[meta-analysis-systematic-review|scoping review]] frames AI governance in higher education through anticipatory-governance and leadership lenses. - **Student co-design of policy:** [[guided-inquiry-genai-course-policy-2026|Hingle & Johri]] show how a guided inquiry activity in which students co-designed a GenAI course policy surfaced student values — prioritizing training, standardized [[ai-use-disclosure|disclosure]] procedures, stronger institutional support, and greater involvement in decision-making. This positions [[pedagogical-partnerships|students as partners]] in governance rather than passive subjects, complementing institutional-level policy with bottom-up student voice. - **The techno-solutionist trap and the policy deficit:** A systematic review of 65 AI × [[social-emotional-learning|social-emotional learning]] studies ([[policy-deficit-ai-sel-2026|Tran, Liu & Nguyen 2026]]) finds that research often foregrounds technical potential while under-specifying the institutional conditions for responsible implementation — a "techno-solutionist" trap. The authors link policy [[student-engagement|engagement]] to publication venue and propose a "WH-question" framework to make governance implications actor-specific, echoing the knowledge base's broader finding that institutional governance lags AI adoption. ## How AI governance appears in the research - **Institutional adoption at scale:** [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of|The AIDA study at the Open University]] shows how an institution designed, implemented, and evaluated a GenAI assistant, identifying that responsible system-level deployment requires governance structures (AI Steering Group), senior leadership sponsorship, and alignment with institutional strategy — not just technical capability. - **Leadership and systemic change:** [[leveraging-complex-systems-leading-for-transformative-change|SPARK]] frames governance within Complexity Leadership Theory, arguing leaders must balance administrative stability with emergent innovation, embedding governance mechanisms (policies, assessment regimes, accountability frameworks) so adaptive-space innovations can be sustained and scaled. - **Policy ambiguity and student experience:** [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination|Students' engagement with GenAI]] found 12/23 students noted the lack of explicit institutional AI policies ("University doesn't have a clear and unified policy yet"), arguing governance ambiguity shapes students' practices, norms, and [[self-regulated-learning|self-regulation]] — supporting a shift toward transparent institutional guidance. - **AI literacy as a governance capacity ("the 18th SDG"):** [[ai-literacy-sdg-governance-framework-2026|Islam, Morshed, and Islam (2026)]] reconceptualize AI literacy as a systemic governance mechanism rather than a classroom skill, proposing a six-level AIRE Taxonomy (Recognize → Comprehend → Apply → Analyze → Integrate → Govern) that extends Bloom's hierarchy with ethical synthesis and strategic foresight, and an AI–SDG Nexus mapping literacy competencies onto all seventeen [[sustainability|Sustainable Development Goals]]. Framing AI literacy as an "18th SDG" heuristic — a cross-cutting capacity that channels learning into governance and sustainable development — the study's survey of 300 professionals found governance literacy the strongest predictor of AI–SDG nexus awareness (β = 0.64, r = 0.67 with nexus awareness) and identified ethical reasoning and reflective thinking as the strongest predictors of trustworthy AI use. This ties institutional governance to the cultivation of public [[ai-literacy]], echoing the knowledge base's finding that responsible AI alignment requires [[stakeholders|policymakers]] and citizens who can critically interpret algorithmic systems, not merely compliance-oriented technical control. - **Academic integrity and assessment:** Governance is central to how institutions handle AI-related [[academic-integrity]] concerns and redesign [[assessment]] — moving from prohibition/policing toward guidance, AI literacy, and process-oriented designs, as seen in research on [[student-rationalization-ai-writing|student rationalization]] and [[beyond-detection-authentic-assessment-ai-2025|authentic assessment redesign]]. - **The instrument mix of written-down governance:** [[institutional-ai-policy-health-informatics-2026|Eldredge et al. (2026)]] audited AI-related documents from all 48 CAHIIM-accredited health informatics and health information management master's programs in the United States. Forty of the 48 (83%) had at least one publicly available document, and across the 40 documents analyzed governance was realized mostly as guidance rather than binding rule: 21 guidelines (53%), 9 informational documents (23%), and only 7 formal policies (18%). Most documents addressed faculty and students together (20, 51%) rather than students alone (6, 15%), and neither policy type nor audience varied with delivery mode (Fisher's exact P = .85 and P = .71), suggesting the instrument mix tracks institutional habit rather than the demands of online, campus, or hybrid programs. Content clustered on academic integrity, responsible AI use, and student conduct, with little guidance covering AI use in applied learning, research, and simulated environments, precisely where curricular and data governance concerns intersect. The authors also report that regional accreditors supplied most of the corpus (Higher Learning Commission n = 14 and SACSCOC n = 13, together about 60%, with none from the WASC region) and argue accreditation is the external lever best placed to reduce variation across programs. - **Governance as a distributed support ecosystem:** [[qian-governing-genai-higher-ed-policy-2026|Qian (2026)]] coded the AI guidance of the 50 US universities ranked most innovative and found governance realized less as rule than as interpretation and support: only 6 of 50 framed their principal page as a "policy", 33 of 50 positioned instructors as the primary translators of institutional expectations into course rules, and the support for doing so clustered across four interlocking units — teaching and learning centers supplying syllabus language and course-policy menus, libraries setting citation and provenance standards, IT and enterprise-governance offices operating data-classification rules and vetted tool stacks, and academic integrity offices carrying due process and education-first remediation. Qian argues it is the [[scaffolding]], not the stance, that reduces ambiguity, and reports it as unevenly distributed: 33 of 50 institutions published syllabus or course-policy materials while only 4 foregrounded student-facing guidance. This is the instrument-mix finding above in organizational form — governance capacity located in units rather than documents. - **Ethics, privacy, and bias:** Governance mechanisms operationalize the ethical principles ([[ethics]], [[privacy]], [[bias-mitigation]]) that are often recognized but not enforced, connecting to responsible AI and [[regulation|regulatory]] debates in education. - **A value/norm matrix for governance.** [[agarwal-ethical-values-norms-aied-2026|Agarwal et al. (2026)]], a [[meta-analysis-systematic-review|systematic review]] of 25 articles, consolidate AIED ethics into six main ethical values (non-discrimination, data stewardship, [[human-in-the-loop-ai|human oversight]], goodwill, explicability, educational aptness) and map the ethical norms extracted from the literature onto a stakeholder-by-value matrix. The mapping makes norms actionable rules for realizing specific values and offers a foundation for building detailed ethical frameworks and regulation for AIED — giving educational institutions, developers, and regulators concrete norms to implement. Norms for human oversight cluster on educational institutes and end users, educational aptness on educational institutes and regulators, and goodwill norms aimed at regulators are far more numerous (nine) than for any other set, signaling regulators' role in ensuring AIED benefits learners through policy and legislation. ### Governance education AI governance operates at two levels the knowledge base treats together: the *institutional rules* that govern AI use (policies, acceptable-use frameworks, [[assessment]] and declaration requirements) and *education about those rules* (preparing people to navigate them). Governance without education risks being unenforced or opaque; education without governance lacks teeth. Research in this strand includes [[genai-policies-higher-ed-computing|institutional GenAI policy analysis]], [[genai-declaration-frameworks-higher-education|AI declaration frameworks]], [[genai-assessment-governance|assessment governance]], [[ai-uk-higher-education-policy-2026|UK AI higher-education policy]], and [[raza-farooq-aied-review-2020-2025|comprehensive AIED reviews]] that situate governance in the wider policy landscape. ### Assessment governance A central arena of AI governance is **how institutions govern assessment** — the rules that determine what counts as acceptable AI use, how AI-assisted work is declared, and how summative measures are safeguarded. This includes the design and enforcement of [[ai-use-disclosure|AI use and disclosure statements]]: research shows that mandatory declarations fail when they feel punitive or ambiguous ([[gonsalves-student-non-compliance-ai-declarations-2025|Gonsalves 2025]], [[vetter-hidden-cost-disclosure-genai-2026|Vetter et al. 2026]]), and that clear, consistent, trust-based policy is what actually fosters disclosure. The knowledge base's research distinguishes between *detection-based* governance (policing AI use, e.g., via [[ai-detection]]) and *design-based* governance (redesigning [[summative-assessment|summative]] and [[authentic-assessment|authentic]] assessment so AI use is expected, declared, and scrutinized). [[genai-assessment-governance|Evidence-centered governance of generative AI in assessment]] and [[beyond-detection-authentic-assessment-ai-2025|Beyond Detection]] argue that governance must pair any detection with assessment redesign, while the choice of AI-resistant summative formats (oral exams, proctored/closed-book measures, code-review interviews — see [[summative-assessment]]) is itself a governance decision. Large-scale evidence [[stromberg-generative-ai-learning-penalty-secondary-2026|(Strömberg, Lei, & Wu 2026)]] underscores the importance of governing summative measures, since ungoverned homework can be inflated by AI while actual learning declines. The scope of a prohibition is itself a governance decision: [[wright-transcription-not-generation-2026|Wright (2026)]] shows that rules barring "generative AI" without distinguishing generation from format conversion capture assistive transcription tools and turn policy misapplication into misconduct allegations, burdening disabled and [[equity-in-ai-education|equity]]-exposed students — a governance failure that [[legal-issues-and-risks|legal issues and risks]] treats as an equalities question as much as an integrity one. Evidence that [[ai-refusal-higher-education-diagnostic-non-use-2026|Zagami (2026)]] gathers from high-stakes arenas points to uneven governance rather than simple refusal: partial adoption and incomplete policy in assessment, admissions and disciplinary processes, with institutional delay operating as a governance stance rather than a failure. The institutional task is to interpret refusal well enough to improve governance, asking which decisions require human review, which systems require audit, which uses require disclosure, and which procurement choices require public justification. [[weidlich-inference-at-risk-assessment-validity-2026|Weidlich (2026)]] adds the assessment-specific corollary: detector scores are conditional probabilistic signals that cannot by themselves establish misconduct, so detection-centered governance is an insufficient basis for defending assessment claims. ### Governance across levels AI governance operates at multiple levels — from **national/regulatory** (government policy, the OECD framework, state AI guidelines) to **institutional** (university policies, AI steering groups, ethical review boards) to **classroom** (instructor guidelines, syllabus statements, assignment design). Effective governance aligns these levels: national frameworks set expectations, institutions translate them into policies and support structures, and educators implement them in ways that build students' AI literacy and agency. The knowledge base's research emphasizes that governance is not merely about restriction but about creating the conditions for responsible, equitable, and learning-supportive AI integration — including [[educational-development|faculty development]], transparent guidance, and ongoing evaluation. **Faculty governance and the case for deliberate review.** Where the levels above are aligned through policy, some sectors reach them through shared governance instead. [[gutowski-hurley-genai-policy-legal-education-2025|Gutowski and Hurley (2025)]] describe [[legal-education|law schools]] as a hard case for exactly this reason: US law faculties hold unusual individual authority over course standards, so policy has to be built with them rather than announced to them, and most schools writing prohibitive rules also reserved discretion to individual instructors and committed to revision as the technology moved. Their recommendations — flexibility by design, periodic review, proactive training, and self-regulation through information sharing rather than waiting for an accreditor to dictate policy — describe governance as a continuing process rather than a document, the same point the [[crompton-governing-genai-higher-ed-delphi-2026|Delphi consensus]] makes in recommending scheduled review cycles and a standing multidisciplinary committee. Institution-wide, the **AIGEM framework** ([[tan-aigem-ai-educational-management-2026|Tan et al. 2026]]) positions responsible AI as a strategic organizational capability for educational management, integrating AI strategic leadership, responsible governance, decision intelligence, human-AI collaborative intelligence, competency development, and sustainable value creation -- and links responsible implementation to the SDGs. It underscores that governance is not only a compliance layer over teaching but an executive function of educational institutions.### Connections to related concepts AI governance connects to [[ethics]] (the principles it operationalizes), [[higher-ed]] (the institutional context), [[privacy]] and [[bias-mitigation]] (specific governance concerns), and [[academic-integrity]] (a primary governance arena). It is central to [[change-management|institutional change]] and responsible AI, and intersects with [[ai-literacy]] (governance supports the development of critical, informed use). It also connects to [[learning-analytics]] (data governance) and [[student-experience]] (governance shapes how students navigate acceptable use). **Governing the cognitive commons.** The Cognitive Commons framework ([[cognitive-commons-ai-expertise-regeneration|Lovett 2026]]) frames expertise regeneration as a profession-level collective-action problem requiring Ostrom-style governance (boundary definition, monitoring, graduated sanctions, collective choice). AI governance is thus not only about tool regulation but about sustaining the shared expertise pool professions need — at organizational, professional-association, and policy levels. ### Relationship to educational policy Governance is distinct from — but inseparable from — [[educational-policy-ai|educational AI policy]]. **Policy is the content**: the formal rules and statements (what AI use is allowed, what must be disclosed, what assessment is permitted). **Governance is the machinery** that produces, implements, enforces, and revises those rules: who sets them, how they are resourced and communicated, how compliance and appeals are handled, and how they adapt as AI evolves. Where the [[educational-policy-ai|policy]] page catalogs the *rules themselves* and their maturity gaps, this page focuses on the *structures and practices* that make rules real — steering groups, ethical review, assessment governance, and accountability across levels. A rule on paper is policy; a rule that is owned, monitored, and enforced is governance. The two are mutually dependent: policy without governance is unenforced, and governance without policy lacks direction. [[learning-analytics-to-educational-interventions-2026|Svetec, Divjak & Kadoić (2026)]] identify **ethics & data governance** as one of seven enablers of trustworthy LA-based educational interventions — policies for the ethical use of LA and AI, data privacy, security, and accountability — and position [[trust|trustworthiness]] (including leadership and governance that support implementation) as the prerequisite for meaningful data-informed interventions. ## Connected Concepts - [[pedagogical-partnerships]] — Pedagogical Partnerships - [[ai-use-disclosure]] — AI use and disclosure statements - [[remote-proctoring]] - [[ethics]] - [[higher-ed]] - [[privacy]] - [[bias-mitigation]] - [[academic-integrity]] - [[ai-literacy]] - [[learning-analytics]] - [[student-experience]] - [[educational-policy-ai]] - [[regulation]] - [[summative-assessment]] - [[ai-misuse-learning-harm]] - [[trust-calibration]] - [[generative-ai]] - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) ## Connected Articles - [[ai-literacy-sdg-governance-framework-2026]] — AI literacy as a governance capacity for sustainable development: the AIRE Taxonomy and AI–SDG Nexus (Islam, Morshed & Islam 2026) - [[tan-aigem-ai-educational-management-2026]] — AIGEM framework for AI governance in educational management - [[institutional-ai-policy-health-informatics-2026]] — AI policy and guidance documents across 48 CAHIIM-accredited health informatics programs: governance as non-binding, integrity-centric guidance (Eldredge et al. 2026) - [[guided-inquiry-genai-course-policy-2026]] — Students co-designing GenAI course policies via guided inquiry (Hingle & Johri 2026) - [[learning-analytics-to-educational-interventions-2026]] — From learning analytics to educational interventions: enablers of trustworthy LA-based interventions (Svetec, Divjak & Kadoić 2026) - [[gonsalves-student-non-compliance-ai-declarations-2025]] — Student non-compliance with AI use declarations - [[vetter-hidden-cost-disclosure-genai-2026]] — The hidden cost of disclosure - [[chang-should-i-tell-my-teacher-ai-disclosure-2026]] — Student AI disclosure, stigma, and self-regulated learning - [[weidlich-inference-at-risk-assessment-validity-2026]] — Which inference is at risk: assessment validity reasoning and generative AI (Weidlich 2026) - [[ai-adaptation-gap-higher-education-2026]] — The AI Adaptation Gap in Higher Education - [[crompton-governing-genai-higher-ed-delphi-2026]] — Global Delphi on GenAI governance and policy - [[qian-governing-genai-higher-ed-policy-2026]] — Governance by guidance: instructor-set syllabus rules and a four-unit support ecosystem across 50 innovative US universities (Qian 2026) - [[gutowski-hurley-genai-policy-legal-education-2025]] — Faculty governance, instructor discretion and periodic review in law school GenAI policy (Gutowski & Hurley 2025) - [[wright-transcription-not-generation-2026]] — Over-inclusive "AI" prohibitions, format conversion and the reasonable-adjustment problem (Wright 2026) - [[ai-refusal-higher-education-diagnostic-non-use-2026]] — Refusal as evidence: uneven governance, the duty to understand and the diagnostic value of non-use (Zagami 2026) - [[baroudi-anticipatory-governance-ai-higher-ed-2026]] — Anticipatory governance and leadership for AI - [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of]] — Implementing AIDA at the Open University - [[leveraging-complex-systems-leading-for-transformative-change]] — SPARK: Leading for Transformative Change - [[students-engagement-with-generative-ai-in-academic-learning-a-self-determination]] — Students' Engagement With GenAI (SDT) - [[oecd-digital-education-outlook-2026]] — OECD Digital Education Outlook 2026 - [[beyond-detection-authentic-assessment-ai-2025]] — Beyond Detection: Authentic Assessment Redesign - [[genai-policies-higher-ed-computing]] — Institutional GenAI policy in computing - [[genai-declaration-frameworks-higher-education]] — AI declaration frameworks - [[genai-assessment-governance]] — Assessment governance under GenAI - [[ai-uk-higher-education-policy-2026]] — AI in UK higher-education policy - [[raza-farooq-aied-review-2020-2025]] — Comprehensive review of AIED research - [[cognitive-commons-ai-expertise-regeneration]] — The tragedy of the cognitive commons: AI and expertise regeneration - [[policy-deficit-ai-sel-2026]] — The Policy Deficit in AI × SEL Research - [[agarwal-ethical-values-norms-aied-2026]] — Ethical values and norms for AI in education - [[perrotta-zero-shot-governance-2026]] — Zero-shot governance: general-purpose AI in policy (Perrotta 2026) --- ## [Change Management](https://edtechdev.github.io/aied/concepts/change-management/) > **Change management** — the deliberate set of processes, governance structures, and leadership practices through which [[higher-ed]] institutions plan, implement, and sustain the integration of AI into teaching, learning, and [[assessment]]. Because [[generative-ai|generative AI]] arrived as an "arrival technology" that entered classrooms before [[pedagogy|pedagogical]] evidence accumulated, change management in [[ai-education|AI education]] must support continuous adaptation under uncertainty rather than one-time adoption, balancing institutional stability with emergent innovation and shared governance across [[administrator|administrators]], [[teacher-role|faculty]], students, and [[educational-policy-ai|policymakers]]. ## Questions to Consider - Change management here is described as governing continuous adaptation under uncertainty, because generative AI arrived in classrooms before pedagogical evidence accumulated. What does it mean to manage change when you can't wait for best practices but also can't responsibly scale untested innovations? - A recurring obstacle is the gap between top-down policy intent and classroom reality — for instance, policies that encourage GenAI while course syllabi prohibit it, leaving instructors to improvise. Why do you think policy and practice drift apart so easily around AI? - [[research-methods-aied|Research]] finds that a faculty member's pedagogical orientation — what they believe AI means for disciplinary knowledge — is the strongest predictor of adoption, while institutional initiatives and demographics are surprisingly weak predictors. What does that suggest about where change efforts should focus their energy? - The page notes that only 7% of institutions have created senior AI leadership roles despite 49% viewing AI as a strategic priority. What do you think it signals when an institution calls something strategic but doesn't resource it with leadership? - A coordination-game model explains why policy statements alone fail: [[student-ai-interaction|student AI use]] is a collective norm-formation process, and small, well-calibrated changes to assessment incentives can trigger rapid cohort-wide shifts toward responsible use. How might changing assessment incentives do more than issuing a policy ever could? - The page warns that fragmented adoption widens existing gaps — inclusion, equity, and sustainability are often overlooked even where core ethical principles are embraced. Which [[learners]] or institutions do you suspect lose out when change is managed unevenly, and how would you keep them central? ## Introduction The empirical literature consistently shows that effective AI change management is a socio-technical, not merely technical, endeavor. The [[institutional-change-framework-ai|institutional change framework]] adapts classic change models to generative AI, arguing that institutions "cannot wait for best practices, but cannot responsibly scale unjustified innovations," and calling for humble local inquiry, reform organized around pedagogical approaches rather than ephemeral tools, and students engaged as partners in reform. This framing resonates with the [[leveraging-complex-systems-leading-for-transformative-change|complex-systems leadership]] argument that education is a complex adaptive system where leaders must judge when to reinforce administrative stability and when to enable "adaptive space" — diffusing high-risk innovations like AI through trusted social networks (complex contagions) rather than linear knowledge flow. A recurring obstacle to change is the gap between top-down policy intent and classroom reality. The [[genai-policies-higher-ed-computing|comparative policy analysis]] found that while 63% of U.S. R1 university policies encourage GenAI use, half of computing-course syllabi outright prohibit it, leaving instructors to improvise inconsistent local rules. Similarly, the [[ai-uk-higher-education-policy-2026|UK higher-education policy review]] documents a divide between research-intensive and teaching-led institutions, with national guidance criticized for lacking specificity and enforcement power, while the [[institutional-governance-ai-universities|institutional governance analysis]] shows university-wide policies emphasize data security and risk mitigation while [[k-12|school]]-level policies (when they exist) focus on pedagogy — a structural misalignment that falls short of accreditation expectations for unified integration of [[curriculum-design|curriculum]], policy, assessment, and infrastructure. ## Change management in AI education research Scholarship converges on several change-management levers. **Governance frameworks** provide consensus-driven templates: the [[crompton-governing-genai-higher-ed-delphi-2026|global Delphi study]] proposes an eight-area GenAI governance framework (academic integrity, ethical use, [[privacy]], equitable access, literacy, integration, [[human-in-the-loop-ai|human oversight]], institutional support) plus a six-part review mechanism to keep policy current, positioning policies as enabling structures within an interconnected institutional ecosystem. **Anticipatory leadership** is the complementary thread: the [[baroudi-anticipatory-governance-ai-higher-ed-2026|scoping review]] finds institutions must shift from reactive to foresight-driven governance, with empowering and distributive leadership increasing adoption — yet only 7% of institutions have created senior AI leadership roles despite 49% viewing AI as a strategic priority. **Adoption drivers** reveal why change stalls or succeeds. The [[alrahmi-org-drivers-ai-adoption-he-2026|organizational drivers study]] shows that internal organizational culture and technological attributes (compatibility, relative advantage, low complexity) catalyze AI adoption, with government [[regulation]] as an external enabler — context-sensitive dynamics that individual-level [[technology-acceptance-model|technology acceptance]] models miss. Conversely, the [[ai-pedagogical-orientation|faculty orientation study]] finds that a faculty member's epistemic interpretation of AI — their pedagogical orientation toward what AI means for disciplinary knowledge — is the strongest predictor of adoption, while institutional initiatives and demographics are surprisingly weak predictors. This cautions that top-down strategic plans have limited impact unless they engage faculty beliefs, and that bottom-up peer networks (department colleagues were the top information source) matter more than central initiatives. **Institutional transformation case studies** show how these levers play out in practice. [[ai-digital-transformation-liberal-arts-lingnan-2026|Qin (2026)]] analyzes Lingnan University's strategic change across four dimensions — instructional upgrade through AI, prioritization of irreplaceable human competencies, curriculum renewal, and retention of ethical/cultural values — illustrating how a long-established institution can manage [[generative-ai|GenAI]]-driven change without relinquishing its identity. It frames the transformation as intellectual rather than technocentric, giving institutional leaders a staged model for mandating AI literacy while preserving humanistic mission. **Assessment reform** is a central change-management battleground. The [[ai-assessment-scale-reform|AI Assessment Scale study]] shows framework implementation hampered by departmental inconsistencies, workload pressures, and uncertainty, with staff describing the process as "a bit of chaos and madness." The [[ethical-ai-higher-ed-game-theory|coordination game model]] offers a formal account of why policy statements alone fail: student AI use is a collective norm-formation process, and small, well-calibrated changes to reflective assessment incentives can trigger rapid cohort-wide shifts toward responsible use, whereas weak or misaligned incentives allow opportunistic practices to persist. This supports pedagogy-led governance over surveillance. The [[ai-adaptation-gap-higher-education-2026|AI adaptation gap survey]] adds a [[stakeholders|stakeholder]] dimension: students report higher AI-use intensity and perceived usefulness than faculty and administrative staff, while the latter report stronger [[academic-integrity]] concerns — and perceived usefulness drives [[trust]] (β = 0.402) more strongly than institutional policy clarity (β = 0.223). ## Implications Change management in AI education carries both positive and negative implications. **Positively**, structured frameworks give institutions a path from reactive crisis management to proactive, participatory governance; [[crompton-governing-genai-higher-ed-delphi-2026|Delphi consensus]] and [[baroudi-anticipatory-governance-ai-higher-ed-2026|anticipatory governance]] both support treating [[ai-literacy]] and [[ethics]] as cross-cutting institutional capabilities rather than isolated rules. **Negatively**, fragmented adoption widens existing gaps: the [[adarkwah-genai-unesco-policy-2026|UNESCO framework analysis]] of 30 leading universities finds core ethical and [[governance]] principles widely embraced but [[inclusive-learning|inclusion]], equity, and [[sustainability]] (internet access, gender parity, environmental impact) often overlooked — and national AI-preparedness rankings do not predict robust institutional policy. The [[league-ethical-governance-student-data-2026|LEAGUE framework]] extends this concern to [[learning-analytics]], arguing that governance must move beyond compliance (FERPA/GDPR) toward lawfulness, equity, [[agency]], utility, and ethics by design, reviewing student-data practices transparently and educationally. **For [[administrator|administrators]] and [[educational-policy-ai|policymakers]]**, the evidence argues for participatory governance, infrastructure investment, and capacity-building over aspirational strategy documents, translating national policy into concrete instructor support rather than issuing top-down mandates. **For instructors**, change management means moving from adopter to inquiry-driven experimenter, engaging [[pedagogical-partnerships|students as partners]], and anchoring reform in durable pedagogical principles rather than tools that may be obsolete within months — with the caveat that professional development must target a full design cycle (needs assessment, feedback) rather than tool use alone, as the [[dot-framework-survey-2026|DOT framework survey]] demonstrates. ## Connections to other concepts Change management is the institutional complement to classroom-level integration. It operationalizes the systemic conditions — governance, faculty development, stakeholder [[student-engagement|engagement]], and equity safeguards — that allow pedagogical innovation to take hold, connecting [[educational-policy-ai]] policy design to [[governance]], [[educational-development]], and [[equity-in-ai-education]] outcomes. ## Connected Concepts - [[governance]] — the policy and oversight structures change management operationalizes - [[educational-policy-ai]] — national and institutional AI policy intent that change management must translate into practice - [[educational-development]] — building faculty capacity and pedagogical orientation for AI integration - [[administrator]] — leadership and anticipatory governance as change agents - [[ai-literacy]] — cross-cutting capability underpinning responsible institutional adoption - [[equity-in-ai-education]] — ensuring change benefits accrue evenly across institutions and learners - [[technology-acceptance-model]] — adoption drivers that explain how change spreads - [[academic-integrity]] — central regulatory anchor around which assessment reform is organized ## Connected Articles - [[institutional-change-framework-ai]] — six-dimension framework for adapting institutional change models to AI as an arrival technology - [[leveraging-complex-systems-leading-for-transformative-change]] — SPARK framework and complexity leadership for transformative change - [[crompton-governing-genai-higher-ed-delphi-2026]] — global Delphi consensus on eight-area GenAI governance - [[baroudi-anticipatory-governance-ai-higher-ed-2026]] — scoping review of anticipatory governance and leadership - [[ai-adaptation-gap-higher-education-2026]] — stakeholder gaps in AI use, attitudes, and trust across students, faculty, and staff - [[ai-uk-higher-education-policy-2026]] — national policy intent vs. institutional capacity in UK higher education - [[alrahmi-org-drivers-ai-adoption-he-2026]] — organizational and technological drivers of AI adoption - [[adarkwah-genai-unesco-policy-2026]] — UNESCO framework analysis of institutional GenAI policies - [[ai-digital-transformation-liberal-arts-lingnan-2026]] — Digital transformation of a liberal arts university toward a research-intensive model in the GenAI era (Qin 2026) - [[beyond-the-algorithm-academic-developers-digital-mediators-2026]] — Digital mediators translating institutional AI strategy into grounded practice under infrastructural scarcity --- ## [AI Regulation in Education](https://edtechdev.github.io/aied/concepts/regulation/) > **AI regulation** — the laws, policies, and governance frameworks that control how AI is developed and deployed in educational settings. Regulation in the knowledge base spans government policy, institutional governance, and industry [[self-regulated-learning|self-regulation]]. ## Questions to Consider - AI tools are deployed in classrooms far faster than rules can be written. Before you read, who do you think is actually setting the effective rules right now — lawmakers, institutions, developers, or teachers improvising on the spot? - The page distinguishes regulation (binding laws and rules) from governance (the broader norms and structures). Why does this distinction matter? What does an institution with strong governance but weak regulation look like, and is that a stable situation? - Regulation 'both constrains and enables' — it sets boundaries while creating conditions for equitable, safe integration. Can you think of a rule that would simultaneously limit misuse and expand responsible use, or is the tension unavoidable? - Ethics frameworks increasingly harden into binding rules, and safety requirements act as de facto regulation. Do you see ethical principles becoming enforceable rules as progress, or as a way to appear accountable without real teeth — and how would you tell the difference? - The knowledge base documents a persistent 'governance gap' between deployment speed and regulatory maturity, uneven across jurisdictions and educational levels. As a [[teacher-role|teacher]] or developer, how does an inconsistent regulatory environment affect your day-to-day decisions about what AI to use? - Student regulatory awareness [[research-methods-aied|research]] asks whether [[learners]] actually know and follow the rules. Before you read, how well do you think most students understand the AI rules they're bound by — and whose responsibility is it when they don't? ## Introduction Regulation is the legal and policy layer of AI [[governance]]: it sets the binding rules, standards, and enforcement mechanisms that institutional governance translates into practice. Where governance is the broad framework of norms and structures, regulation provides the authoritative rules — from national AI laws and [[privacy|data-protection]] statutes to institutional acceptable-use policies and professional guidelines. A recurring theme in the knowledge base's research is that regulation lags behind AI deployment, leaving institutions to improvise governance in the gap. ### Regulatory landscape - **Government policy:** [[educational-policy-ai]] research examines national and regional [[ai-education|AI education]] policies. and [[ai-lifelong-learning-policy|lifelong learning policy]] address regulatory gaps, while [[ai-uk-higher-education-policy-2026|UK AI higher-education policy]] and [[oecd-digital-education-outlook-2026|the OECD Digital Education Outlook]] situate national approaches in comparative and international perspective. - **Institutional governance:** [[governance|AI governance frameworks]] and [[genai-policies-higher-ed-computing|institutional policy analysis]] document how universities develop internal AI rules, while [[genai-declaration-frameworks-higher-education|AI declaration frameworks]] and [[genai-assessment-governance|assessment governance]] regulate AI use in assessed work. [[qian-governing-genai-higher-ed-policy-2026|Qian (2026)]] finds the sector's internal rules are mostly guidance rather than binding policy — 44 of 50 innovative US universities published guidelines, principles or resource hubs while only 6 framed their principal page as a "policy" — with instructor-set syllabus rules doing the operative regulatory work in individual courses. [[watson-rainie-ai-challenge-faculty-survey-2026|Watson and Rainie's (2026)]] survey of 1,057 US faculty measures how far below the institutional layer the effective rules sit: 87% of respondents wrote their own assignment-level policies while only 48% reported written institutional guidelines and 35% departmental ones, against a structural response that is thin at the top — a task force or oversight group in 55% of cases but AI literacy adopted as a general education outcome in only 13%. - **Safety regulation:** [[pedagogical-safety]], [[child-safety-genai|child safety]], and [[eduzone-llm-safety-k12|K-12 safety frameworks]] represent de facto regulation through safety requirements. [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] shows why assessment tooling belongs in that category too: in a red-team test, two of five indirect prompt injections hidden in a submission file raised a failing grade without any warning to the marker — reported success rates of 100% and 94% — and the paper's sector-level asks are clear AI policy, [[educational-development|professional development]], and standardized, domain-agnostic assessments of prompt injection resilience so that the attack surface is measured rather than assumed. - **Ethics as regulation:** [[ethics]] frameworks increasingly serve regulatory functions — [[ai-ethics-education-public-discourse|public discourse on AI ethics]] shapes policy expectations, and [[league-ethical-governance-student-data-2026|ethical governance of student data]] shows how ethics principles harden into binding rules. - **Equalities and reasonable-adjustment law:** Over-inclusive AI rules can collide with statutory duties. [[wright-transcription-not-generation-2026|Wright (2026)]] argues that prohibitions barring "[[generative-ai|generative AI]]" without distinguishing content generation from format conversion capture AI transcription tools and may engage the reasonable-adjustment duty under the UK Equality Act 2010, the US Americans with Disabilities Act and the Australian Disability Discrimination Act 1992 — making the drafting of a prohibition a regulatory question, not only an academic-integrity one (see [[legal-issues-and-risks|legal issues and risks]]). [[li-genai-assessment-language-equity-2026|Li (2026)]] extends the same collision to language background: assessment rules that do not separate legitimate language support from substantive substitution impose cohort-skewed compliance burdens on students using English as an additional language, and the framework designed to correct this is defended through indirect-discrimination reasoning plus the administrative-law expectation that a decision-maker can state the rule applied, the evidence relied on and why the outcome was proportionate — the standard a challenge on review applies to an administrative decision. - **Compliance and accountability:** [[student-regulatory-awareness-genai|student regulatory awareness]] examines whether learners actually know and follow AI rules, and [[dot-framework-survey-2026|technology-adoption frameworks]] explore how regulatory and ethical concerns influence adoption decisions. ### The governance gap The knowledge base documents a persistent gap between AI deployment speed and regulatory maturity. [[institutional-change-framework-ai|Institutional change frameworks]] and regulation research argue for proactive [[governance]] rather than reactive policy. Studies of and [[raza-farooq-aied-review-2020-2025|comprehensive AIED reviews]] highlight that regulation is uneven across jurisdictions and educational levels, creating an inconsistent operating environment for teachers, students, and developers. **The gap is one of evidence as well as timing.** [[gutowski-hurley-genai-policy-legal-education-2025|Gutowski and Hurley (2025)]] characterize one professional sector as making policy under time pressure without an evidence base to make it with: most ABA-approved US [[legal-education|law schools]] took generally prohibitive positions while reserving discretion to individual instructors, and the authors report no consensus on disclosure or citation practice and only a ~15% response rate to the ABA's 2024 policy survey. Their normative response — clear guidelines whatever the stance, [[stakeholders|stakeholder]] involvement in drafting, and governance designed to be flexible and reviewed periodically — matches the [[crompton-governing-genai-higher-ed-delphi-2026|global Delphi consensus]], which likewise treats policy maintenance as a recurring institutional mechanism rather than a one-time task. [[coates-governing-academic-integrity-indicators-2025|Coates, Croucher and Calderon (2025)]] add the reverse dependency: their governance reform program concludes that institutional development is unlikely to pay out without external affordance from regulation, benchmarking and cross-institutional competition, making the quality and regulatory agencies — and the comparison they force between institutions — the condition under which internal governance reform takes hold. ### Connections Regulation connects to [[educational-policy-ai]], [[governance]], [[ethics]], [[privacy]], [[pedagogical-safety]], and [[academic-integrity]]. It is the institutional layer that shapes how all other AI education practices operate. Regulation both constrains and enables: it sets the boundaries of acceptable AI use while creating the conditions — through [[ai-literacy]] and responsible-use expectations — for [[equity-in-ai-education|equitable]], safe integration. ## Connected Concepts - [[educational-policy-ai]] - [[governance]] - [[ethics]] - [[privacy]] - [[pedagogical-safety]] - [[academic-integrity]] - [[equity-in-ai-education]] - [[higher-ed]] - [[k-12]] - [[ai-literacy]] - [[trust-calibration]] - [[legal-issues-and-risks]] — legal exposure when regulation and governance are unclear or over-broad ## Connected Articles - [[ethical-ai-higher-ed-game-theory]] — Coordination game framework for ethical AI use in higher education (Ogbo et al. 2026) - [[institutional-change-framework-ai]] - [[genai-policies-higher-ed-computing]] - [[ai-lifelong-learning-policy]] - [[ai-ethics-education-public-discourse]] - [[ai-uk-higher-education-policy-2026]] — AI in UK higher-education policy - [[oecd-digital-education-outlook-2026]] — OECD Digital Education Outlook 2026 - [[genai-declaration-frameworks-higher-education]] — AI declaration frameworks - [[genai-assessment-governance]] — Assessment governance under GenAI - [[league-ethical-governance-student-data-2026]] — Ethical governance of student data - [[student-regulatory-awareness-genai]] — Student regulatory awareness of GenAI - [[dot-framework-survey-2026]] — Technology-adoption frameworks - [[raza-farooq-aied-review-2020-2025]] — Comprehensive review of AIED research - [[generative-ai-reduced-study-time-math]] — Age gradient and proctoring findings inform AI policy - [[policy-deficit-ai-sel-2026]] — The Policy Deficit in AI × SEL Research - [[qian-governing-genai-higher-ed-policy-2026]] — Guidance over binding policy: internal AI rules across 50 innovative US universities (Qian 2026) - [[gutowski-hurley-genai-policy-legal-education-2025]] — Law school GenAI policy scored on five dimensions: prohibitive by default, instructor discretion, periodic review (Gutowski & Hurley 2025) - [[wright-transcription-not-generation-2026]] — Over-inclusive AI prohibitions and the reasonable-adjustment duties they may engage (Wright 2026) - [[li-genai-assessment-language-equity-2026]] — Language background in assessment rules: the support–substitution boundary as an equalities question (Li 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Prompt-injection attacks on AI graders and the case for standardized resilience testing (Humble 2026) - [[coates-governing-academic-integrity-indicators-2025]] — External regulatory pressure as the condition for governance reform (Coates, Croucher & Calderon 2025) - [[watson-rainie-ai-challenge-faculty-survey-2026]] — 87% of faculty write their own rules against a thin institutional policy layer (Watson & Rainie 2026) --- ## [Equity](https://edtechdev.github.io/aied/concepts/equity-in-ai-education/) > **Equity** — the principle that AI should serve all learners fairly, and the study of systemic disparities in access to, representation within, and benefits from AI educational tools. Equity [[research-methods-aied|research]] in the knowledge base examines access gaps and the digital divide, bias and fairness in [[ai-technologies|AI systems]], culturally responsive and linguistically inclusive design, accessibility for learners with disabilities, and the distribution of AI's benefits and harms across groups. It connects the technical (bias mitigation, fair algorithms) with the structural (infrastructure, policy) and the [[pedagogy|pedagogical]] (culturally relevant [[teacher-role|teaching]]). ## Questions to Consider - Providing AI tools to a school or classroom does not automatically close [[learning-gains|achievement gaps]] — in fact, access alone can widen them. If 'access is not enough,' what else has to be in place for AI to actually serve all learners fairly? - Even the data used to *simulate* learners carries bias: when LLMs generated student vignettes, different models produced more Global North or Global South profiles and gendered pronouns. How much should we trust AI-generated representations of learners when the models themselves encode uneven priors? - Equity in AI education is often framed around three concerns: who gets the tools (access), who is represented in them (representation), and who benefits (outcomes). Can you think of a situation where a group gets access but still doesn't benefit? What explains the gap? - Research on 'structural silence' argues that speakers of underrepresented languages are disadvantaged by AI infrastructure — training corpora, tokenization, benchmarks — *before any model is even trained*. If the disadvantage is baked into the infrastructure, where does fixing it start? ## Introduction Equity in [[ai-education|AI education]] addresses three overlapping concerns: who *gets* AI tools (access), who and what is *represented* in AI systems (representation), and who *benefits* (outcomes). AI can both widen and narrow existing disparities depending on design, infrastructure, and policy. Equity is therefore a cross-cutting lens applied to [[bias-mitigation|algorithmic fairness]], [[digital-divide|digital access]], [[language-learning|linguistic inclusion]], [[accessibility]], and [[culturally-relevant-pedagogy|culturally relevant teaching]]. ## Access and infrastructure equity - **The digital divide:** [[digital-divide|Unequal access]] to AI-powered learning tools across socioeconomic lines, regions, and nations is a foundational barrier. documents how [[generative-ai|generative AI]] benefits are distributed unevenly across countries and institutions. - **Access is not enough:** [[access-not-enough-ai-tutoring-2026|Providing AI tools without addressing structural barriers]] does not close gaps — access must be paired with skills, support, and conditions that enable genuine use. - **Faculty expect the divide to widen.** [[watson-rainie-ai-challenge-faculty-survey-2026|Watson and Rainie (2026)]] surveyed **1,057 US college and university faculty** and found **81%** expected generative AI to widen digital inequities (**58%** saying a lot) — the strongest equity expectation in the report and one that sits beside **68%** who said their institutions had not prepared faculty to use the tools for teaching and mentoring. The same survey shows how unevenly the tools are taken up: **26%** of respondents do not use generative AI at all, rising to **40%** of arts and [[humanities-education|humanities]] faculty and **28%** of social scientists, so non-use and unpreparedness concentrate in particular [[discipline-specific-aied|disciplines]]. The report is a non-probability sample its authors state is not generalizable, so these are the sector's expressed concerns rather than measured effects. - **Infrastructure disadvantage:** [[structural-silence-underrepresented-language-ai-2026|Structural Silence]] shows that AI *infrastructure* — training corpora, tokenization, [[benchmark|benchmarks]], deployment architectures — systematically disadvantages speakers of underrepresented languages *before a model is trained*, reframing dataset scarcity as a structural rather than incidental problem. - **Model-specific demographic priors in synthetic data:** [[lopez-pernas-llm-appropriate-student-support-2026|López-Pernas et al. (2026)]] found that when LLMs generated student vignettes, each model imposed distinct demographic tendencies — GPT produced more Global North profiles and used they/them pronouns, Qwen produced more [[global-south|Global South]] profiles, and Mistral skewed toward she/her. Even the *construction* of learner data by an [[llm]] thus carries regional and gendered priors that can propagate into downstream recommendations, an under-examined equity risk. - **Socioeconomic gradients:** [[ai-lifelong-learning-policy|AI and lifelong-learning policy]] and [[generative-ai-education-productivity-gaps|productivity-gap experiments]] examine how AI can either narrow or widen gaps among different learner groups. - **Bridging divides for disabled learners:** [[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]] found that GenAI levels the playing field for visually impaired undergraduates across digital, geographic, and socioeconomic divides, framing inclusion as both an infrastructural and a cultural matter — extending digital equity discourse beyond access to belonging, voice, and representation. Preparation matters more than preference, and the right to refuse is unevenly distributed. Students with strong academic confidence can refuse AI without penalty, while students who need language support, accessibility support or rapid feedback experience refusal as a loss of opportunity, and casual staff may feel pressure to adopt tools that cut preparation time without cutting responsibility. Where a duty to understand is imposed without training, secure infrastructure and clear policy, it becomes hidden workload, and refusal turns into a predictable response to institutional under-preparation rather than resistance to the technology ([[ai-refusal-higher-education-diagnostic-non-use-2026|Zagami 2026]]). ## Representational equity - **Bias in training data and outputs:** AI training data largely reflects dominant cultural perspectives. [[gender-bias-transfer-llm-writing|Gender bias transfer research]] shows LLM-assisted writing can contaminate student work with gender bias; [[paternalistic-filter-llm-history-education|history-education filters]] and [[ai-scoring-language-bias-physics|AI scoring]] can encode Western-centric and linguistically biased assumptions. - **Marginalized knowledges:** [[genai-minoritized-knowledges-disability|Research on minoritized knowledges]] examines how generative AI marginalizes non-dominant knowledge systems and disability perspectives in [[higher-ed|higher education]]. - **[[curriculum-design|Curriculum]] diversification:** Teachers increasingly use LLMs to diversify curriculum materials (Wang et al., 2025, found 78% did so), yet AI-curated reading lists still underrepresent BIPOC authors, and [[stem-education|STEM]] [[intelligent-tutoring|AI tutors]] default to Western-centric problem contexts. ## Outcome equity - **Differentiated impact:** AI tools may widen gaps if designed without an equity lens — [[genai-higher-education-systematic-review-2026|systematic reviews]] and [[ai-scoring-language-bias-physics|scoring-bias studies]] show uneven benefits and harms across learner groups. - **Bias amplification:** AI suggestions and [[ai-feedback-quality|automated feedback]] can reinforce (not challenge) existing teacher and systemic biases. [[marked-pedagogies-linguistic-bias-writing-feedback|Marked Pedagogies]] shows LLM writing-feedback tools systematically shift toward stereotype-aligned praise and withheld critique when feedback is personalized with a student's race, language, disability, achievement, or motivation — even on identical essays — making "[[personalized-learning|personalization]]" a concrete bias vector in automated feedback. - **Fairness-aware systems:** [[bias-mitigation]] and [[ground-truth-reliability-aied|ground-truth reliability]] research develop methods for detecting and correcting bias in AI tutors, scorers, and recommenders. - **Fairness regularizers may not generalize to new learners:** [[student-attention-estimation-fairness-2026|Fragkiadakis et al. (2026)]] added gender- and age-targeted error-gap regularization to a [[multimodal]] transformer predicting real-time student attention, and found it narrowed demographic disparities on *validation* data but these gains did not consistently transfer to held-out subjects or repeated subject-level splits (the regularized model reduced the gap in only 4 of 10 training runs). Certified fairness on a single split can therefore evaporate on genuinely new learners — educational AI needs leave-subjects-out, repeated-seed evaluation rather than aggregate metrics alone. - **Student agency:** ensuring AI empowers rather than replaces [[student-experience|student voice]] and [[agency]], especially for historically marginalized learners. - **Psychological vs. cognitive equity:** [[school-ai-education-readiness-gaps-agency-2026|Liang et al. (2026)]] found a year of school AI instruction in Hong Kong secondary schools **narrowed psychological AI-readiness gaps (confidence, [[motivation]], [[ethics|ethical]] awareness) but not cognitive ones** — objective [[ai-literacy]] gaps between self-initiated ("high-agency") learners and their peers persisted, a Matthew-effect pattern where curricula "raised the floor but did not level the playing field." Access to a curriculum alone, without sustained self-initiated [[student-engagement|engagement]], may foster psychological but not full cognitive parity. - **Automated marking, attainment and language.** [[opraise-automated-marking-ai-assessment-2026|The OpRaise comparison of three frontier models on 761 authentic essays]] found that AI–human disagreement varied with students' attainment level and with surface language features (vocabulary range, connectives, sentence complexity) in ways human marking did not, and that accuracy differed across three UK institutions whose cohorts differ — the authors connect this directly to institutions' duties under the UK Equality Act, on the reasoning that some student groups may be affected more than others, and note a right to explanation under GDPR Article 22 where automated marking decisions affect students. Because AI marks compressed toward the middle of the distribution, the students most exposed are those at the top and bottom of the attainment range. - **Assessment rules can convert linguistic disadvantage into integrity risk.** [[li-genai-assessment-language-equity-2026|Li (2026)]] argues that integrity rules treating GenAI as a single category of unauthorized assistance impose higher compliance burdens on students who use English as an additional language, whose legitimate use of language support is more frequent and iterative, and concentrate suspicion on writers whose surface fluency has shifted most — a rule-design rather than a behavior problem, and one that raises the risk of selective enforcement on weak evidence. The proposed boundary is purpose- and construct-based instead of tool-based: grammar, punctuation and sentence-level clarity edits, translation for comprehension, and first-pass drafting that the student substantively rewrites count as permitted support because they add no new ideas, no new sources, and no material re-ordering of the analysis, whereas generating arguments or counter-arguments, applying disciplinary rules to facts, restructuring the analytical sequence, or generating citations count as substitution. Because construct statements often reward fluency and idiom as proxies for reasoning, EAL students meet construct-irrelevant variance in their scores; the remedies are ex ante specificity about permitted and prohibited functions, [[ai-use-disclosure|disclosure]] calibrated so that routine support costs less to declare than a bibliographic entry, staged submissions and source trails rather than [[ai-detection|detection-led]] inference, and cohort-level monitoring of referrals and sanctions. The framework is normative and untested — Li makes no claim about rates of GenAI use by EAL students. - **Students themselves judge the support-substitution line accurately, but hesitantly.** [[reed-ai-literacy-ethical-judgment-scenarios-2026|Reed et al. (2026)]] put six ethical vignettes to 531 undergraduates and found 57.6% classified all six correctly (mean 88.54%): verification-supported uses were widely accepted — summarizing notes with professor verification (91.7%), grammar improvement on a self-written paper (84.9%), AI-generated search terms followed by an independent literature review (91.1%) — while substitutionary uses were overwhelmingly rejected. Uncertainty clustered on the cases a rule must actually resolve, the unverifiable AI-generated references (9.9% "unsure") and minimally edited AI-generated text (9.6%), which the authors read as evidence that the boundary between assistance and meaningful authorship is not clear to students. Objective [[ai-literacy]] predicted classification accuracy only weakly (Spearman's ρ = 0.234), the sample was predominantly White, female and first-year at one public Midwestern university, and the authors stress that correct judgment is not ethical conduct. - **Prompt privilege:** [[prompt-privilege-equitable-ai-access-2026|Jin et al.]] document "prompt privilege" — users who phrase requests skillfully systematically obtain better LLM output than users expressing the same intent less adroitly — making [[prompt-engineering|prompting]] skill a silently uneven resource. Their Prompt Equity Transformer shifts prompt optimization into the system, treating equitable output as an accessibility property rather than demanding expert prompting from novices. - **The interaction-management gap.** [[brunnstrom-ai-interaction-literacy-srl-2026|Brunnström and Palmqvist (2026)]] reach an ambivalent conclusion about GenAI as a leveller: because productive use requires recognizing an over-abstract answer, requesting simplification, and structuring a session around small goals, unguided GenAI "may be most beneficial to already advantaged students" — those with strong study habits and confidence in directing an AI — while students with weaker study skills or lower academic [[self-efficacy]] meet added complexity and frustration. Rather than substituting for missing academic conversation partners, the tool introduces a new competence whose acquisition creates its own gap; the authors conclude the responsibility for teaching it cannot rest with the student alone ([[ai-literacy]], [[self-regulated-learning]]). - **Skill-gap and resource-gap mechanisms are not the same problem.** [[kumar-genai-computing-education-systematic-review-2026|Kumar, Wongsirichot and Nanthaamornphong (2026)]] separate two mechanisms their [[meta-analysis-systematic-review|systematic review]] of 72 computing-education studies found the literature tends to conflate. The **skill gap** operates within a single classroom: students with stronger [[prior-knowledge|prior knowledge]] convert AI assistance into durable skill while struggling students use it as a crutch that removes [[desirable-difficulties|productive struggle]], widening the competence distribution by semester's end — addressed by graduated access tied to demonstrated competence. The **resource gap** operates across institutions and national contexts: reliable internet and paid API subscriptions sustain more capable tool use than students without them — addressed by institutional investment in shared tool access and policies that do not assume universal availability. Equity is the thinnest of the review's three framework requirements (six studies), and the authors read that thinness as the finding: the absence of equity-focused intervention research is itself the equity problem ([[assessment-validity]], [[scaffolding]]). - **Tutoring quality shifts with learner demographics.** EduFair-Bench pairs a fixed [[simulating-students|LLM student]] with each tutor across nine demographic levels spanning gender, immigration background, first language and socioeconomic status, and scores five turn-level pedagogical metrics. In [[intelligent-tutoring|LLM tutoring]] the largest deviations appear in explicit demographic conditions — step scaffolding correlating up to |r| = 0.144 in mathematics and corrective tone exceeding 0.10 for every model in [[chemistry-education|chemistry]] (0.102-0.168) and [[physics-education|physics]] (0.129-0.294) — and in 11 of 15 model-by-domain cells the wrong-answer condition scored higher than the correct one, indicating that tone tracked the student's demographic label rather than the quality of the student's reasoning. Pedagogy-specific [[reinforcement-learning|reinforcement learning]] redistributed rather than removed these gaps. ([[edufair-bench-pedagogical-fairness-llm-tutors-2026]]) - **Implicit demographic signals are a less controllable bias channel than stated attributes.** [[demographic-signals-llm-student-assessment-2026|Rooein, Benedetto and Hovy (2026)]] held each task input fixed while varying only the demographic context across six instruction-tuned [[llm|LLMs]] and three educational tasks, producing 192,480 inference calls, and separated *explicit* signals (stated student attributes) from *implicit* ones carried by a ten-prompt [[conversational-ai|conversation]] history. In [[automated-essay-scoring|automated essay scoring]] most models were comparatively stable under explicit conditioning, while implicit conditioning inflated scores — Llama-70B scored 1.57 points above its own default (p < 0.001). In metalinguistic question answering the implicit condition drifted the other way: responses for lower education levels received less positive sentiment, a 0.3 average gap between the lowest and the higher education levels on a 0-4 scale against a within-item standard deviation of 0.07. The equity difficulty is structural — the cue is not a stated attribute that a policy can forbid or a prompt field that an audit can inspect, but a property of the interaction itself. ## Linguistic, cultural, and disability inclusion - **Language:** Most AI tools prioritize English, marginalizing [[multilingual-learning|multilingual]] learners. [[genai-linguistic-diversity-academic-writing|Linguistic diversity in academic writing]], [[structural-silence-underrepresented-language-ai-2026|underrepresented languages]], and [[language-learning]] research address this. - **Culture:** [[culturally-relevant-pedagogy|Culturally relevant pedagogy]] and [[culturally-aware-aied-community-learning|community-centered AIED]] call for AI that reflects learners' cultural contexts rather than imposing dominant norms. - **Disability and neurodiversity:** [[inclusive-learning|Accessible learning]], [[universal-design-for-learning|universal design]], [[neurodiversity]], and [[special-education|special education]] research examines how AI can support or exclude learners with disabilities — [[neurodivergent-computing-students|neurodivergent computing students]], [[dyslexlens-dyslexic-learners-ai|dyslexic learners]], and [[inclusive-learning|accessible educational materials]] are illustrative. - **Language, culture, and cost together:** Bashir and Afzal (2026) frame equitable design in [[well-being]] AI as a single question of language, culture, and cost — English-only tools fail students who express distress in Urdu or Roman Urdu, Western datasets fail to capture locally salient stressors, and expensive commercial systems fail less well-equipped universities ([[culturally-aware-student-stress-chatbot-2026|Sukoon]]). Their response combined an [[open-source]] multilingual LLM accessed through a hosted API for low resource requirements, a bilingual assessment interface, and a classifier trained on a validated 20-feature instrument — while acknowledging that the training data is not representative of the target population, that "some areas... tools are available but often expensive," and that the adaptation is prompt-level rather than validated with students, a candid account of the gap between equitable intent and demonstrated equity. ## Special populations and global equity - **Special populations:** [[special-education]], [[neurodivergent-computing-students|neurodivergent learners]], [[dyslexlens-dyslexic-learners-ai|dyslexic learners]], and [[inclusive-learning|learners with disabilities]] represent groups whose needs are often overlooked in AI system design. - **Global South perspectives:** [[suacode-african-students-motivations|African student motivations]], [[connected-ai-lesson-planning-vietnam|Vietnamese AI lesson planning]], and [[pre-service-science-teachers-ai-perceptions-2026|Ghanaian teacher acceptance]] provide Global South perspectives often absent from Western-centric AIED research. At the institutional level, [[adeniranye-ai-integration-nigerian-higher-education-2026|Adeniranye et al. (2026)]] show AI integration across 45 Nigerian universities is driven by institution age and geography rather than governance type, with reinforcing network ties letting well-connected institutions compound advantage — evidence that equity gaps are reproduced structurally, not just through individual access. - **Global capacity:** documents how generative AI benefits are distributed unevenly across countries and institutions, and [[ai-lifelong-learning-policy|AI and lifelong-learning policy]] addresses structural socioeconomic gradients. - **Disability and Global South intersection:** [[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]] — a [[qualitative-research|qualitative]] case study of 21 visually impaired undergraduates across three Palestinian universities — shows GenAI bridging digital, geographic, and socioeconomic divides while extending [[technology-acceptance-model|technology acceptance models]] to disability contexts, where [[usability-research|usability]], affordability, and accessibility are mutually reinforcing. - **Gender equity in computing:** equity-oriented uses of [[generative-ai|GenAI]] remain underexplored. [[all-girls-genai-makerspace-gender-equity-2026|An all-girls GenAI makerspace initiative in Europe]] combined two GenAI tools with feminist pedagogy to address persistent gender inequities in girls' representation in computing, with practitioners enacting specific steps to support girls' participation and engagement — an example of equity-focused GenAI design. Deliberately gendered AI can also function as the *intervention itself*: [[ada-female-coded-chatbot-gender-stereotypes-2026|Rücker and Becker-Genschow (2026)]] showed that a female-coded, [[discipline-specific-aied|domain-specific]] math [[conversational-ai|chatbot]] modeled on Ada Lovelace's persona (and deployed as both a role model and a [[intelligent-tutoring|learning assistant]]) significantly reduced gender-stereotypical beliefs about mathematical ability and [[math-education|mathematics]] as a male domain among ninth graders — in *both* genders, with high and gender-neutral technological acceptance. This reframes representation as a design lever, not just a bias to audit: systematically designed AI personas can counter, rather than merely avoid reproducing, [[gender-bias-transfer-llm-writing|gender bias]]. ## Implications for AI in education - **Fairness is design, not afterthought:** [[bias-mitigation|bias mitigation]] and fairness-aware algorithms must be built into AI tutors, scorers, and recommenders, and evaluated for equity alongside accuracy. - **Infrastructure is equity:** addressing the [[digital-divide|digital divide]] and underrepresented-language infrastructure is a precondition for equitable AI, not a secondary concern. - **Representation matters in content and assessment:** AI-curated materials and [[automated-assessment|automated assessment]] must reflect and not penalize diverse learners, cultures, languages, and knowledge systems. - **Pair access with support:** providing tools is insufficient; learners need skills, conditions, and culturally relevant [[scaffolding]] to benefit. - **Reach the audiences formal education misses:** the [[ai-literacies-young-adults-2025|AI Literacies framework for public service media]] argues that provision will keep reaching the already-advantaged unless it is designed otherwise, names young people who are digitally or otherwise marginalized as those with the fewest opportunities through formal education, and proposes targeted partnerships plus national-reach provision — not universal publication — as the remedy. - **Policy and governance:** institutional AI policy ([[educational-policy-ai]], [[governance]]) must embed equity as a guiding principle. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[learners]] — Learners: the umbrella for the learner-side concepts - [[digital-divide]] — Unequal access to AI tools and infrastructure across socioeconomic lines, regions, and nations - [[bias-mitigation]] — Methods for detecting and correcting bias in AI tutors, scorers, and recommenders - [[accessibility]] — Design that makes AI learning tools usable by learners with disabilities - [[assistive-technology]] — Tools that support learners with disabilities in AI-mediated settings - [[culturally-relevant-pedagogy]] — Teaching that reflects learners' cultural contexts rather than imposing dominant norms - [[language-learning]] — Linguistic inclusion of multilingual and underrepresented-language learners - [[inclusive-learning]] — Accessible and equitable learning for all learners - [[universal-design-for-learning]] — Designing for learner variability from the outset - [[neurodiversity]] — Supporting neurodivergent learners in AI education - [[special-education]] — Meeting the needs of learners with disabilities in AI system design - [[ai-literacy]] — The skills learners need to benefit from AI equitably - [[educational-policy-ai]] — Institutional policy embedding equity as a guiding principle - [[governance]] — Oversight and accountability for equitable AI - [[agency]] — Ensuring AI empowers rather than replaces student voice - [[stakeholders]] — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers) - [[parents-and-families]] ## Connected Articles - [[ai-literacies-young-adults-2025]] — Equity as a delivery problem: reaching young people formal education misses - [[typology-generative-ai-tools-education-2026]] — Free-tier availability as the gate on which tools reach educators - [[opraise-automated-marking-ai-assessment-2026]] — OpRaise report: AI marking of 761 university essays across three UK universities - [[kumar-genai-computing-education-systematic-review-2026]] — Skill-gap vs resource-gap: two equity mechanisms requiring different remedies - [[brunnstrom-ai-interaction-literacy-srl-2026]] — Unguided GenAI may widen gaps: the interaction-management competence (Brunnström & Palmqvist 2026) - [[adeniranye-ai-integration-nigerian-higher-education-2026]] — Institutional structures, digital inequality, and AI integration in Nigerian higher education - [[student-attention-estimation-fairness-2026]] — Fairness-Aware Multimodal Transformer Modeling for Real-Time Student Attention Estimation - [[school-ai-education-readiness-gaps-agency-2026]] — School AI education narrows psychological but not cognitive readiness gaps - [[prompt-privilege-equitable-ai-access-2026]] — Prompt Privilege: measuring & mitigating accessibility disparities in LLM access - [[ai-scoring-language-bias-physics]] — Language bias in AI-based scoring - [[gender-bias-transfer-llm-writing]] — Gender bias transfer in LLM-assisted writing - [[genai-minoritized-knowledges-disability]] — Generative AI and the marginalization of minoritized knowledges - [[structural-silence-underrepresented-language-ai-2026]] — Structural silence: underrepresented languages in AI infrastructure - [[genai-higher-education-systematic-review-2026]] — GenAI in higher education: systematic review - [[ai-lifelong-learning-policy]] — AI and lifelong-learning policy - [[generative-ai-education-productivity-gaps]] — Does generative AI narrow education-based productivity gaps? - [[genai-linguistic-diversity-academic-writing]] — Linguistic diversity in AI-mediated academic writing - [[access-not-enough-ai-tutoring-2026]] — Access is not enough - [[ada-female-coded-chatbot-gender-stereotypes-2026]] — Female-coded chatbot as role model reducing math gender stereotypes - [[paternalistic-filter-llm-history-education]] — Paternalistic filtering in LLM-based history education - [[dyslexlens-dyslexic-learners-ai]] — DyslexLens: AI support for dyslexic learners - [[ground-truth-reliability-aied]] — Ground-truth reliability in AIED - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Marked Pedagogies: stereotype-aligned feedback bias across student attributes - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[all-girls-genai-makerspace-gender-equity-2026]] — All-girls GenAI makerspace workshops and gender equity in computing - [[khlaif-assistive-genai-visually-impaired-2026]] — Assistive GenAI for visually impaired learners - [[culturally-aware-student-stress-chatbot-2026]] — An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning - [[demographic-signals-llm-student-assessment-2026]] — Implicit (conversation-history) demographic signals shift LLM scoring, feedback and answering (Rooein, Benedetto & Hovy 2026) - [[watson-rainie-ai-challenge-faculty-survey-2026]] — 1,057 US faculty on AI's present and future: 81% expect wider digital inequities, 26% do not use the tools - [[li-genai-assessment-language-equity-2026]] — Drawing the support-substitution line: GenAI assessment rules and EAL students' compliance burden - [[ai-refusal-higher-education-diagnostic-non-use-2026]] — The right to refuse is unevenly distributed: refusal as evidence of institutional under-preparation (Zagami 2026) - [[reed-ai-literacy-ethical-judgment-scenarios-2026]] — Scenario-based ethical judgment and AI literacy among 531 undergraduates --- ## [Differential Effects Across Learner Groups](https://edtechdev.github.io/aied/concepts/differential-effects-across-learner-groups/) > **Differential effects across learner groups** — what [[ai-education|AI in education]] research finds about how the *use* of AI and its *effects* differ across kinds of learners: [[special-education|students with disabilities]] and [[neurodiversity|neurodivergent students]], second-language and [[multilingual-learning|multilingual learners]], girls and boys, minoritized students, students from lower-income backgrounds, rural students, first-generation students, international students, and [[adult-learning|adult learners]]. The pattern worth carrying away is uneven in two directions at once: some strands have real evidence (disability, language) while others are close to empty (first-generation, international, refugee), and even the strong strands rarely establish that a *group* differs — they establish that a tool helped or harmed a sample of that group, which is a different claim. This page maps what exists, what it shows, and the methodological reasons a group average is not a prediction about a learner. ## Questions to Consider - Your tool mostly works for the students in your class, and a subgroup of five struggled. Would a subgroup difference of that size be detectable at all in a study of this kind — and would you want to change the tool for the whole class on that evidence? - Research on students with disabilities reports effects that vary by disability category. If two categories sit at very different effect sizes in the same [[meta-analysis-systematic-review|meta-analysis]], what does that say about treating "disability" as one group in a design decision? - Several studies find AI tools that do *not* differ by gender in their effects, and other work shows that [[ai-feedback-quality|AI feedback]] and AI-assisted writing *do* reproduce gender stereotypes when personas or prompts carry them. How can both be true at once? - Most fairness work tests simulated student personas rather than real learners, because real demographic attributes are not available to the model at inference in most deployments. What can a counterfactual persona audit establish, and what can it not? - Reviewing the strands on this page, which learner groups are actually represented in the studies behind your tool's validation — and what would you do about a group that is absent? - An intervention that raises outcomes for everyone but closes a gap (helping the students who were behind the most) has a different equity story from one that raises the average. When you evaluate a pilot, are you measuring the gap or the mean? ## Introduction Two questions hide inside "does AI work for *this* kind of student?". The first is about effects: whether a tool produces different outcomes for different groups. The second is about use: whether different groups adopt, access, or interact with the same tool differently, which can shape outcomes without any differential effect at all. This page covers both, group by group, and states plainly where the literature stops. It is deliberately not a duplicate of its neighbours. [[equity-in-ai-education|Equity in AI Education]] carries the normative and structural frame — access, representation, and outcome equity, and the argument for what AI in education *should* do. [[inclusive-learning|Inclusive Learning]] carries the design frame, including [[universal-design-for-learning|Universal Design for Learning]]. [[digital-divide|Digital Divide]] covers the access-and-skills layers. [[neurodiversity]], [[special-education|Special Education]], [[accessibility]], and [[multilingual-learning|Multilingual Learning]] go deep on single groups. What is left for this page is the cross-group evidence map and the appraisal question: how these effects are estimated, which groups are studied at all, and what a group-level finding licenses you to do. ## How group differences get reported, and why most of it cannot **A single-group study is not a differential-effect study.** Most of the literature in this area measures one group in the absence of a comparison group. [[zhang-ai-students-disabilities-meta-analysis-2024|Zhang et al. (2024)]] pooled 29 (quasi-)experimental studies of AI for students with disabilities and found a medium positive effect (Hedge's g = 0.588, 95% CI [0.349, 0.826]) — with no neurotypical comparator. That tells you an intervention helped, not that it helped this group differently. **Subgroup analyses are usually too small to answer the question.** The [[ai-tutoring-micro-rct-gcse-science-2026|GCSE science micro-RCT]] is unusually explicit: its treatment-by-status interaction was 0.57 marks (95% CI −2.25 to 3.39), with stratified estimates of g = 0.28 (95% CI −0.04 to 0.59) for one group and g = 0.35 (95% CI 0.18 to 0.52) for the other. A subgroup interval that overlaps zero is a question for a local pilot, not a basis for a class-wide rule. **Demographic attributes are usually simulated.** Real learner demographics are rarely attached to model inputs, so audits supply personas instead. [[demographic-signals-llm-student-assessment-2026|Rooein, Benedetto and Hovy (2026)]] ran six instruction-tuned models across three educational tasks under model-default, explicit-persona, and implicit-history conditions — 192,480 inference calls — and found that both explicit personas and *implicit* conversation histories move model behavior. Their distinction is worth keeping: awareness of learner differences can be desirable (adjusting feedback for a learner's first language), while the same sensitivity is a harm when it shifts the judgment of identical work. **Fairness fixes may not generalize.** [[student-attention-estimation-fairness-2026|Fragkiadakis et al. (2026)]] reduced gender- and age-targeted error gaps in student-attention estimation on [[assessment-validity|validation]] data, but the gains did not consistently transfer to held-out subjects or repeated subject-level splits. **Group labels hide the variation inside them.** In the same disability meta-analysis, students with specific learning disabilities, intellectual and developmental disabilities, or who are deaf showed a larger effect (g = 0.952) than students with autism spectrum disorder (g = 0.368) — a difference the authors report as not statistically significant across 239 effect sizes from 41 independent samples. "Disability" is not one group, and neither is "L2 learner". ## Disability and neurodivergence This is the deepest strand in the knowledge base, and the one where the reporting is strongest. - **Effects are real but unevenly distributed by outcome.** In [[zhang-ai-students-disabilities-meta-analysis-2024|Zhang et al. (2024)]], [[learning-gains|academic performance]] showed the largest effect (k = 80, g = 0.929), ahead of daily-life and other skills (g = 0.766) and social-emotional skills (k = 144, g = 0.382). Publication bias was present (Egger's test β = 2.837, p < .001); trim-and-fill reduced the pooled estimate to g = 0.2694, which remained statistically significant. Read the headline effect and the bias-adjusted one together. - **The field has reorganized around [[generative-ai|generative AI]].** The [[assistive-tech-neurodivergent-higher-ed-review-2026|scoping review of digital assistive technologies for neurodivergent students]] screened 766 records to include 40 empirical studies, of which 15 used generative AI and 11 used immersive formats. The strand is small and recent rather than mature. - **Some AI supports equalize rather than differentiate.** [[adhd-video-segmentation-computing-education|Pimenova, Begel and colleagues]] segmented [[video-education|instructional videos]] into single-instruction chunks with fixed pauses in a within-participants study (17 ADHD, 10 non-ADHD); everyone improved, and ADHD participants' errors and hesitations fell to parity with their non-ADHD peers. That is the strongest available shape of evidence for [[universal-design-for-learning|Universal Design for Learning]] through automated content transformation: a general change that closes a gap, rather than a group-targeted fix. - **What neurodivergent students say they need is often mundane.** [[neurodivergent-computing-students|A survey of 24 neurodivergent computing students and 20 neurotypical peers]], with four interviews, found significant discomfort with assignments that lack clear structure or carry ambiguous expectations — an accommodation that costs nothing to provide. - **Design choices can also exclude epistemically.** [[genai-minoritized-knowledges-disability|Tali-Otmani (2026)]] argues that Anglophone, Western-centric training data marginalizes non-hegemonic ways of knowing and puts the situation of disabled learners at the center of that critique. ## Language: second-language, multilingual, and English learners By article count this is the largest strand, and it splits cleanly into tool effects and tool harms. - **Tool effects are promising and noisy.** [[robot-assisted-language-learning-meta-analysis-2026|Wang, Zhang and Zou (2026)]] meta-analyzed 11 studies (17 effect sizes, N = 595) of robot-assisted [[language-learning|language learning]] and found a positive overall effect (g = 0.83, 95% CI [0.46, 1.21]) with high heterogeneity (I² = 84.4%); of six moderators tested, only robot-learner interaction type was significant. An evidence base this size supports a provisional [[benchmark]], not a procurement decision. - **The infrastructure itself is uneven before any tool is used.** [[structural-silence-underrepresented-language-ai-2026|Roy and Roy (2026)]] document the corpus gap with Bengali as the case: under 0.5% of global web content against roughly 49.5% for English, despite Bengali speakers being nearly 4% of the world's population. - **Language of instruction changes outcomes across 294 higher-education students.** The same review reports that foreign-language content yielded lower outcomes than native-language instruction, and that bilingual [[cs-education|programming]] instruction outperformed English-only instruction. - **Detectors penalize second-language writers.** [[hadra-ai-detector-accuracy-efl-2026|Hadra, Cambridge and Mesbah (2026)]] tested Turnitin and Originality against 192 texts: one detector classified 48 of 48 professionally authored texts correctly but misclassified four of 48 EFL student texts (91.6%). The asymmetry is the point — the error lands on the group whose writing is already being scrutinized. ## Gender Gender research here divides into whether tools treat learners differently, and whether learners are exposed to stereotype-laden tooling. - **A deliberately gender-neutral design showed no gender difference.** [[ada-female-coded-chatbot-gender-stereotypes-2026|A quasi-experimental study of 195 ninth-grade students]] tested ADA, a female-coded [[conversational-ai|chatbot]] grounded in Ada Lovelace: situational interest rose for both genders with no gender difference in emotional response, cognitive load, or academic performance. A role-model persona can be built without triggering stereotype threat. - **But prompt content transfers bias into student work.** [[gender-bias-transfer-llm-writing|A controlled study with 123 participants]] had students write career-plan essays for paired profiles differing only in gender, under no-AI, neutral-AI, and gender-biased-AI conditions; the biased condition transferred gender-differentiated language into student writing and suppressed female [[agency]] asymmetrically. The researchers first confirmed the effect on 1,600 generated essays. - **Space and framing matter as much as the tool.** [[all-girls-genai-makerspace-gender-equity-2026|An all-girls generative AI makerspace case study]] found girls valued the single-gender setting as safer and more relaxed, and warns against "girlification" — surface-level adaptation that leaves power relations untouched. - **Model sensitivity is a model property, not a constant.** In [[edufair-bench-pedagogical-fairness-llm-tutors-2026|EduFair-Bench]], five tutors from 7B to 70B were audited across nine demographic levels: Qwen2.5-7B exceeded the |r| ≥ 0.10 bias threshold in 7 of 12 domain-by-dimension cells, while LLaMA-3.1-8B exceeded it once. Pedagogy-specific training reduced some biases and increased others. ## Race, ethnicity, and minoritized students This strand is small in article count and strong in mechanism, because the evidence is about what AI does *with* a demographic attribute once it has one. - **[[personalized-learning|Personalization]] is a bias vector.** [[marked-pedagogies-linguistic-bias-writing-feedback|Marked Pedagogies]] shows [[llm]] writing-feedback tools shifting toward stereotype-aligned praise and withheld critique when feedback is personalized with a student's race, language, disability, achievement, or [[motivation|motivation]] — on identical essays. - **The demographic cue can be implicit.** The counterfactual audit above found that conversation histories, not just explicit personas, moved scoring and feedback behavior — the same [[demographic-signals-llm-student-assessment-2026|192,480 inference calls]] result, and a reason a disclosure rule about *declared* demographics is insufficient. - **Migration and language gaps can dominate.** In EduFair-Bench, the 70B model paired the largest language and immigration gaps with the smallest pedagogy gaps, so a tutor that looks pedagogically strong can be the one most sensitive to who the student appears to be. - **Epistemic exclusion, not just error:** see [[genai-minoritized-knowledges-disability|the marginalization of minoritized knowledges]] above. ## Socioeconomic status, geography, and age - **Digital literacy, not AI usage, is the mediator.** [[ai-divide-ses-personality-primary-education-2026|Wang and colleagues (2026)]] modeled survey and national registry data from 4,497 Grade 6 students in the Netherlands and found the link between personality traits and academic performance ran through digital literacy rather than AI usage intensity, with differences in digital literacy driven more by personality than by socioeconomic status — and SES advantages operating independently of AI engagement. The classic SES-only framing is incomplete. - **Geography can be the binding constraint.** [[arc-hubs-k12-ai-robotics-rural-2026|ARC's account]] of [[k-12]] robotics and AI education reports that rural FIRST LEGO League participation fell in the 2020 remote season and never recovered while urban participation gradually did, and identifies sustained local technical mentorship — not kits or curriculum — as the constraint that is distributed geographically. - **Adult learners are a separate design case,** covered by [[adult-learning|Adult Learning]] and the knowledge base's andragogy work rather than by K-12 studies. - **Access still gates everything else:** see [[digital-divide|Digital Divide]] and the finding that [[access-not-enough-ai-tutoring-2026|access to AI tutoring is not enough]] without [[pedagogy|pedagogical]] integration. ## Where the evidence is missing The honest finding of this survey is how thin some groups are. Each of these is a real gap in the evidence base, not a gap in this page. - **First-generation students:** one study reports a usage difference rather than an effect. In [[student-ai-inquiry-types-cs2-2026|a CS2 inquiry study]], continuing-generation students treated the AI as an active [[problem-solving]] partner, while first-generation students took a confirmatory, validation-oriented role and asked fewer questions overall. - **International students:** one [[mixed-methods-research|mixed-methods]] study (survey n = 60, interviews n = 14) on [[international-students-conversational-ai-adaptation|cross-cultural adaptation support]]. - **Gifted and high-achieving students:** effectively unstudied as a group in this corpus. - **Refugee, immigrant, and displaced learners:** no studies. - **Gender is still analysed as binary** in most of the work above, and disability categories vary across studies, so cross-study comparison of "the same" group has limits. ## Using this evidence without overfitting a group label - **Treat group means as hypotheses about a population, never as predictions about a person.** Every differential claim above is a distributional statement. - **Ask whether your learners were in the validation sample at all** before trusting a tool's claimed inclusivity; the neurodivergence and language strands both show that representation in development is the exception. - **Prefer equalizing designs where the evidence allows.** The ADHD video-segmentation result — everyone improves, the gap closes — is a better-fitting goal for a general classroom than a group-targeted add-on. - **Be careful with personalization that consumes demographic attributes.** The [[marked-pedagogies-linguistic-bias-writing-feedback|marked pedagogies]] finding is directly about that mechanism, and it applies to [[well-being|wellbeing]] check-ins, [[affective-computing|affective tutors]], and [[recommender-systems-and-learning-paths|recommender systems]] that profile learners. - **Test locally, on the group you care about, with an outcome measured without the tool.** See [[interpreting-and-applying-aied-research|Interpreting and Applying AIEd Research]] for the appraisal habits and [[research-methods-aied|Research Methods in AI in Education]] for the designs that make a local test defensible. ## Connected Concepts - [[equity-in-ai-education]] - [[inclusive-learning]] - [[digital-divide]] - [[neurodiversity]] - [[accessibility]] - [[special-education]] - [[multilingual-learning]] - [[language-learning]] - [[bias-mitigation]] - [[culturally-relevant-pedagogy]] - [[universal-design-for-learning]] - [[assistive-technology]] - [[global-south]] - [[personalized-learning]] - [[learners]] - [[learner-identity]] - [[interpreting-and-applying-aied-research]] ## Connected Articles - [[zhang-ai-students-disabilities-meta-analysis-2024]] — 29 studies of AI for students with disabilities: g = 0.588, and g = 0.269 after bias adjustment - [[assistive-tech-neurodivergent-higher-ed-review-2026]] — 766 records to 40 studies, 15 of them generative AI - [[adhd-video-segmentation-computing-education]] — A general design change that brought ADHD participants to parity - [[neurodivergent-computing-students]] — 24 neurodivergent students on structure, ambiguity, and collaboration - [[genai-minoritized-knowledges-disability]] — Epistemic marginalization in AI, with disability as the case - [[robot-assisted-language-learning-meta-analysis-2026]] — g = 0.83 with I² = 84.4% in a small L2 evidence base - [[structural-silence-underrepresented-language-ai-2026]] — Bengali's under 0.5% of web content against English's 49.5% - [[hadra-ai-detector-accuracy-efl-2026]] — Detectors misclassify EFL writing while scoring professional texts perfectly - [[ada-female-coded-chatbot-gender-stereotypes-2026]] — A 195-student test of female-coded role models with no gender difference - [[gender-bias-transfer-llm-writing]] — Gender-biased prompts transferring into student essays - [[all-girls-genai-makerspace-gender-equity-2026]] — Single-gender space valued, with a warning against girlification - [[edufair-bench-pedagogical-fairness-llm-tutors-2026]] — Tutor fairness varying by model, domain, and behavior dimension - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Stereotype-aligned feedback on identical essays - [[demographic-signals-llm-student-assessment-2026]] — 192,480 calls: explicit personas and implicit histories both move models - [[student-attention-estimation-fairness-2026]] — Fairness regularization that did not generalize - [[ai-divide-ses-personality-primary-education-2026]] — 4,497 students: digital literacy, not AI use, mediates the gap - [[arc-hubs-k12-ai-robotics-rural-2026]] — Rural participation that never recovered, and mentorship as the constraint - [[ai-tutoring-micro-rct-gcse-science-2026]] — A subgroup interaction whose confidence interval crosses zero - [[student-ai-inquiry-types-cs2-2026]] — The corpus's one first-generation usage finding - [[international-students-conversational-ai-adaptation]] — The corpus's one international-student study --- ## [Digital Divide](https://edtechdev.github.io/aied/concepts/digital-divide/) > **Digital divide** — the unequal distribution of access to, skills for, and benefits from digital (and increasingly AI) [[ai-technologies|technologies]] across individuals, communities, and nations. In AI education, the digital divide is a central equity concern: [[generative-ai|generative AI]] is rapidly reshaping learning, and the gap between those who can use it effectively and critically and those who cannot threatens to deepen existing educational inequalities. ## Questions to Consider - The digital divide is often described in three levels: access, skills, and who actually benefits. Which level do you think most people are thinking about when they say 'closing the digital divide' — and why might that be incomplete? - Giving every student a device and internet access doesn't automatically mean they can use AI effectively or critically. What separates having access from being able to benefit? - AI adds new layers to inequality: algorithmic bias can disproportionately harm marginalized [[learners]], and AI literacy itself determines whether the technology widens or narrows gaps. How is this different from the divides of earlier technologies? - The divide also extends to WHICH communities, languages, and perspectives are represented in and served by AI systems. How is representation itself a form of access — or exclusion? - If closing the divide is 'a question of justice and participation,' who bears the responsibility: platforms, schools, governments, or all of them — and what would a fair distribution of AI's benefits actually look like? ## Introduction The digital divide is commonly understood as operating across **three levels** (van Deursen & van Dijk, 2014): the *first-level* divide concerns access to technologies and infrastructure (connectivity, devices, supportive environments); the *second-level* divide concerns skills and competencies (the uneven capacity to use tools effectively and meaningfully); and the *third-level* divide concerns outcomes and benefits (who actually benefits from technology use, with AI potentially exacerbating social, cultural, and economic disparities). [[framing-ai-use-for-students|Framing AI]] literacy through this lens makes clear that equity requires more than closing the device-and-infrastructure gap — it requires building the skills to use AI effectively and critically so that its benefits are distributed fairly rather than reinforcing existing inequalities. ### How the digital divide appears in the research - **AI literacy as a mechanism for equity:** [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl|The SAIL framework]] was explicitly designed to address second- and third-level divides, providing a scaffolded, age-agnostic pathway for equitable AI literacy across all stages of education, grounded in the argument that AI literacy is inseparable from equity and participation. - **Policy and infrastructure:** [[oecd-digital-education-outlook-2026|OECD Digital Education Outlook 2026]] situates the digital divide within national [[educational-policy-ai|education policy]], examining how access to digital and AI technologies varies and what systems can do to close gaps. - **Responsible-use and [[prompt-engineering|prompting]] literacy:** [[aaai2026-prompting-literacy-k12|K-12 prompting-literacy research]] addresses the second-level divide by [[teacher-role|teaching]] students the skills to use AI [[conversational-ai|chatbots]] responsibly, recognizing that access alone does not confer the ability to [[ai-literacy|use AI well]]. - **Representation and structural silence:** [[structural-silence-underrepresented-language-ai-2026|Research on underrepresented languages]] highlights how the digital divide extends to *which* communities, languages, and perspectives are represented in and served by AI systems — a cultural and epistemic dimension of inequality. ### AI deepens (and can close) divides AI adds new layers to the equity implications of technology. Algorithmic bias can disproportionately impact learners from marginalized communities, and [[ai-literacy|AI literacy]] — the ability to understand, critically evaluate, and mitigate AI's biases and risks — is itself a key factor in whether AI widens or narrows gaps. [[research-methods-aied|Research]] shows educators with higher AI literacy are more effective at identifying and mitigating biased outcomes. The digital divide in the AI era is therefore not simply a technical provision problem but a question of justice and participation: who can access AI, who can use it critically, and who benefits. **Personality, not just SES, shapes the AI-era divide.** [[ai-divide-ses-personality-primary-education-2026|Wang et al. (2026)]] analyzed survey and national registry data from **4,497 Grade 6 students** in the Netherlands, separating two mediating pathways — AI usage and digital literacy — linking student background and personality to [[learning-gains|academic performance]]. Their key finding reframes the classic divide: **digital literacy, not AI usage intensity, mediates** the link between personality and performance, and a new digital-skills divide emerges that is driven more by **personality traits than by socioeconomic status**. SES advantages on performance operated independently of AI [[student-engagement|engagement]]. This complicates the access-and-SES framing of the digital divide, pointing to skills formation and dispositional support as equity-relevant levers alongside device and tool access. **The divide operates at the institutional level too.** [[adeniranye-ai-integration-nigerian-higher-education-2026|Adeniranye et al. (2026)]] show that in Nigeria's higher education system, AI integration capacity concentrates in older, South-West-region institutions and compounds through mutually reinforcing network ties (international collaborations × industry partnerships, r = 0.74) — meaning institutional "have-nots" (typically newer state universities) face structural barriers to entering the very networks that would help them catch up. Digital inequality is thus reproduced not only across individual learners but across the [[governance|institutional]] structures that shape who can participate in an AI-transformed knowledge economy. **The divide also has a geography, and mentorship is its mechanism.** [[arc-hubs-k12-ai-robotics-rural-2026|Jacobson et al. (2026)]] document a recovery asymmetry in Indiana FIRST LEGO League participation: both urban and rural participation fell in the 2020 remote season, but only urban participation recovered, and rural participation stayed near its post-2020 level through 2025–2026. The mechanism they name is access to technical mentorship — people with enough programming and [[educational-robotics|robotics]] knowledge to start and sustain a team — which rural schools may lack even when students and teachers are interested, making it a precondition for the robotics and AI pathway rather than a feature of it. Their response is to engineer the propagation of that mentorship: college primary hubs train undergraduates and host workshops, mature school programs become secondary hubs that mentor nearby schools, and a spatial Markov simulation of Indiana's 1,925 public schools projects 992 programs after 40 years under moderate assumptions against 161 without ARC, including 341 rural programs against 60. Read alongside the institutional-network result above, the pattern is that divides persist through the *structure of who can supply expertise where*, and that supplying it deliberately — rather than assuming proximity to a university — is the policy lever. **Access and epistemic hierarchy are different problems.** [[beyond-the-algorithm-academic-developers-digital-mediators-2026|Sithole (2026)]] draws the distinction sharply from interviews with [[educational-development|academic developers]] at two South African Historically Disadvantaged Institutions: digital inequality is distributive — devices, connectivity, budgets, digital literacy — and answerable in principle through redistribution, whereas **algorithmic coloniality** is epistemic and persists even under conditions of full access, because it inheres in what the systems encode and whose knowledge they center. The study's participants experience both at once, described as being "asked to build a digital future on analogue foundations": the foundations name the material register of the divide, and the imported future arrives pre-loaded with the epistemic assumptions of the contexts that designed it. The practical implication for equity work is that closing an access gap does not by itself unsettle the hierarchy — the two phenomena operate at different registers and require different responses. ### Connections to related concepts The digital divide is a core concern of [[equity-in-ai-education]] research, closely tied to [[ai-literacy]] (which is positioned as a central mechanism for addressing structural barriers), and to [[ethics]] and [[bias-mitigation]] (since algorithmic bias disproportionately affects marginalized groups). It connects to [[ai-education]] and [[higher-ed]] as the settings where access and capability gaps manifest, and relates to [[student-experience]] as it shapes who can participate meaningfully in AI-shaped learning. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[remote-proctoring]] - [[equity-in-ai-education]] - [[ai-literacy]] - [[ethics]] - [[bias-mitigation]] - [[ai-education]] - [[higher-ed]] - [[student-experience]] - [[parents-and-families]] ## Connected Articles - [[adeniranye-ai-integration-nigerian-higher-education-2026]] — Institutional structures, digital inequality, and AI integration in Nigerian higher education - [[ai-divide-ses-personality-primary-education-2026]] — SES, personality, and AI divides in primary education (Wang et al. 2026) - [[school-ai-education-readiness-gaps-agency-2026]] — School AI education narrows psychological but not cognitive readiness gaps - [[academic-dishonesty-automated-proctoring-ai-2026]] - [[prompt-privilege-equitable-ai-access-2026]] — Prompt Privilege: measuring & mitigating accessibility disparities in LLM access - [[the-scaffolded-ai-literacy-sail-framework-results-of-a-delphi-study-for-equitabl]] — The Scaffolded AI literacy (SAIL) framework - [[oecd-digital-education-outlook-2026]] — OECD Digital Education Outlook 2026 - [[aaai2026-prompting-literacy-k12]] — Teaching Responsible Use of AI Chatbots to K-12 Students - [[structural-silence-underrepresented-language-ai-2026]] — Structural Silence and Underrepresented Languages - [[sec-ai-literacy-narrative-review-2026]] — Social-Emotional Competence in AI Literacy - [[bilingual-llm-lecture-companion-srl-2026]] - [[multilingual-adaptive-learning-nigeria-2026]] — AI-Based Adaptive Learning Platform for Multilingual Low-Resource Contexts - [[ai-science-chemistry-education-systematic-review-2025]] — Systematic review of AI in science/chemistry education - [[unesco-ai-guidelines-chemical-education-2026]] — UNESCO AI guidelines translated to chemical education; epistemic drift - [[lodge-loble-cognitive-offloading-2026]] — AI, cognitive offloading and implications for education (Lodge & Loble 2026) - [[mechanical-compliance-human-flourishing-ai-literacy-2026]] — Socialist humanist AI literacy + fair use - [[arc-hubs-k12-ai-robotics-rural-2026]] — ARC: rural robotics access follows mentorship geography, not device access (Jacobson et al. 2026) - [[beyond-the-algorithm-academic-developers-digital-mediators-2026]] — Digital inequality as distributive problem vs. algorithmic coloniality as epistemic one, in South African HDIs --- ## [Bias Mitigation](https://edtechdev.github.io/aied/concepts/bias-mitigation/) > **Bias mitigation in AI education** — the identification, measurement, and reduction of unfair, identity-patterned behavior in [[intelligent-tutoring|AI tutors]], scorers, recommenders, and educational systems. Bias can enter at any stage of the AI pipeline — training data, model behavior, prompts, scoring, and deployment — and manifest as differential treatment of learners based on language, gender, race, culture, or other identity characteristics. Mitigation spans data curation, debiasing algorithms, [[prompt-engineering|prompt design]], fair-scoring methods, explainability, and evaluation. It is the technical counterpart to [[equity-in-ai-education]] and a core concern of [[ethics]] in AI education. ## Questions to Consider - Bias can enter at any stage of the AI pipeline — training data, model behavior, prompts, scoring, and deployment. Before reading, where in that chain did you expect bias to live? This page suggests it can appear almost anywhere. Where is one place you hadn't considered? - [[research-methods-aied|Research]] shows AI [[physics-education|physics]] scoring systematically underestimates students whose text-based explanations are of lower linguistic quality — the AI scores the language, not the understanding. Why might a system that agrees well with human raters overall still consistently penalize non-native or less fluent writers? - One study found that a gender-biased prompt induces students' essays to display a larger 'agentic gap' and more gender-stereotypic content — bias transferred from the tool into the learner's own work. What does this say about bias as not just an unfair score but a force that can reshape what students produce and who they see themselves as? - Another study showed LLMs shift feedback in stereotype-aligned ways when personalized with student attributes — overusing praise and withholding critique for 'marked' students even on identical essays. How might 'nicely' biased feedback be more harmful than an obviously wrong score, because it's harder to detect? - Mitigation spans data curation, debiasing algorithms, neutral prompt design, fair-scoring methods, explainability, and human oversight. Which single mitigation lever do you think would make the biggest difference in an AI system you rely on, and what would you need to audit to know it worked? - A neutral prompt largely avoids inducing gender-differentiated language — suggesting prompt design is a practical mitigation. But if bias can be reintroduced through data, scoring, or deployment, why might fixing the prompt alone be an incomplete answer? ## Introduction Bias mitigation matters because [[ai-education|AI in education]] is not neutral: systems trained on dominant language and cultural data can systematically disadvantage marginalized learners, from AI-based scoring that penalizes non-native writers to [[llm|LLM]] tutors that answer differently for different groups. Bias is a cross-cutting concern that appears in [[automated-assessment|Automated Grading]], [[automated-essay-scoring]], [[knowledge-tracing]], recommendation systems, and [[conversational-ai|conversational AI]] tutors. ## Sources of bias The knowledge base's research documents bias entering at multiple points in the pipeline: - **Language and scoring bias:** [[ai-scoring-language-bias-physics|AI-based physics scoring]] systematically underestimates the conceptual understanding of students whose text-based explanations are of lower linguistic quality — the AI scores the language, not the understanding, penalizing non-native or less fluent writers. This is a direct validity and fairness failure in [[automated-assessment|Automated Grading]]. - **Gender bias transfer in LLM-assisted writing:** [[gender-bias-transfer-llm-writing|Contaminated Collaboration]] shows that when students write with a gender-biased LLM prompt, their essays display a significantly larger agentic gap and more gender-stereotypic occupation suggestions (N=123); bias transfer is asymmetric, suppressing agency in female-target essays. A verification study (N=1,600 LLM essays, R²=.399) confirms a gender-biased prompt induces gender-differentiated language. - **Differential refusals and epistemic injustice:** [[paternalistic-filter-llm-history-education|The Paternalistic Filter]] audits four LLMs as history tutors (1,800 responses) and exposes a "paternalistic filter": models differentially refuse, soften, or reframe sensitive content for different learners — an epistemic injustice with direct equity implications. - **Selection bias in [[learning-analytics|learning analytics]]:** [[temporal-smoothness-debiased-kt|Debiased knowledge tracing]] addresses selection bias arising from non-random exercise recommendations: training on observed logs with standard empirical risk produces biased mastery estimates that compound errors in adaptive recommendation loops. - **Data and annotation bias:** [[data-annotations-pedagogical-hints|data annotations]] and [[ground-truth-reliability-aied|ground-truth reliability]] research examine how the labels and inter-rater reliability underlying AI models carry bias — arguing against treating κ > 0.8 as a binary stamp of approval. - **Marginalized knowledges:** [[genai-minoritized-knowledges-disability|Generative AI and minoritized knowledges]] documents how training data and model behavior marginalize non-dominant knowledge systems and disability perspectives. - **Stereotype-aligned automated feedback (Marked [[pedagogy|Pedagogies]]):** [[marked-pedagogies-linguistic-bias-writing-feedback|Tan et al. (2026)]] show four widely used LLMs systematically shift writing feedback in stereotype-aligned ways when feedback is personalized with student attributes — race, ethnicity, ELL designation, learning disability, achievement, or motivation — producing positive feedback bias and feedback withholding bias (overuse of praise, less substantive critique, assumptions of limited ability) for marked students even on identical essays. The "Marked Words" concentration metric offers a concrete method for auditing such bias in automated feedback. - **Visual bias in text-to-image tools:** [[bias-representation-text-to-image-education-2026|Alon, Hadar Shoval, and Levkovich (2026)]] [[meta-analysis-systematic-review|systematically review]] 31 peer-reviewed studies (2023–2025) on bias and representation in educational uses of AI-generated text-to-image. Using a six-part analytic framework (gender; race, ethnicity, and SES; culture and religion; age; body and (dis)ability; content), they find biased representation pervasive — images frequently centered white, male, Western, thin, and non-disabled figures, while diversity related to age, body, and ability was largely overlooked. Most studies relied on image audits and [[qualitative-research|qualitative]] methods, with few experimental or intervention-based designs, revealing significant blind spots in how educational research measures and responds to visual bias. - **Non-discrimination as a core ethical value.** [[agarwal-ethical-values-norms-aied-2026|Agarwal et al. (2026)]], a [[meta-analysis-systematic-review|systematic review]] of 25 articles, identify non-discrimination (definitions using bias/discrimination/diversity) as one of six main ethical values for [[ai-education|AI in education]], alongside data stewardship, human oversight, goodwill, explicability, and educational aptness. The review notes the values are tightly coupled and can conflict — e.g., non-discrimination vs. data stewardship — producing ethical dilemmas, and that no norms on non-discrimination address end users directly, leaving learners largely passive in the ethical literature. ## Mitigation approaches The knowledge base's research illustrates several complementary strategies: - **Fairness-aware modeling:** [[fair-explainable-edu-recommendations|The Hybrid HKG-GRU framework]] integrates **Group Distributionally Robust Optimization (GroupDRO)** for fairness alongside explainability and counterfactual stability, evaluated on Moodle logs (152 students, ~150k interactions). It demonstrates that recommendation systems can be trained to be fair and transparent, not just accurate. - **Debiasing estimators:** [[temporal-smoothness-debiased-kt|Temporal Smoothness Doubly Robust (TSDR) learning]] combines a propensity model with an error-imputation model, retaining unbiasedness if either is correct, to remove selection bias from knowledge-tracing mastery estimates. - **Prompt-level mitigation:** [[gender-bias-transfer-llm-writing|the gender-bias study]] shows a neutral prompt largely avoids inducing gender-differentiated language, so prompt design is a practical mitigation lever. - **Validated, language-independent scoring:** addressing [[ai-scoring-language-bias-physics|scoring bias]] requires scoring that separates conceptual understanding from linguistic quality, and auditing scores for language bias. - **Explainability:** [[xai-education-framework|XAI in education]] provides transparency into why a system produced a given score or recommendation, enabling detection and correction of biased behavior and supporting [[trust]]. - **Pipeline-wide auditing:** [[antiskillbench-persona-skills-privacy-2026|persona-skills auditing]] and systematic audits like the paternalistic-filter study show the value of auditing models across identity conditions before deployment. ## Mitigation across the AI pipeline Bias mitigation is not a single fix but an ongoing process spanning the pipeline: 1. **Data curation** — diversify training data and audit labels for identity-based gaps and unfair annotations. 2. **[[pedagogical-llm-training|Model training]]** — apply debiasing and fairness-aware objectives (e.g., GroupDRO, doubly robust estimators). 3. **Prompt and system design** — design neutral prompts and systems that do not differentially respond to [[learner-identity|learner identity]]. 4. **Scoring and assessment** — validate that automated scoring measures understanding rather than language or demographic proxies. 5. **Evaluation and auditing** — audit models across identity conditions (language, gender, culture) and require explainability to surface bias. 6. **Human oversight** — retain [[human-in-the-loop-ai|human-in-the-loop]] review, especially for low-confidence or high-stakes cases. ## Relationship to related concepts Bias mitigation is the technical mechanism through which [[equity-in-ai-education|Equity]] is operationalized, and a core requirement of [[ethics]] and responsible AI design. It connects to [[ai-ed-evaluation]] (bias as an evaluation criterion), [[educational-measurement]] and [[assessment-validity]] (fairness in scoring), and [[privacy]] (as a related responsible-AI concern). It also connects to [[cognitive-offloading|Over-Reliance]] (since biased systems are especially harmful when over-trusted) and [[ai-literacy]] (helping users recognize and question biased AI). ## Implications for AI in education - **Audit the whole pipeline:** bias can enter at data, model, prompt, scoring, and deployment stages — mitigate across all of them. - **Test across identity conditions:** evaluate AI tutors, scorers, and recommenders for differential behavior across language, gender, culture, and disability. - **Separate understanding from language in scoring:** automated scoring must not penalize non-native or less fluent writers for conceptual understanding they demonstrate. - **Make systems explainable:** transparency into AI decisions is essential for detecting and correcting bias. - **Combine technical and human mitigation:** pair debiasing algorithms with human-in-the-loop oversight, especially for high-stakes or low-confidence cases. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[explainable-ai]] - [[guardrails]] - [[equity-in-ai-education]] - [[ethics]] - [[ai-ed-evaluation]] - [[automated-assessment]] - [[automated-essay-scoring]] - [[educational-measurement]] - [[knowledge-tracing]] - [[llm]] - [[generative-ai]] - [[privacy]] - [[human-in-the-loop-ai]] - [[trust]] - [[cognitive-offloading]] - [[ai-literacy]] - [[student-experience]] - [[ai-education]] - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning]] - [[zhan-chapman-genai-cs-education-2026]] - [[ai-online-education-engagement-satisfaction-2026]] - [[prompt-privilege-equitable-ai-access-2026]] — Prompt Privilege: measuring & mitigating accessibility disparities in LLM access - [[nspa-neuro-symbolic-pedagogical-alignment-2026]] — Neuro-symbolic pedagogical alignment (NSPA) - [[ai-scoring-language-bias-physics]] — Language bias in AI-based scoring - [[gender-bias-transfer-llm-writing]] — Gender bias transfer in LLM-assisted writing - [[paternalistic-filter-llm-history-education]] — The paternalistic filter and differential refusals - [[fair-explainable-edu-recommendations]] — Fair and explainable educational recommendations - [[temporal-smoothness-debiased-kt]] — Debiased knowledge tracing - [[ground-truth-reliability-aied]] — Modernizing ground truth for AI reliability - [[data-annotations-pedagogical-hints]] — Data annotations as pedagogical hints - [[xai-education-framework]] — Explainable AI in education - [[antiskillbench-persona-skills-privacy-2026]] — Persona-skills privacy and bias auditing - [[genai-minoritized-knowledges-disability]] — GenAI and the marginalization of minoritized knowledges - [[genai-higher-education-systematic-review-2026]] — GenAI in higher education: systematic review - [[marked-pedagogies-linguistic-bias-writing-feedback]] — Marked Pedagogies: stereotype-aligned biases in automated writing feedback - [[lopez-pernas-llm-appropriate-student-support-2026]] — Can AI deliver appropriate support for diverse student profiles? A large-scale evaluation - [[bias-representation-text-to-image-education-2026]] — Bias and representation in AI-generated text-to-image: systematic review (Alon et al. 2026) - [[agarwal-ethical-values-norms-aied-2026]] — Ethical values and norms for AI in education - [[llm-grade-bands-calibration-bias-2026]] — Can large language models reproduce higher education grade bands? Cross-model study of calibration and grading bias in authentic student writing --- ## [Culturally Relevant Pedagogy](https://edtechdev.github.io/aied/concepts/culturally-relevant-pedagogy/) > **Culturally relevant pedagogy** — introduced by Gloria Ladson-Billings (1995), centers marginalized students' cultural references in [[curriculum-design|curriculum design]]. It rests on three pillars: **academic success** (rigorous standards that honor cultural identity), **cultural competence** (critical consciousness about culture and power), and **[[critical-pedagogy|sociopolitical consciousness]]** (empowering students to challenge inequitable systems). As AI tools enter classrooms, CRP has become a central lens for evaluating whether [[generative-ai|AI]] amplifies or erases non-dominant cultural knowledge. ## Questions to Consider - Culturally relevant pedagogy rests on academic success, cultural competence, and sociopolitical consciousness. Which of these pillars is hardest to achieve with AI tools — and why? - One study found 94% of AI-generated lesson plans contained no discernible multicultural content, and almost none reached the level of transformation or social action. If AI defaults to monocultural output, whose responsibility is it to inject the missing perspectives? - AI training data is predominantly Western and Anglophone. What does it mean for a system to 'actively marginalize' ways of knowing — and how does that differ from simply lacking access? - Community-based AI learning proposes that learners' own lived epistemologies should be the standard for evaluating AI output, with refusal and non-use treated as valid responses. How would you center community knowledge as the judge of an AI's relevance and harm? - Culturally grounded data can dramatically improve an AI's relevance — one Indian-knowledge dataset moved a small model from near zero to rivaling a far larger general-purpose one. If better data is the fix, who should build and own it? - A cross-cultural study found that identical AI-use behaviors were judged [[ethics|ethical]] in one country and unethical in another, regardless of written policy. What does that tell you about trying to govern AI use with uniform rules? ## Introduction ### AI's Double-Edged Role in CRP AI can support teachers in making instruction culturally responsive, but its default outputs also risk **reinforcing dominant narratives** when left unprompted. - **Support for teachers:** Wang et al. (2025) built **CulturAIEd**, an [[llm|LLM]]-powered system that helps [[k-12|K-12]] teachers design culturally responsive [[ai-literacy]] activities by combining student demographic information with rubric-driven guidance (a CRT checklist layered into generation). In a four-teacher pilot it **enhanced teachers' confidence** in spotting opportunities for cultural responsiveness and in modifying existing activities, with 78% finding AI suggestions helpful for diversifying materials. The tool directly targets the time, training, and resource barriers that block CRP implementation. - **Risk of monocultural output:** Trust et al. (2025) analyzed 310 AI-generated civics lesson plans (2,230 activities): **94% contained no discernible multicultural content**, and of the 144 that did, 137 sat at the lowest "Additive" level — **only one reached "Transformation" and none reached "Social Action."** All three [[conversational-ai|chatbots]] produced structurally identical, monocultural lesson templates. This is concrete evidence that AI defaults to homogenized curricula unless [[teacher-ai-competency|teachers]] actively intervene. ### Epistemic Marginalization in AI Systems Beyond lesson generation, CRP connects to a deeper critique: AI training data and design processes encode Western, Anglophone epistemic frameworks that marginalize other ways of knowing. - **Epistemic coloniality:** Tali-Otmani (2026) argues that GenAI systems are not epistemically neutral — predominantly Western-centric training data **actively marginalizes minoritized knowledges**, producing a "double marginalization" for disabled learners whose epistemologies are both underrepresented in training data and excluded from design. This extends the [[equity-in-ai-education|equity]] conversation from *access* to *whose knowledge is validated*. - **Redistributing epistemic authority:** Ojeda-Ramirez, Gyles & Peppler (2026) propose **community-based AI learning**, a framework that repositions learners' lived and community-based epistemologies as the evaluative standard over AI outputs. Its three commitments — **epistemic fine-tuning**, **redistribution of authority**, and **[[situated-learning|situated]] discernment** — calibrate trust against local histories and community expertise, treating refusal and strategic non-use as valid CRP responses to AI. ### Culturally Grounded Data and Evaluation A body of knowledge base-sourced work addresses the *content* and *evaluation* gaps behind CRP. - **Non-Western training data:** IKS-Instruct provides a **24,795-example [[multilingual-learning|multilingual]] instruction dataset** for [[teacher-role|teaching]] LLMs Indian Knowledge Systems across seven [[language-learning|languages]] and 41 [[pedagogy|pedagogical]] techniques. A compact domain-tuned 7B model reached a median judge score of 6.39 (vs. 6.54 for a far larger general-purpose model) — while the base model scored **near zero** on IKS-specific dimensions, showing how much culturally grounded data improves relevance. - **[[global-south|Global South]] [[benchmark|benchmarks]]:** The **NSMQ Riddles** benchmark draws 1.8K scientific/mathematical riddles from 11 years of Ghana's National Science and Maths Quiz — one of the first Global South educational benchmarks — and found state-of-the-art LLMs **underperform the best student contestants**, exposing geographic bias in how models are evaluated. - **Culture over policy:** A cross-cultural survey of Canadian and South Korean [[cs-education|computing]] students found that **culture, not policy text, drove perceptions of AI-use ethicality** — identical behaviors were judged differently across cohorts, reinforcing the need for culturally aware communication rather than abstract rules. - **Cultural adaptation as design moves:** Bashir and Afzal (2026) operationalize cultural relevance in an AI [[well-being]] support system ([[culturally-aware-student-stress-chatbot-2026|Sukoon]]) through three moves: a bilingual 20-question assessment with parallel English and Urdu labels; a system prompt instructing the model to respond consistently with Pakistani social and cultural norms and to use Urdu and Roman Urdu expressions where appropriate; and explicit sensitivity to locally salient stressors (family expectations, financial pressure, hierarchical teacher-student relationships). Their justification is empirical as well as ethical — [[explainable-ai|feature importance]] placed teacher-student relationship second among stress predictors — but they concede the adaptation lives in the prompt rather than in the NLP pipeline, and that cultural appropriateness was assessed only by informal testing, not by the students the system targets. ### Practical Guidance Grounded in the knowledge base's own articles, educators and designers can apply CRP to AI: - **Treat AI as a draft generator, not an authority.** The Trust et al. civics findings show teachers must inject [[critical-thinking|higher-order thinking]] and multicultural perspectives the AI omits; [[human-in-the-loop-ai|human judgment]] remains essential for cultural authenticity and community alignment. - **Layer demographic and cultural context into prompts and tools.** CulturAIEd and [[connected-ai-lesson-planning-vietnam|ConnectED]] (a Vietnamese, curriculum-aligned lesson-planning system) show that structured, locally grounded prompt templates plus teacher validation gates improve cultural fit over generic generation. - **Center community knowledge as the evaluative standard.** Following community-based AI learning, have learners judge AI outputs against locally grounded criteria of relevance, harm, and usefulness, and honor context where refusal or non-use is the right call. - **Adopt and evaluate culturally grounded datasets.** IKS-Instruct and NSMQ Riddles illustrate that [[discipline-specific-aied|domain-specific]], non-Western data meaningfully improves both relevance and honest evaluation. ## Connected Concepts - [[equity-in-ai-education]] - [[curriculum-design]] - [[ai-literacy]] - [[k-12]] - [[teacher-ai-competency]] - [[teacher-role]] - [[bias-mitigation]] - [[critical-pedagogy]] - [[human-in-the-loop-ai]] - [[student-experience]] - [[higher-ed]] - [[language-learning]] - [[cs-education]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education ## Connected Articles - [[llm-cultural-relevance-k12]] — LLMs for Culturally Relevant K-12 Pedagogy - [[civic-education-ai-lesson-plans]] — AI-Generated Lesson Plans in Civic Education - [[ojeda-ramirez-community-based-ai-learning]] — Community-Based AI Learning - [[genai-minoritized-knowledges-disability]] — Generative AI and the marginalization of minoritized knowledges - [[iks-instruct-dataset-indian-knowledge]] — IKS-Instruct: Indian Knowledge Systems Dataset - [[nsmq-riddles-science-math-benchmark]] — NSMQ Riddles: Ghana STEM Benchmark - [[cross-cultural-student-perceptions-genai-computing]] — Cross-Cultural Perceptions of GenAI Use - [[international-students-conversational-ai-adaptation]] — International Students and Conversational AI - [[connected-ai-lesson-planning-vietnam]] — ConnectED: Vietnamese Lesson Planning - [[culturally-aware-aied-community-learning]] — Culturally-Aware AI for Community Learning - [[taklif-ai-interest-based-personalized-assignments]] — Taklif: Interest-Based Personalized Assignments - [[multilingual-adaptive-learning-nigeria-2026]] — AI-Based Adaptive Learning Platform for Multilingual Low-Resource Contexts - [[culturally-aware-student-stress-chatbot-2026]] — An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning --- ## [Multilingual Learning](https://edtechdev.github.io/aied/concepts/multilingual-learning/) > Multilingual learning in AI education concerns how educational [[ai-technologies|technologies]] and LLM-based systems support learners across languages, dialects, and low-resource linguistic contexts — and the risks of linguistic exclusion when AI systems are built primarily for dominant languages. ## Questions to Consider - Most AI models are trained primarily on high-resource languages like English. If you think, study, or are assessed in a different language, how might that systematically disadvantage you — even if the tool seems to 'work' in English? - The page warns that unaddressed monolingual bias in AI deepens the digital divide and undermines equity, especially across the Global South. What does genuine equity require beyond simply translating AI content into another language? - Automated assessment can exhibit language bias, penalizing non-native speakers even for the same reasoning. If you were implementing AI scoring, what would you check to ensure it is fair across languages rather than just accurate in one? - The page shows low-resource languages can be served by fine-tuning models on curated corpora, even under practical hardware constraints. What trade-offs would you expect between efficiency and how faithfully the model handles a low-resource language? - Multilingual AI must go beyond translation to reflect culturally relevant pedagogy — content that is linguistically *and* contextually appropriate. How might content that is perfectly translated still fail a learner if it ignores local context and culture? ## Introduction Multilingual learning concerns education for learners who study or think in languages other than the dominant ones, and it is a core equity dimension of AI in education: [[generative-ai|generative AI]] is overwhelmingly trained and tuned on high-resource languages, which can systematically disadvantage everyone else. The theme spans technical work (adapting models to low-resource languages and dialectal corpora, [[rag|retrieval]] in non-dominant languages), pedagogical concerns ([[culturally-relevant-pedagogy|locally grounded instruction]]), and structural equity (who can access useful educational AI at all) — which makes it inseparable from the [[digital-divide]] and [[equity-in-ai-education]]. ## Overview Multilingual learning is a core equity dimension of [[ai-education|AI in education]]. Generative AI and [[llm|LLMs]] are overwhelmingly trained and tuned on high-resource languages, which can systematically disadvantage learners who study or think in other languages. The theme spans technical challenges (adapting models to low-resource languages, dialectal corpora, [[rag|RAG]] in non-dominant languages), [[pedagogy|pedagogical]] concerns (culturally relevant and locally grounded instruction), and structural equity (who gets access to useful educational AI at all). ## Technical approaches - **Fine-tuning for low-resource languages:** Nwogo et al. (2026) [[multilingual-adaptive-learning-nigeria-2026|fine-tuned an instruction-tuned LLM on a curated Nigerian Pidgin corpus]] within an [[adaptive-learning]] platform, and systematically analyzed quantization (4/5/8-bit) trade-offs between semantic fidelity and computational efficiency — showing that low-resource languages can be served with practical hardware constraints. See also the [[bilingual-llm-lecture-companion-srl-2026|bilingual LLM lecture companion]] for [[self-regulated-learning|self-regulated learning]]. - **Corpus and data equity:** building curated corpora (e.g., Nigerian Pidgin, Indian knowledge systems via [[iks-instruct-dataset-indian-knowledge|IKS-Instruct]]) is a recurring strategy for enabling model output in learners' own languages. - **Voice-first and oral contexts:** [[kutti-ai-voice-first-learning-companion|voice-first companions]] and [[structural-silence-underrepresented-language-ai-2026|structural-silence analyses]] address contexts where text-based AI fails speakers of underrepresented languages. ## Equity and pedagogy Multilingual AI must go beyond translation to reflect [[culturally-relevant-pedagogy|culturally relevant pedagogy]] — generating content that is linguistically and contextually appropriate. Studies of [[llm-cultural-relevance-k12|LLM cultural relevance in K-12]] and [[scaffolding-critical-engagement-genai-minority-students|critical engagement with GenAI among minority students]] show that linguistic and cultural alignment shapes whether students actually benefit. Unaddressed, monolingual bias in AI deepens the [[digital-divide]] and undermines [[equity-in-ai-education|Equity]] across the [[global-south|Global South]]. ## Assessment bias Multilingual concerns also affect [[automated-assessment|automated assessment]]: [[ai-scoring-language-bias-physics|AI scoring can exhibit language bias]] (e.g., in [[physics-education|physics]]), penalizing non-native speakers. Ensuring assessment tools are fair across languages is part of [[assessment-validity]]. ## Implications for instructors in multilingual contexts - **Extend AI to learners' own languages, not just English.** Fine-tune or configure models for low-resource and non-dominant languages ([[multilingual-adaptive-learning-nigeria-2026|Nigerian Pidgin platform]]) rather than forcing English-only tools; pair AI with [[rag|RAG]] and local corpora where possible. - **Guard assessment against language bias.** [[ai-scoring-language-bias-physics|AI scoring]] can penalize non-native speakers — use language-aware or human-moderated evaluation to protect [[assessment-validity]] and [[equity-in-ai-education|fairness]]. - **Reflect culture and context, not just translation.** Multilingual AI must go beyond translation to [[culturally-relevant-pedagogy|culturally relevant pedagogy]] — generate content that is linguistically and contextually appropriate ([[llm-cultural-relevance-k12|K-12 cultural relevance]]). - **Pair AI with multilingual support structures.** Use voice-first and oral modes ([[kutti-ai-voice-first-learning-companion|voice-first companions]]) where text-based AI fails, and support [[self-regulated-learning|self-regulation]] in bilingual contexts ([[bilingual-llm-lecture-companion-srl-2026|bilingual lecture companion]]). - **Watch the digital divide.** Monolingual bias in AI deepens the [[digital-divide]] and undermines access across the [[global-south]] — plan for equitable infrastructure and access alongside tool choice. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[language-learning]] - [[llm]] - [[equity-in-ai-education]] - [[global-south]] - [[digital-divide]] - [[culturally-relevant-pedagogy]] - [[inclusive-learning]] - [[generative-ai]] ## Connected Articles - [[llm-comparative-judgment-writing-screening-2026]] — Validity of Large Language Model Comparative Judgment for Universal Writing Screening - [[multilingual-adaptive-learning-nigeria-2026]] — AI-Based Adaptive Learning Platform for Nigeria - [[bilingual-llm-lecture-companion-srl-2026]] — Bilingual LLM Lecture Companion - [[structural-silence-underrepresented-language-ai-2026]] — Structural Silence: Underrepresented Languages - [[llm-cultural-relevance-k12]] — LLM Cultural Relevance in K-12 - [[scaffolding-critical-engagement-genai-minority-students]] — Critical Engagement with GenAI Among Minority Students - [[iks-instruct-dataset-indian-knowledge]] — IKS-Instruct: Indian Knowledge Systems Dataset - [[kutti-ai-voice-first-learning-companion]] — Voice-First Learning Companion - [[ai-scoring-language-bias-physics]] — AI Scoring Language Bias in Physics --- ## [Inclusive Learning](https://edtechdev.github.io/aied/concepts/inclusive-learning/) > **Inclusive Learning** — the design and delivery of educational experiences that accommodate diverse learner needs, spanning physical, cognitive, sensory, and situational differences. In AI in education, inclusive learning [[research-methods-aied|research]] examines both how AI tools can remove barriers for disabled and neurodivergent learners and how [[ai-technologies|AI systems]] themselves must be designed to avoid creating new accessibility gaps. ## Questions to Consider - An accessible tool does not guarantee inclusive instruction, and assistive tech does not guarantee meaningful agency. What is the difference between removing a barrier to a format and designing education so everyone can meaningfully participate? - The page distinguishes inclusive learning, accessibility, assistive technology, special education, and universal design. Where does AI in your own context sit — and which question are you actually trying to answer? - One study found AI-segmented videos with fixed pauses eliminated the performance gap between ADHD and non-ADHD learners. How might designing for one group's needs improve learning for everyone? - AI systems are described as risking new accessibility gaps even as they remove old ones. What kind of learner might a text-based, visual, always-online AI tool silently exclude? - Inclusive assessment research exposes a tension between anti-cheating measures and accommodating learners with visual-processing needs. When security and accessibility conflict, how should the trade-off be decided — and by whom? - Several tools invert the assumption that [[edtech-platform|edtech]] must be visual — e.g., voice-first companions for visually impaired learners. What assumptions about the 'default' learner might your own tools or materials be making? ## Introduction Inclusive learning is the design commitment that education should be built so that all learners can participate meaningfully, rather than adapted after the fact for those who struggle. It functions here as an umbrella over adjacent but distinct concepts — [[accessibility]] (can everyone perceive and operate the format?), [[assistive-technology]] (what tools bridge an individual's [[digital-divide|access gap]]?), [[universal-design-for-learning]] (how should the design anticipate variability?) — and it engages [[neurodiversity]] and [[special-education]] as design contexts rather than exceptions. AI enters as both a promise (personalization, accommodation, translation) and a new source of exclusion (cost, data, language coverage, and the assumptions built into models). ## How the related concepts fit together Inclusive learning is the **umbrella** concept; the pages below sit inside it, each answering a different question. They overlap but are not interchangeable — knowing which one a claim belongs to keeps the knowledge base precise: | Concept | Core question it answers | Typical focus | |---|---|---| | **Inclusive Learning** *(this page)* | How do we design education so all learners can meaningfully participate? | The broad design of instruction across learner variability | | **[[accessibility]]** | Can everyone perceive and operate the *format/medium*? | Captions, alt text, transcripts, contrast, keyboard/screen-reader compat, WCAG | | **[[assistive-technology]]** | What tools/equipment bridge an individual's access gap? | Screen readers, TTS/STT, braille/tactile, captioning, AI accommodations | | **[[special-education]]** | How do we deliver instruction to learners with diagnosed disabilities? | IEPs, individualized accommodations, disability-specific tutoring — **primarily a [[k-12]] term (IDEA/entitlement)** | | **[[universal-design-for-learning]]** | How do we proactively build in flexibility from the start? | Multiple means of [[student-engagement|engagement]], representation, action/expression | In practice: **UDL** is the design philosophy that *prevents* barriers; **accessibility** is the property that removes *format* barriers; **assistive technology** is the *tool* layer individuals use; **special education** is the *instructional* domain for diagnosed disabilities — and it is primarily a **K-12** term, whereas in **[[higher-ed|higher education]]** (and increasingly K-12 too) the more common framing is [[universal-design-for-learning|Universal Design for Learning]]. **Inclusive learning** is the umbrella that holds them together around the shared goal of equitable participation. An accessible tool does not guarantee inclusive instruction, and assistive tech does not guarantee meaningful agency — which is why the umbrella must span all of them. Inclusive learning sits at the intersection of [[equity-in-ai-education]], [[learning-design]], and [[special-education]], and is supported by the concrete tool layer of [[assistive-technology]] and the design property of [[accessibility]]. Unlike narrow accommodations that retrofit access onto existing systems, the inclusive learning perspective — grounded in [[universal-design-for-learning|Universal Design for Learning]] — argues that environments should be designed for the full range of human diversity from the start. The articles in this knowledge base explore how AI can enable this through automated content transformation, adaptive assessment interfaces, and tools designed with [[neurodiversity|neurodivergent users']] lived experience as the starting point. ### Key research themes **AI-powered content accessibility** demonstrates how automated pipelines can reduce barriers. **[[adhd-video-segmentation-computing-education|Pimenova et al.]]** showed that AI-segmented [[video-education|instructional videos]] with fixed pauses eliminated the performance gap between ADHD and non-ADHD learners — strong evidence for Universal Design for Learning via automated content transformation. The study connects to [[neurodivergent-computing-students]] research on how [[collaborative-learning|collaborative learning]] structures affect neurodivergent comfort. **[[llm-question-generation-deaf-hard-of-hearing-2026|Chen et al.]]** designed an [[llm]]-powered question-generation system for Deaf and Hard of Hearing learners, introducing Visual and Emotion question strategies that target moments of visual or emotional difficulty in video — while revealing the persistent mismatch between text-based AI prompts and DHH learners' sign-based first languages, underscoring the need for language- and culture-aware AI design. **[[text-simplification-its|MuTSE]]** tackles a complementary barrier — reading level — by evaluating LLM-based text simplification for [[intelligent-tutoring]], matching content complexity to each learner's current level via a [[human-in-the-loop-ai|human-in-the-loop]] evaluation framework rather than relying on linguistic metrics that miss [[pedagogy|pedagogical]] quality. **Sensory accessibility: blind, low-vision, and Deaf learners.** Several articles invert the assumption that edtech must be visual. **[[kutti-ai-voice-first-learning-companion|Kutti AI]]** makes spoken conversation the primary and sufficient modality for visually-impaired children — real-time struggle detection, [[multilingual-learning|multilingual]] answer matching, and offline-first on-device ASR remove both the visual dependency and the connectivity requirement. **[[tactile-statistical-graphs-accessibility|Obiuwevwi et al.]]** built a reusable pipeline that generates tactile 3D-printed statistical graphs for blind/low-vision students in under 250ms, with optional LLM-based chart extraction from images. **[[pepper-robot-sign-language-lis-2025|Bolla et al.]]** explored whether the Pepper social robot can produce intelligible Italian Sign Language, co-designing 52 signs with a Deaf student and expert interpreter — extending [[educational-robotics]] into communicative accessibility for Deaf learners while highlighting the challenge of reproducing the non-manual components (facial expression, posture) crucial to meaning. **[[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]]** extend this line of work to higher education, finding in a [[qualitative-research|qualitative]] case study of 21 visually impaired undergraduates in Palestine that GenAI tailors pace, content, and delivery to individual profiles and converts complex academic texts across modalities — with learners viewing GenAI as complementing rather than replacing teachers, preserving human connection while enabling participation. **Inclusive assessment design** grapples with the tension between security and accessibility. **[[behaviorally-adaptive-visual-diversion-assessment-2026|BAVD]]** proposes a theoretical framework for adaptive visual diversion that resists screen-capture cheating while accommodating learners with visual-processing needs — explicitly modeling the trade-off between anti-cheating measures and inclusive-learning principles. This connects to broader [[academic-integrity]] and [[assessment]] concerns. **Neurodivergent learner experiences** center the voices of disabled and neurodivergent students. **[[neurodivergent-computing-students|Zastudil et al.]]** found that neurodivergent computing students need structured assignments, small consistent teams, and explicit role definitions — preferences that [[intelligent-tutoring|AI tutoring]] and collaboration tools must accommodate. **[[dyslexlens-dyslexic-learners-ai|DysLexLens]]** analyzed dyslexic learners' forum discussions, revealing that while they value AI for literacy support, they face significant accessibility barriers from inconsistent output quality and lack of equitable accommodations. Both connect to [[special-education]] and [[student-experience]]. **[[cognitive-offloading|Cognitive offloading]] and the access-vs-development trade-off.** [[seung-basham-cognitive-offloading-swld-2026|Seung & Basham (2026)]] show that for students with learning disabilities, the same GenAI that lowers barriers to reading and writing access (text leveling, summarizing, drafting support) can, if unguarded, substitute for the comprehension, planning, and monitoring practice these learners need most — an equity tension central to inclusive learning. Inclusive design must therefore consider not just whether a tool is *accessible* but whether it preserves the learner's opportunity to develop the very skills access is meant to enable. **AI for dyslexia: detection, support, and [[personalized-learning|personalized learning]].** A 2026 interdisciplinary [[meta-analysis-systematic-review|systematic review]] (Dabaghi, D'Urso & Sciarrone, PRISMA-guided, 2018–2024, n=72) finds AI supporting students with dyslexia across detection, assistive support, and personalized learning — but with these strands evolving in parallel rather than in integration, driven more by technological opportunity than by consolidated educational theory. ML-based help-education tools span five areas (specific applications, engagement, personalization, recommendation, generic support) yet emphasize technical performance and classification accuracy while overlooking ecological validity and practical classroom deployment. Detection research (EEG, eye-tracking, ML models) shows diagnostic promise for early intervention but often requires specialized equipment and controlled environments, limiting scalability and accessibility in typical school settings. Open challenges include limited experimental validation, scalability, [[ethics]]/privacy concerns with sensitive student data, limited [[teacher-role|teacher]] support and training, and language/cultural barriers (most research targets English-speaking populations) — underscoring that inclusive learning must pair technical capability with validated, scalable, and ethically grounded deployment. **Disability-centered AI critique** examines how AI systems can marginalize rather than include. **[[genai-minoritized-knowledges-disability|Tali-Otmani]]** argues that [[generative-ai|generative AI]] systems in higher education actively marginalize disability-centered ways of knowing due to Anglophone, Western-centric training data — connecting to [[equity-in-ai-education]] concerns about epistemic justice. **Accessible tools in practice** shows how AI can expand participation. **[[suacode-african-students-motivations|SuaCode]]** demonstrated that smartphone-based coding courses reach students in low-resource African contexts where fewer than 1% have coding skills. **[[embodied-string-learning-blindness-low-vision-musicians|Pimenova et al.]]** worked with blind and low-vision musicians to develop non-visual learning strategies, centering disability-led [[embodied-learning|embodied]] design. **[[ludia-udl-ai-thought-partner-2026|LUDIA]]** provides a no-cost, private, multilingual AI thought partner connecting educators with UDL principles. **[[special-r1-rl-special-education|Special-R1]]** extends [[reinforcement-learning|reinforcement learning]] to model cognitive and communicative diversity across disability profiles. ### Connections to related concepts Inclusive learning is deeply connected to [[equity-in-ai-education]] — accessibility is not merely a technical concern but a question of who gets to participate in learning. It connects to [[accessibility]] as its concrete access layer and [[assistive-technology]] as the tool layer, to [[universal-design-for-learning]] as its theoretical foundation, to [[special-education]] for disability-specific approaches, to [[learning-design]] for how courses and tools are structured, and to [[neurodiversity]] as the lens that reframes difference as diversity rather than deficit. Work on sign-language robots and tactile tools links accessibility to [[educational-robotics]] and [[educational-nlp]], while text simplification connects it to [[sociocultural-learning]] and [[adaptive-learning]]. The [[ai-education]] and [[generative-ai]] connections highlight both the promise (automated content adaptation) and peril (AI systems that reproduce exclusion). ## Implications for instructors designing inclusive learning - **Design for the excluded modality first, not last.** Building for blind/low-vision users from the start ([[kutti-ai-voice-first-learning-companion|Kutti AI]], [[tactile-statistical-graphs-accessibility|tactile graphs]]) produces tools that also work offline and in low-resource settings — accessibility as a catalyst, not a retrofit. - **Co-design with the target community.** [[pepper-robot-sign-language-lis-2025|Sign-language robots]] and [[llm-question-generation-deaf-hard-of-hearing-2026|DHH question generation]] show community involvement surfaces barriers (e.g., sign-based first languages) designers can't anticipate — involve learners and communities in design. - **Evaluate pedagogical quality, not just linguistic metrics.** [[text-simplification-its|MuTSE]] shows LLM output variability requires human-in-the-loop evaluation so that simplification helps rather than oversimplifies. - **Use AI to close performance gaps.** [[adhd-video-segmentation-computing-education|AI-segmented videos]] eliminated the ADHD performance gap — deploy adaptive AI where evidence shows it equalizes outcomes. - **Treat the security/accessibility trade-off explicitly.** [[behaviorally-adaptive-visual-diversion-assessment-2026|BAVD]] models how anti-cheating measures can inadvertently exclude learners with visual-processing needs — weigh integrity against access. - **Guard against AI reproducing exclusion.** [[genai-minoritized-knowledges-disability|Disability-centered critique]] warns that Anglophone, Western-centric training data marginalizes disabled ways of knowing — audit AI tools for epistemic justice alongside [[equity-in-ai-education]]. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[equity-in-ai-education]] - [[accessibility]] — the concrete access layer (captions, alt text, assistive-tech compatibility) - [[assistive-technology]] — the tool layer students use to access content - [[special-education]] - [[learning-design]] - [[universal-design-for-learning]] - [[neurodiversity]] - [[student-experience]] - [[ai-literacy]] - [[higher-ed]] - [[k-12]] - [[cs-education]] - [[assessment]] - [[academic-integrity]] - [[privacy]] - [[generative-ai]] - [[ai-education]] - [[educational-robotics]] - [[educational-nlp]] - [[sociocultural-learning]] - [[adaptive-learning]] - [[speech-and-voice-technologies]] ## Connected Articles - [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning]] - [[prompt-privilege-equitable-ai-access-2026]] — Prompt Privilege: measuring & mitigating accessibility disparities in LLM access - [[adhd-video-segmentation-computing-education]] - [[llm-question-generation-deaf-hard-of-hearing-2026]] — LLM-powered question generation for Deaf and Hard of Hearing learners - [[text-simplification-its]] — Text Simplification for Intelligent Tutoring - [[kutti-ai-voice-first-learning-companion]] — Kutti AI: voice-first companion for visually-impaired children - [[tactile-statistical-graphs-accessibility]] — Tactile 3D-printed statistical graphs - [[pepper-robot-sign-language-lis-2025]] — Pepper robot supporting sign language communication - [[behaviorally-adaptive-visual-diversion-assessment-2026]] - [[dyslexlens-dyslexic-learners-ai]] - [[neurodivergent-computing-students]] - [[genai-minoritized-knowledges-disability]] - [[embodied-string-learning-blindness-low-vision-musicians]] - [[suacode-african-students-motivations]] - [[ludia-udl-ai-thought-partner-2026]] - [[special-r1-rl-special-education]] - [[bilingual-llm-lecture-companion-srl-2026]] - [[multilingual-adaptive-learning-nigeria-2026]] — AI-Based Adaptive Learning Platform for Multilingual Low-Resource Contexts - [[gemini-lualatex-physics-video-transcription-2026]] — Gemini+LuaLaTeX math-accessible physics video transcription - [[khlaif-assistive-genai-visually-impaired-2026]] — Assistive GenAI for visually impaired learners - [[dabaghi-ai-dyslexia-education-review-2026]] — AI to help people with dyslexia in education - [[assistive-tech-neurodivergent-higher-ed-review-2026]] — Generative AI, virtual reality, and beyond: A scoping review of digital assistive technologies for neurodivergent students in higher education --- ## [Accessibility](https://edtechdev.github.io/aied/concepts/accessibility/) > **Accessibility** — the design of educational technology, content, and interfaces so that they can be perceived, operated, and understood by people with disabilities and diverse needs. In [[ai-education|AI in education]], accessibility covers concrete, operational barriers to the *medium* of learning: video captions, alt text, transcripts, screen-reader and keyboard compatibility, color contrast, text simplification, tactile output, sign-language support, and compatibility with assistive [[ai-technologies|technologies]]. ## Questions to Consider - When you last designed or chose a digital learning tool, did you check whether its captions, alt text, keyboard navigation, and color contrast worked before you considered its [[pedagogy]]? Why might that ordering matter? - A video with accurate captions is 'accessible,' while a course that structures discussion around a Deaf learner's communication needs is 'supporting that learner.' Where would you draw the line between removing a technical barrier and meaningfully serving a student? - [[research-methods-aied|Research]] shows AI-segmented [[video-education|instructional videos]] with fixed pauses eliminated the performance gap between [[neurodiversity|ADHD]] and non-ADHD [[learners]]. Can you recall a 'fix designed for one learner' that ended up benefiting everyone in a class you were part of? - Some argue accessibility is necessary but not sufficient — an accessible tool is not automatically an inclusive or disability-just one. What's the difference between being able to use a tool and being meaningfully served by it? - Many AI tools are trained largely on English, Western-centric data. How might that limit how well they serve learners whose first language is sign, or whose ways of knowing differ from the mainstream? - AI can automate accessibility at scale — generating captions, simplifying text, producing tactile output. What would you want to verify by hand before trusting that automated accessibility, and why? ## Introduction Accessibility is distinct from, but closely related to, three neighboring concepts in this knowledge base. **[[inclusive-learning]]** is the broader umbrella for designing education for all learner variability (physical, cognitive, sensory, situational). **[[special-education]]** is the instructional domain for learners with diagnosed disabilities, including individualized accommodations. **[[universal-design-for-learning]]** is the proactive design framework (multiple means of [[student-engagement|engagement]], representation, action/expression). **Accessibility** sits inside this constellation as the *technical and procedural layer*: removing barriers to perceiving and operating the format, rather than redesigning the pedagogy. The two can be separated on a spectrum — accessibility asks "can everyone access this content and tool?" while supporting students with disabilities asks "does instruction meaningfully serve each learner, including accommodations and disability-specific support?" Both matter, and AI intersects both. ### Why the distinction matters A video with accurate captions and a properly tagged transcript is *accessible*; a course that structures discussion to include a Deaf learner's communication preferences is *supporting that learner*. They overlap — accessible media is a prerequisite for inclusive instruction — but they require different design moves and draw on different evidence. Accessibility is anchored in standards and law (WCAG, the U.S. [[educational-policy-ai|Assistive Technology Act]] and IDEA), while accessible learning and special education are anchored in pedagogy and [[student-experience|learner experience]]. ### Key research themes **Format accessibility: captions, transcripts, and text.** **[[adhd-video-segmentation-computing-education|AI-segmented instructional videos]]** with fixed pauses eliminated the performance gap between ADHD and non-ADHD learners — accessibility as a catalyst that benefits everyone. **[[text-simplification-its|MuTSE]]** evaluates [[llm]]-based [[intelligent-tutoring|text simplification]] to match content complexity to each learner's reading level, a [[human-in-the-loop-ai|human-in-the-loop]] accessibility layer. **[[llm-question-generation-deaf-hard-of-hearing-2026|Chen et al.]]** build LLM [[automated-question-generation|question generation]] for Deaf and Hard of Hearing learners, confronting the mismatch between text-based AI prompts and sign-based first languages. **Sensory access: non-visual and tactile output.** **[[kutti-ai-voice-first-learning-companion|Kutti AI]]** makes spoken conversation the primary modality for visually impaired children, removing visual dependency. **[[tactile-statistical-graphs-accessibility|Tactile 3D-printed graphs]]** turn visual statistical data into touchable output for blind and low-vision students. **[[pepper-robot-sign-language-lis-2025|Sign-language robots]]** extend [[educational-robotics]] into communicative accessibility for Deaf learners. **[[generative-ai|Generative AI]] for visually impaired learners.** **[[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]]** — a [[qualitative-research|qualitative]] case study of 21 visually impaired [[higher-ed|undergraduates]] across three Palestinian universities — found GenAI tailors pace, content, and delivery to individual learning profiles and simplifies complex academic texts while converting content across modalities (text, audio, visual), making materials usable that were previously inaccessible. Learners framed immediacy as a foundational accessibility requirement rather than a convenience, and six interdependent technological attributes — interactivity, user-friendliness, affordability, multimodality, integration, and scalability — determined whether GenAI was genuinely accessible in a [[global-south|low-resource context]], with [[usability-research|usability]], affordability, and accessibility mutually reinforcing rather than separate design considerations. **Policy and accommodations for students with disabilities.** **[[shin-ai-policies-sld-2026|Shin et al.]]** analyze U.S. AI policy documents to reveal a void in guidance for students with specific learning disabilities, proposing accommodations and [[educational-policy-ai|policy]] recommendations grounded in the [[assistive-technology|Assistive Technology]] Act and IDEA. **[[zhang-ai-students-disabilities-meta-analysis-2024|Zhang et al.]]** meta-analyze 29 studies of AI-based interventions for students with disabilities, finding a medium positive effect on [[learning-gains|learning outcomes]] (g = 0.588) — and argue AI must do more than ensure accessibility: it must enable [[agency|agentic]] participation. **Over-inclusive AI prohibitions and assistive transcription.** **[[wright-transcription-not-generation-2026|Wright (2026)]]** argues that blanket prohibitions on "AI use" are over-inclusive because they fail to distinguish speech-to-text transcription and OCR from generative drafting: recognition technologies convert the format of content the student has already authored rather than producing new content, yet a policy written around platform identity captures both alike. Students with conditions affecting fine motor control, handwriting legibility or typing accuracy — including autism spectrum conditions, dyspraxia, cerebral palsy and repetitive strain injuries — have relied on standalone voice-to-text and OCR, and reports suggest several standalone voice-to-text products have been discontinued or degraded, leaving AI-powered transcription to fill the functional gap. Treating that substitution as misconduct raises fairness concerns under the reasonable-adjustment duty in the UK Equality Act 2010, the anticipatory Public Sector Equality Duty, the US Americans with Disabilities Act and the Australian Disability Discrimination Act 1992, though the paper does not claim the characterization has been tested in a tribunal. It notes that the intersection of disability, assistive technology and AI misconduct policy is underexplored, that the scale of this displacement is unmeasured, and that the same imprecision produces differential false-positive risk, since detectors read the low-perplexity text of non-native English writers as machine authorship. The legal exposure this over-inclusion creates is mapped on [[legal-issues-and-risks]]. **AI for dyslexia: detection, support, and [[personalized-learning|personalized learning]].** A 2026 interdisciplinary [[meta-analysis-systematic-review|systematic review]] (Dabaghi, D'Urso & Sciarrone, PRISMA-guided, 2018–2024, n=72) finds AI supporting students with dyslexia across detection, assistive support, and personalized learning — but with these strands evolving in parallel rather than in integration, driven more by technological opportunity than by consolidated educational theory. ML-based help-education tools span five areas (specific applications, engagement, personalization, recommendation, generic support) yet emphasize technical performance and classification accuracy while overlooking ecological validity and practical classroom deployment. Detection research (EEG, eye-tracking, ML models) shows diagnostic promise for early intervention but often requires specialized equipment and controlled environments, limiting scalability and accessibility in typical school settings. Open challenges include limited experimental validation, scalability, [[ethics]]/[[privacy]] concerns with sensitive student data, limited [[teacher-role|teacher]] support and training, and language/cultural barriers (most research targets English-speaking populations) — reinforcing that accessibility must be validated, scalable, and ethically grounded, not merely technically demonstrated. **The limits of accessibility alone.** **[[genai-minoritized-knowledges-disability|Critical work]]** warns that AI trained on Anglophone, Western-centric data can marginalize disability-centered ways of knowing. Accessible formats do not guarantee inclusive or just instruction — reinforcing that accessibility is necessary but not sufficient, and must connect to [[equity-in-ai-education]]. ## Implications for practice - **Prioritize the format barrier first.** Captions, transcripts, alt text, contrast, and keyboard operability are the gatekeeping layer — without them nothing else matters for learners who need them. - **Use AI to automate accessibility at scale.** AI can generate captions, simplify text, produce tactile/audio alternatives, and adapt presentation — but evaluate output quality with human-in-the-loop checks. - **Treat accessibility as necessary but not sufficient.** An accessible tool is not automatically an inclusive or disability-just tool; pair accessibility with [[inclusive-learning]] design and [[special-education]] support. - **Ground accommodations in law and policy.** Reference standards (WCAG) and statutes (Assistive Technology Act, IDEA) when designing or procuring AI tools. - **Math-accessible transcription of [[physics-education|physics]] videos (2026):** An AI workflow using Gemini (audio + 1 fps video sampling) and LuaLaTeX compiles instructional physics videos into PDF/UA-2 and ISO 32005 math-accessible PDFs that routinely pass accessibility validation — a practical, free path to making equation-heavy video content screen-readable for blind and low-vision students ([[gemini-lualatex-physics-video-transcription-2026]]). ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[inclusive-learning]] — broader umbrella for designing education across learner variability - [[special-education]] — instructional domain for learners with diagnosed disabilities - [[universal-design-for-learning]] — proactive design framework - [[equity-in-ai-education]] - [[educational-policy-ai]] - [[neurodiversity]] - [[assistive-technology]] - [[learning-design]] - [[generative-ai]] - [[educational-robotics]] - [[intelligent-tutoring]] - [[adaptive-learning]] - [[agency]] - [[virtual-and-augmented-reality]] — headsets, motion sickness and device access decide who can use it - [[speech-and-voice-technologies]] - [[legal-issues-and-risks]] — the umbrella page for over-broad rules, defective evidence and reasonable adjustment - [[arts-design-and-media-education]] ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Accessibility as a learner-centered requirement - [[seung-basham-cognitive-offloading-swld-2026]] — GenAI cognitive offloading for students with learning disabilities - [[shin-ai-policies-sld-2026]] — AI policies and accommodations for students with specific learning disabilities - [[zhang-ai-students-disabilities-meta-analysis-2024]] — Meta-analysis of AI interventions for students with disabilities - [[adhd-video-segmentation-computing-education]] — AI-segmented videos with fixed pauses - [[text-simplification-its]] — LLM-based text simplification for intelligent tutoring - [[llm-question-generation-deaf-hard-of-hearing-2026]] — LLM question generation for Deaf/Hard-of-Hearing learners - [[kutti-ai-voice-first-learning-companion]] — Voice-first companion for visually impaired children - [[tactile-statistical-graphs-accessibility]] — Tactile 3D-printed statistical graphs - [[pepper-robot-sign-language-lis-2025]] — Pepper robot supporting sign language - [[genai-minoritized-knowledges-disability]] — Critical perspective on AI and disability-centered knowledge - [[gemini-lualatex-physics-video-transcription-2026]] — Gemini+LuaLaTeX math-accessible physics video transcription - [[khlaif-assistive-genai-visually-impaired-2026]] — Assistive GenAI for visually impaired learners - [[dabaghi-ai-dyslexia-education-review-2026]] — AI to help people with dyslexia in education - [[assistive-tech-neurodivergent-higher-ed-review-2026]] — Generative AI, virtual reality, and beyond: A scoping review of digital assistive technologies for neurodivergent students in higher education - [[wright-transcription-not-generation-2026]] — Transcription is not generation: over-inclusive AI prohibitions and the assistive tools they capture --- ## [Assistive Technology](https://edtechdev.github.io/aied/concepts/assistive-technology/) > **Assistive Technology** — devices, software, and services that help people with disabilities perceive, operate, communicate, and participate in learning and daily life. In [[ai-education|AI in education]], assistive technology spans screen readers, speech-to-text and text-to-speech, captioning, braille and tactile output, sign-language tools, and increasingly AI-powered accommodations that adapt content and interaction to individual needs. ## Questions to Consider - Assistive technology is the tool layer of accessibility — the specific devices and software individuals use to bridge access gaps. Before reading, did you think of accessibility and assistive technology as the same thing? How might treating them as distinct change how you design learning environments? - [[research-methods-aied|Research]] finds AI-based interventions yield a medium positive effect on [[learning-gains|learning outcomes]] for students with disabilities (g = 0.588). Yet the page warns that access tools don't by themselves ensure inclusive instruction or learner agency. What's the difference between giving a student access and genuinely including them? - The page notes that US AI policy documents largely fail to address assistive technology and accommodations for students with specific learning disabilities. Why do you think accommodations so easily fall through the cracks of AI policy — and who loses when that happens? - [[generative-ai|Generative AI]] can auto-caption, simplify text, and generate tactile alternatives — lowering the cost of adaptation. But the page cautions you to evaluate the quality of AI-generated alternatives for [[pedagogy|pedagogical]] accuracy. What could go wrong if a 'simplified' or 'tactile' version misrepresents the content it's meant to make accessible? - AI is expanding assistive tools from voice-first learning companions to tactile graphs and sign-language tools. As an educator or designer, which learner's specific barrier would you want to address first with AI — and what would you need to know about that learner before choosing a tool? ## Introduction Assistive technology is the concrete *tool layer* of [[accessibility]]. Where accessibility is the design property of an environment (can everyone access it?), assistive technology is the specific equipment and software that individuals use to bridge access gaps. It is foundational to [[special-education]] and [[inclusive-learning]] — students with specific learning disabilities, visual or hearing impairments, and motor challenges rely on assistive tools to access the [[curriculum-design|curriculum]]. In the U.S., the [[educational-policy-ai|Assistive Technology Act (2004)]] and the Individuals with Disabilities Education Improvement Act (IDEA, 2004) provide the legal basis for providing these tools to students with disabilities. ### Key research themes **AI is expanding assistive technology.** Generative AI and LLMs are transforming assistive tools — [[text-simplification-its|LLM-based text simplification]] adapts reading level in [[intelligent-tutoring]], [[kutti-ai-voice-first-learning-companion|voice-first AI]] removes visual dependency for blind and low-vision learners, and [[tactile-statistical-graphs-accessibility|AI-generated tactile graphs]] convert visual data to touchable output. **[[zhang-ai-students-disabilities-meta-analysis-2024|Zhang et al.]]** find AI-based interventions (robots, software, intelligent VR) yield a medium positive effect on the learning outcomes of students with disabilities (g = 0.588). **[[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]]** add a [[qualitative-research|qualitative]] case study of 21 visually impaired undergraduates in Palestine, showing GenAI functions as an assistive layer that tailors pace, content, and delivery, simplifies complex texts, and converts content across modalities — with learners consistently viewing it as complementing rather than replacing teachers. **Policy and provision.** **[[shin-ai-policies-sld-2026|Shin et al.]]** document that U.S. AI policy documents largely fail to address assistive technology and accommodations for students with specific learning disabilities, calling for policy guidance grounded in the Assistive Technology Act and IDEA. **AI for dyslexia across detection, support, and [[personalized-learning|personalized learning]].** A 2026 interdisciplinary [[meta-analysis-systematic-review|systematic review]] (Dabaghi, D'Urso & Sciarrone, PRISMA-guided, 2018–2024, n=72) maps AI support for students with dyslexia, finding AI used for detection, assistive support, and personalized learning — but with these strands evolving in parallel rather than in integration, driven more by technological opportunity than by consolidated educational theory. ML-based help-education tools fall into five areas (specific applications, [[student-engagement|engagement]], personalization, recommendation, generic support) yet emphasize technical performance and classification accuracy while overlooking ecological validity and practical classroom deployment. Detection research (EEG, eye-tracking, ML models) shows diagnostic promise for early intervention but often requires specialized equipment and controlled environments, limiting scalability and accessibility in typical school settings. Open challenges include limited experimental validation, scalability, [[ethics]]/privacy concerns with sensitive student data, limited [[teacher-role|teacher]] support and training, and language/cultural barriers (most research targets English-speaking populations) — a reminder that assistive tools must be validated, scalable, and ethically grounded to genuinely bridge access gaps. **The limits of assistive tools.** Assistive technology enables access but does not by itself ensure inclusive instruction or learner [[agency]]. **[[genai-minoritized-knowledges-disability|Critical research]]** and the push for [[agency|agentic]] roles for students with disabilities remind us that access must pair with meaningful participation. A 2026 scoping review of digital assistive [[ai-technologies|technologies]] for [[neurodiversity|neurodivergent]] students in [[higher-ed|higher education]] maps the decade's output: 766 records screened across five databases, 40 studies included, with AI-based tools in 15 of them and [[virtual-and-augmented-reality|virtual reality]] in 11. Its organizing finding is a mismatch between what the tools target and where the barriers sit — 27 studies supported learning directly, 13 addressed reading and writing and 12 study management, while attention (n = 4) and social communication (n = 5) were comparatively neglected and only 6 addressed multiple barriers ([[assistive-tech-neurodivergent-higher-ed-review-2026|Rempel et al., 2026]]). ## Implications for practice - **Match the tool to the learner and task.** Screen readers, captions, speech, and tactile output each address different barriers — choose based on the individual's needs and the content format. - **Leverage AI to lower the cost of assistive adaptations.** AI can auto-caption, simplify text, and generate alternatives, but evaluate output quality for pedagogical accuracy. - **Ground provision in policy.** Reference the Assistive Technology Act, IDEA, and WCAG when procuring or building AI tools. ## Connected Concepts - [[accessibility]] — the design property that assistive technology operationalizes - [[inclusive-learning]] - [[special-education]] - [[universal-design-for-learning]] - [[equity-in-ai-education]] - [[educational-policy-ai]] - [[neurodiversity]] - [[learning-design]] - [[speech-and-voice-technologies]] ## Connected Articles - [[shin-ai-policies-sld-2026]] — AI policies and accommodations for students with specific learning disabilities - [[zhang-ai-students-disabilities-meta-analysis-2024]] — Meta-analysis of AI interventions for students with disabilities - [[kutti-ai-voice-first-learning-companion]] — Voice-first AI for visually impaired children - [[tactile-statistical-graphs-accessibility]] — AI-generated tactile statistical graphs - [[text-simplification-its]] — LLM-based text simplification for intelligent tutoring - [[llm-question-generation-deaf-hard-of-hearing-2026]] — LLM question generation for Deaf/Hard-of-Hearing learners - [[gemini-lualatex-physics-video-transcription-2026]] — Gemini+LuaLaTeX math-accessible physics video transcription - [[khlaif-assistive-genai-visually-impaired-2026]] — Assistive GenAI for visually impaired learners - [[dabaghi-ai-dyslexia-education-review-2026]] — AI to help people with dyslexia in education - [[assistive-tech-neurodivergent-higher-ed-review-2026]] — Generative AI, virtual reality, and beyond: A scoping review of digital assistive technologies for neurodivergent students in higher education - [[adapted-stories-social-story-intervention-2026]] — AI-Assisted Social Story Intervention for Special Education: The Design of AdaptED Stories --- ## [Neurodiversity](https://edtechdev.github.io/aied/concepts/neurodiversity/) > **Neurodiversity** — the framing that neurological differences such as autism, ADHD, dyslexia, and dyspraxia are natural variations in [[cognitive-psychology|human cognition]] rather than deficits to be corrected. In education, a neurodiversity-affirming approach designs learning environments that accommodate and leverage these differences rather than forcing conformity to a single cognitive norm. ## Questions to Consider - The page frames neurodiversity as neurological differences being natural variations rather than deficits to fix. How does that shift in framing change what 'helping' a learner with autism, ADHD, or dyslexia should look like in your own setting? - Consider a generative AI tool that produces long text-heavy responses. Which neurodivergent learners might it help, and which might it disadvantage—and can you think of design choices that would tip the balance either way? - AI that reduces [[cognitive-offloading|cognitive load]] can support learners who struggle with executive function or social expectations, yet AI that encourages dependency may undermine them. Where do you draw the line between support and over-reliance for a specific learner? - The page warns that AI tools which assume a dominant communication style can recapitulate equity gaps. Can you recall a time a 'one-size-fits-all' tool or classroom assumed everyone learned the same way, and what it overlooked? - How might a learner's neurotype change how their behavior signals are interpreted in learning analytics? What risk arises when AI reads [[student-engagement|engagement]] or struggle without knowing how a particular brain works? - The strongest evidence on this page is for a general design change that helped everyone and closed a gap, not for a tool built for one diagnosis. Why might universal changes outperform diagnosis-gated ones — and what does that imply about how you would spend a limited accessibility budget? ## Introduction The neurodiversity paradigm shifts the goal of [[special-education]] and [[accessibility]] work from "fix the learner" to "adapt the environment." It overlaps with [[universal-design-for-learning]] and [[inclusive-learning]] but emphasizes affirming identity and strength-based design over accommodation-as-compensation. In the AI era this framing becomes a design test rather than a slogan. [[generative-ai|Generative AI]] offers real promise for neurodivergent learners — alternative means of engagement, representation and expression, support for executive function, and reduced extraneous load — and real risk, since tools that assume a dominant communication style, gate support behind a formal diagnosis, or encourage dependency can disadvantage exactly the learners they claim to serve. This page maps what the evidence actually shows, category by category, and where it runs out. ## What the research covers Two reviews define the shape of the field, and both describe fragmentation rather than consolidation. [[assistive-tech-neurodivergent-higher-ed-review-2026|A PRISMA-ScR scoping review (Rempel et al., 2026)]] searched five databases and a decade of publication (2015–2025), screening 766 records down to 40 empirical studies of digital assistive technologies for neurodivergent students in [[higher-ed|higher education]]. The field has reorganized around [[generative-ai|generative AI]] — 15 of the 40 studies — with immersive formats second at 11. Its central critique is a design critique: individual accommodation and neurotype-specific tools organized around formal diagnosis cannot address functional differences that cut across neurotypes, so it recommends universal design and participatory development instead. It also notes that the most immersive tools are the least scalable. [[dabaghi-ai-dyslexia-education-review-2026|The interdisciplinary review of AI for dyslexia (Dabaghi, D'Urso and Sciarrone, 2026)]] covers 2018–2024 across 72 studies and finds AI used for detection, assistive support, and [[personalized-learning|personalization]], with those strands evolving in parallel rather than in integration, and the field driven more by technological opportunity than by consolidated educational theory. Its detection research (EEG, eye-tracking, [[machine-learning]] models) shows diagnostic promise for early intervention but often needs specialized equipment and controlled conditions, which limits use in ordinary classrooms. That is the practical test this whole strand keeps failing: support has to be embedded in the course a student is actually taking. ## Autism and ADHD in the classroom [[neurodivergent-computing-students|A survey of 24 neurodivergent computing students (autistic and/or ADHD) and 20 neurotypical peers]], with four in-depth interviews, found significant discomfort with assignments that lack clear structure or carry ambiguous expectations — and the same structures that suit neurotypical learners in collaborative [[active-learning|active learning]] can exclude neurodivergent peers. The authors present it as preliminary and among the first studies to center neurodivergent voices in [[cs-education|computing education]], which is also a measure of how little evidence exists. [[adhd-video-segmentation-computing-education|The video-segmentation study (Pimenova, Begel and colleagues, 2026)]] is the strongest single result on this page. Treating [[video-education|instructional videos]] as a post-hoc processing problem — segmenting them into single-instruction chunks with fixed pauses to reduce extraneous load — improved performance for everyone in a within-participants design (17 ADHD, 10 non-ADHD), and the ADHD participants' errors and hesitations fell to parity with their non-ADHD peers. An equalizing change that needs no diagnosis, no disclosure, and no separate tool is the cleanest available demonstration of [[universal-design-for-learning|Universal Design for Learning]] through automated content transformation. ## Specific learning disabilities, dyslexia, and cognitive offloading [[zhang-ai-students-disabilities-meta-analysis-2024|Zhang et al. (2024)]] pooled 29 (quasi-)experimental studies of AI-based interventions for students with disabilities and found a medium positive overall effect (Hedge's g = 0.588, 95% CI [0.349, 0.826]) across 239 effect sizes from 41 independent samples. Two details matter more than the headline. The first is that the effect varies by disability category: students with specific learning disabilities, intellectual and developmental disabilities, or who are deaf showed a larger effect (g = 0.952) than students with autism spectrum disorder (g = 0.368), though the authors report the difference as not statistically significant. The second is publication bias: Egger's test was significant (β = 2.837, p < .001) and trim-and-fill reduced the pooled estimate to g = 0.2694, still positive. "Neurodivergent students" is not one population, and a pooled effect across categories is not a promise to any of them. The risk side has its own literature. [[seung-basham-cognitive-offloading-swld-2026|Seung and Basham (2026)]] argue that generative AI has changed cognitive offloading from a peripheral study aid into a delegation of higher-order cognitive processes, with especially consequential implications for students with learning disabilities — the learners most likely to be offered the shortcut and least likely to be served by it if the delegated process was the point of the task. ## Executive function and reliance [[genai-reliance-executive-functioning-2026|Klarin, Hoff and Daukantaitė (2026)]] define reliance narrowly — preferring generative AI over one's own effort or teacher support, plus difficulty initiating schoolwork without it — and test it in two Swedish community samples of adolescents (849 lower-secondary students, analytic n = 735, and 898 upper-secondary students, analytic n = 839). Executive-functioning difficulties ran through perceived usefulness and habitual use to reliance in both samples, with small indirect effects (β = .10, 95% CI [.06, .14] in the lower-secondary sample; β = .08, 95% CI [.04, .12] in the upper-secondary sample). The practical reading: the students who most need support with initiation are the ones for whom a convenient tool most easily becomes the default route rather than one option — and the effect sizes here are small enough that this is a design concern, not a diagnosis. ## Design cautions - **Diagnosis-gated support excludes the learners it does not name.** The scoping review's critique applies directly to how AI features are deployed: if a feature is unlocked by an accommodation letter, learners with functional differences but no diagnosis never see it. - **Tools can exclude epistemically, not just poorly.** [[genai-minoritized-knowledges-disability|Tali-Otmani (2026)]] argues that Anglophone, Western-centric training data marginalizes non-hegemonic ways of knowing, and puts the situation of disabled learners at the center of that critique — a warning that "inclusive" AI can still encode whose knowledge counts. - **Behavior signals are read by systems that do not know the learner.** Interpretation of [[student-engagement|engagement]], struggle or attention in [[learning-analytics]] and [[student-modeling]] assumes a normative pattern of response; see [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]] for the fairness evidence on that assumption. - **Dependency is a design outcome, not a learner failing.** Given the reliance pathway above, the question to ask of any AI feature is whether it substitutes for the higher-order process the assignment exists to build — the point Seung and Basham make for learning disabilities. ## Where the evidence is thin - **Almost everything here is single-group.** The disability meta-analysis has no neurotypical comparator, so it establishes that interventions helped, not that they helped this group differently. - **Samples are small.** Twenty-four students and four interviews in the computing study; 17 and 10 in the ADHD video study; 72 studies in the dyslexia review with the strands still uncoupled. - **Category boundaries differ across studies,** so cross-study comparison of "the same" neurotype is limited — and the between-category differences in the meta-analysis are reported as non-significant. - **The field's own reviewers describe it as technology-driven,** which means the intervention evidence is thin exactly where the design guidance is strongest. ## Connections Neurodiversity connects to [[special-education]], [[inclusive-learning]], [[universal-design-for-learning]], and [[equity-in-ai-education]]. It informs both how AI is deployed for [[student-experience]] and how assessments and literacy programs are designed to be fair across cognitive variability. For how these findings sit alongside other learner groups — language, gender, socioeconomic status, geography — see [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]]. ## Connected Concepts - [[special-education]] - [[inclusive-learning]] - [[universal-design-for-learning]] - [[equity-in-ai-education]] - [[differential-effects-across-learner-groups]] - [[student-experience]] - [[personalized-learning]] - [[learning-analytics]] - [[generative-ai]] - [[cognitive-offloading]] - [[student-modeling]] ## Connected Articles - [[assistive-tech-neurodivergent-higher-ed-review-2026]] — Generative AI, virtual reality, and beyond: a scoping review of digital assistive technologies for neurodivergent students in higher education - [[zhang-ai-students-disabilities-meta-analysis-2024]] — 29 studies of AI for students with disabilities, and an effect that differs by category - [[adhd-video-segmentation-computing-education]] — Temporal video segmentation that brought ADHD participants to parity - [[neurodivergent-computing-students]] — 24 neurodivergent computing students on structure, ambiguity, and collaboration - [[genai-reliance-executive-functioning-2026]] — Executive-functioning difficulties, perceived usefulness, and the route to AI reliance - [[seung-basham-cognitive-offloading-swld-2026]] — GenAI cognitive offloading for students with learning disabilities - [[dabaghi-ai-dyslexia-education-review-2026]] — AI to help people with dyslexia in education - [[genai-minoritized-knowledges-disability]] — Generative AI and the marginalization of minoritized knowledges - [[tactile-statistical-graphs-accessibility]] — Tactile Statistical Graphs for Accessibility - [[ai-learning-tools-engineering-education-needs]] — Designing Needs- and Attention-Aware AI Learning Tools - [[adapted-stories-social-story-intervention-2026]] — AI-Assisted Social Story Intervention for Special Education: The Design of AdaptED Stories --- ## [Universal Design for Learning](https://edtechdev.github.io/aied/concepts/universal-design-for-learning/) > **Universal Design for Learning (UDL)** — an educational framework that designs instruction to be accessible and effective for the widest range of learners by proactively building in flexible means of [[student-engagement|engagement]], representation, and action/expression, rather than retrofitting accommodations for individuals. ## Questions to Consider - UDL's core claim is that learner variability is the norm, not the exception — so design for many pathways from the start rather than fixing a single path for those who struggle. Do you tend to design for an 'average' learner and add accommodations later? What might be lost in that default approach? - Think of a time a particular format (a dense text, a lecture, a single way of showing work) excluded you or someone you taught. Which of the three UDL principles — engagement, representation, or action/expression — was at stake, and what would a flexible alternative have looked like? - Generative AI can create personalized examples, captions, summaries, and multiple assessment formats — seemingly a gift for UDL. But the page warns AI can also encode bias and reduce [[agency|learner agency]]. Before reading on, how could a tool built to personalize end up *narrowing* rather than widening learning paths for some students? - A common [[misconceptions|misconception]] is that UDL means lowering standards or giving everyone different outcomes. Consider an assignment that lets students submit an essay, a video, a diagram, or code for the same learning outcome. Is that a watering-down, or a fairer way to assess the same ability? Defend your view. - UDL turns 'fix the learner' into 'fix the design.' Pick a frustrating learning experience you've had or designed. If the barrier were a design problem rather than a student problem, what would you change about the design to remove it for everyone? - The page warns against AI that assumes one communication style or penalizes neurodivergent expression. If you use or design AI-assisted tools, where might 'default' styles quietly exclude learners — and what would it take to audit for that? ## Introduction UDL rests on the insight that learner variability is the norm, not the exception. Rather than designing a single path and adding support for those who struggle, UDL designs multiple pathways from the start so that barriers are removed for everyone. It is a core lens for [[inclusive-learning]], [[equity-in-ai-education]], and [[special-education]]. ### The three principles - **Multiple means of engagement** — the "why" of learning: varied ways to motivate and sustain interest, connect to relevance, and support [[self-regulated-learning|self-regulation]]. - **Multiple means of representation** — the "what" of learning: presenting information in varied formats (text, audio, visual, interactive) so all learners can perceive and comprehend it. - **Multiple means of action and expression** — the "how" of learning: offering varied ways for learners to demonstrate what they know (writing, speaking, building, performing). ### UDL in the AI era [[generative-ai|Generative AI]] creates new opportunities and new risks for UDL. AI can personalize representation and provide alternative pathways, supporting [[personalized-learning]] and accessibility. But it can also encode bias, assume dominant communication styles, and — if it reduces learner agency — undermine the engagement principle. [[research-methods-aied|Research]] on [[ai-misuse-learning-harm]] and equity shows that AI tools must be designed with inclusive principles or they recapitulate [[equity-in-ai-education]] gaps. UDL therefore informs both how AI is deployed and how AI-literacy and assessment are designed to be fair across learner variability. Equity-by-design is one concrete way to operationalize this stance. Wenzel, Geiger, and Liening (2026) ground their CAIS-GBL framework for AI [[conversational-ai|conversational agents]] in business [[simulation]] games in an explicit **equity-by-design** approach aligned with universal design for learning, deriving meta-requirements that span cognitive, [[motivation|motivational]], [[affective-computing|affective]], and [[sociocultural-learning|socio-cultural]] engagement. Their design principles and features aim to address individual learner differences in strengths, challenges, and interests — a concrete application of UDL principles to the design of adaptive, AI-supported [[game-based-learning|game-based learning]]. AI also appears in this literature as an obstacle to UDL rather than an instrument of it. In [[faculty-accessible-course-design-ai-2026|Sidhu, Atif and Newland's (2026)]] mixed-methods study of 69 professors at one Ontario university, one of the four key themes is that [[generative-ai|generative AI]] is reversing progress in accessible course design; the title quotation, "AI is reducing options I had used", comes from their interviews, because the take-home and drafting-based assignments that had done inclusion work are the ones an AI can now complete outright. Their faculty-level findings show why that setback lands on an already thin base: 90% of respondents believed they create an accessible learning environment and 87% said they actively consider accessibility, yet fewer than half (46%) thought the university offered sufficient resources for it, and roughly a fifth had never heard of techniques as basic as OCR-readable PDFs or color-independent links. Demand is outrunning provision — Ontario university enrollments grew 17% from 2013 to 2022 while registrations with accessibility services rose 126% — so the case the study makes is that UDL support has to be resourced to compensate for the assessment options AI removed. ### Connections UDL connects to [[inclusive-learning]], [[equity-in-ai-education]], [[special-education]], [[learning-design]], and [[culturally-relevant-pedagogy]]. In assessment, it intersects with [[authentic-assessment]]'s emphasis on representational [[bias-mitigation|fairness]] and with [[reducing-ai-misuse]] as a guardrail against tools that penalize particular communication styles. Because UDL is the framework most commonly invoked in higher-education disability contexts (where "special education" is a [[k-12]] term), it is often the right page to link for college and university disabled-learner research. ## Implications and examples for instructors and instructional designers UDL turns "fix the learner" into "fix the design." For instructors and designers working with AI, the three principles translate into concrete moves: **Engagement — offer multiple ways to spark and sustain motivation.** - Let learners choose how they engage: problem-based, game-based, discussion, or self-paced options. - Use AI to surface relevance — personalized examples, real-world connections, or choice of topic — rather than a single generic task. - *Example:* A course uses AI to generate varied worked examples tied to different learner interests (business, health, arts), letting students pick the context that motivates them. **Representation — present information in multiple formats.** - Offer the same content as text, audio, video, and interactive — AI can auto-generate captions, transcripts, summaries, and alternative explanations at different reading levels. - *Example:* An instructor uses an AI assistant to produce a plain-language summary and an audio version of a dense reading, so learners can choose their entry point. Pair with [[accessibility]] (captions, alt text) so every format is usable. **Action and expression — let learners show what they know in varied ways.** - Provide choice of assessment product (essay, presentation, video, diagram, code) aligned to the same learning outcome. - *Example:* A project allows submission as a written report, an AI-assisted video explainer, or a live demonstration — with AI [[scaffolding|scaffolds]] supporting each mode. In [[assessment]] design, this parallels [[authentic-assessment]]'s representational fairness. **Design for AI-literacy and agency.** - Teach students how and when to use AI, and build checkpoints that keep the learner (not the tool) accountable — see [[ai-literacy]] and [[reducing-ai-misuse]]. - *Example:* A UDL-aligned assignment lets students use AI to draft but requires a [[metacognition|metacognitive]] reflection on their own contribution, preserving the engagement and agency principles. **Use AI to remove barriers, not add them.** - Deploy AI to close performance gaps (e.g., AI-segmented videos with pauses helped ADHD learners) and to lower the cost of accessible formats. - Guard against AI that assumes one communication style or penalizes neurodivergent expression — connect to [[accessibility]], [[equity-in-ai-education]], and [[neurodiversity]]. ## Connected Concepts - [[inclusive-learning]] - [[accessibility]] - [[assistive-technology]] - [[equity-in-ai-education]] - [[special-education]] - [[neurodiversity]] - [[learning-design]] - [[personalized-learning]] - [[culturally-relevant-pedagogy]] - [[authentic-assessment]] - [[student-experience]] - [[generative-ai]] - [[pedagogy]] — Umbrella: pedagogies and teaching strategies in AI education - [[bias-mitigation]] - [[reducing-ai-misuse]] - [[k-12]] ## Connected Articles - [[seung-basham-cognitive-offloading-swld-2026]] — GenAI cognitive offloading for students with learning disabilities - [[authentic-products-authenticated-processes-2026]] — From Authentic Products to Authenticated Processes - [[tactile-statistical-graphs-accessibility]] — Tactile Statistical Graphs for Accessibility - [[neurodivergent-computing-students]] — Neurodivergent Computing Students - [[ai-learning-tools-engineering-education-needs]] — Designing Needs- and Attention-Aware AI Learning Tools - [[gemini-lualatex-physics-video-transcription-2026]] — Gemini+LuaLaTeX math-accessible physics video transcription - [[conversational-agents-business-simulation-gaming-2026]] — CAIS-GBL framework for AI conversational agents in business simulation games (Wenzel et al. 2026) - [[assistive-tech-neurodivergent-higher-ed-review-2026]] — Generative AI, virtual reality, and beyond: A scoping review of digital assistive technologies for neurodivergent students in higher education - [[faculty-accessible-course-design-ai-2026]] — “AI is reducing options I had used”: exploring faculty perceptions of accessible course design in a ChatGPT world --- ## [Global South](https://edtechdev.github.io/aied/concepts/global-south/) The **Global South** refers to countries in Africa, Asia, Latin America, and Oceania that are often economically, politically, and historically marginalized relative to the Global North. In [[ai-education|AI in education]] [[research-methods-aied|research]], Global South contexts are increasingly recognized as underrepresented in the evidence base, yet they raise distinctive questions about [[equity-in-ai-education|equity]], [[culturally-relevant-pedagogy|cultural relevance]], resource constraints, and the epistemic dominance of Western, Anglophone training data. ## Questions to Consider - If an AI system performs well on benchmark tests built from Western, English-language data, how confident should you be that it will work equally well for your students? What would you want to check before trusting it in your own context? - Consider a learning tool whose training data and evaluation metrics were all developed in the Global North. What assumptions about language, knowledge traditions, and educational realities might it silently encode — and who is most likely to be misrepresented by them? - When AI adoption is studied, whose classrooms and institutions tend to dominate the evidence base? How might that skew what we think we know about whether AI 'works' in education? - The page argues for treating learners' lived and community epistemologies as authoritative, not just as add-ons to Western models. What would it take for your own AI-related practice or research to center local knowledge rather than import it from elsewhere? - Resource constraints are a recurring theme in Global South contexts. How might the promise of low-cost, scalable AI support both close and widen existing opportunity gaps, depending on how it is designed and deployed? - How could you tell whether a technology 'adopted' in a particular setting was actually adopted because it fit local conditions — or because it was assumed to transfer? What evidence would distinguish the two? ## Introduction ### Why It Matters for AIED Mainstream AI and educational-technology research has historically been dominated by Western, English-language datasets and [[governance|institutional]] contexts. This creates two problems: (1) [[ai-technologies|AI systems]] trained on such data may underperform or misrepresent learners in Global South settings, and (2) evaluation [[benchmark|benchmarks]] built in the Global North may not reflect the educational realities, languages, or knowledge traditions of other regions. Research from Global South contexts in this knowledge base spans culturally grounded datasets, benchmarks, and technology-adoption studies, with implications for [[ai-literacy]] and [[higher-ed|higher]] and [[k-12|K-12]] education. ### Applications in the Knowledge Base - **Culturally grounded data and benchmarks:** [[iks-instruct-dataset-indian-knowledge|IKS-Instruct]] provides a [[multilingual-learning|multilingual]] Indian Knowledge Systems instruction dataset; [[nsmq-riddles-science-math-benchmark|NSMQ Riddles]] introduces a Ghana-based [[stem-education|STEM]] benchmark, one of the first Global South educational evaluation datasets. - **AI ethics awareness in Ghanaian higher education:** [[ai-ethical-awareness-ghana-students-2026|Acquah et al. (2026)]] validated a three-factor (autonomy, beneficence, fairness) AI ethical awareness scale with 509 [[higher-ed|university]] students in Ghana, finding five latent profiles from uniformly very high awareness to a small cluster with almost none — weakest on beneficence — and no gender differences, providing a Global South measurement anchor for the AI-[[ethics]] literature. - **Teacher ethical reasoning in low-resource contexts:** [[teachers-contextual-ethical-reasoning-ai-2026|Adelana et al. (2026)]] found in-service secondary STEAM teachers in a low-resource context reason about AI ethics through professional identity, classroom realities, and cultural norms rather than formal principle lists, countering "ethics empty" assumptions about the Global South. - **Contextual adoption:** [[socio-cognitive-genai-adoption-engineering-2026|Asag & Al Mamun]] model [[generative-ai|GenAI]] adoption among Bangladeshi [[engineering-education|engineering]] students, and [[connected-ai-lesson-planning-vietnam|ConnectED]] deploys a [[curriculum-design|curriculum]]-aligned lesson-planning system for Vietnamese education. - **Institutional integration and structural inequality:** [[adeniranye-ai-integration-nigerian-higher-education-2026|Adeniranye et al. (2026)]]'s comparative content analysis of 45 Nigerian universities (federal/state/private) found that AI integration is predicted by institution age and geographic location — not governance type — and that international and industry network ties reinforce one another (r = 0.74), so well-connected institutions compound advantage while others fall further behind. It locates [[digital-divide|digital inequality]] not just at the learner level but in the structural capacity of [[higher-ed|higher-education]] institutions themselves. - **[[educational-development|Academic development]] as digital mediation under structural constraint:** [[beyond-the-algorithm-academic-developers-digital-mediators-2026|Sithole (2026)]] interviews twelve academic developers and learning designers across two South African Historically Disadvantaged Institutions and draws an analytical line the Global South literature often blurs: **digital inequality** is distributive and answerable in principle through redistribution and access, while **algorithmic coloniality** is epistemic and persists even under full access because it inheres in what the systems encode and whose knowledge they center. The two are entangled but demand different responses, and the institution is doubly positioned — excluded by infrastructural scarcity, and, when included, subordinated by imported systems that encode other knowledges and norms. Institutional AI rhetoric ("performance more than practice") arrived loosely coupled to the material realities of teaching, so developers translated global digital-transformation discourse into contextually viable practice: testing and deliberately "breaking" tools in peer spaces, insisting on judgment over skill, and absorbing the affective cost of projecting expertise they were still acquiring. - **Epistemic marginalization:** [[genai-minoritized-knowledges-disability|Tali-Otmani]] argues that Western-centric training data marginalizes non-Western and disability-centered knowledges — connecting Global South concerns to [[equity-in-ai-education]] and [[culturally-relevant-pedagogy]]. - **Disability and [[inclusive-learning|inclusion]] in the Global South:** [[khlaif-assistive-genai-visually-impaired-2026|Khlaif et al. (2026)]] — a [[qualitative-research|qualitative]] case study of 21 visually impaired undergraduates across three Palestinian universities — found GenAI bridges digital, geographic, and socioeconomic divides for disabled learners, extending [[technology-acceptance-model|technology-acceptance]] research to disability contexts where [[usability-research|usability]], affordability, and [[accessibility]] are mutually reinforcing. - **Designing for local stressors under resource constraints:** Bashir and Afzal (2026) build [[culturally-aware-student-stress-chatbot-2026|Sukoon]] as a Global South design response to a documented mismatch — mental-health [[conversational-ai|chatbots]] trained mostly on Western datasets and overwhelmingly English-language, while Pakistani students face academic, financial, familial, and relational stressors simultaneously and often cannot raise emotional difficulties with [[parents-and-families|parents]], [[teacher-role|teachers]], or peers because of stigma. The authors' practical constraints are as instructive as their model: a free-access [[open-source]] [[llm]] through a hosted API for low resource requirements, a lightweight Flask deployment for regional universities, "tools are available but often expensive" listed as a barrier, and unequal access to paid models flagged as a general dependency risk. They also note the classifier was trained on a publicly available dataset not representative of Pakistani students, and commit to locally collected DASS-21 data before drawing population conclusions. ### Implications Attending to Global South contexts requires moving beyond assuming Western models and benchmarks transfer directly. It calls for locally grounded datasets, culturally relevant [[pedagogy]], community-centered evaluation standards, and research that treats learners' lived and community epistemologies as authoritative — aligning with frameworks like community-based AI learning and [[technology-acceptance-model|technology-acceptance]] research adapted to local conditions. The scaling record is part of that picture, and it is sobering. Programs that work at pilot scale frequently stop working when governments run them: the Kenyan program that produced substantial gains under non-governmental implementation showed no detectable gain at government scale (Bold et al.), effect sizes tend to fall as programs grow (Vivalt), and implementation quality dilutes as sites multiply (Al-Ubaydli, List & Suskind). Latin America's own precedent is the Peruvian One Laptop per Child program — about 800,000 laptops distributed with no detectable effect on [[math-education|mathematics]] or reading — from which the lesson drawn was that access to technology is not instruction. [[el-salvador-ai-tutoring-selection-claim-2026|Restrepo Morales et al. (2026)]] add a reporting lesson to the scaling lesson: in the 2026 El Salvador episode an AI-tutoring pilot in 171 schools was announced as comparable to Germany and Sweden while the same country's representative PISA 2025 sample showed no movement in two of three subjects, and the pilot data were too thin (7.0 assessed students per school against 25.4 nationally) to exclude selection as the explanation. For Global South systems adopting AI at scale the implication is double: the evidence that a reform works is usually about the components it bundles rather than the technology in its headline, and a phased rollout designed before deployment identifies effects at almost no cost — which matters most where connectivity dictates the phasing anyway. ## Connected Concepts - [[differential-effects-across-learner-groups]] - [[equity-in-ai-education]] - [[generative-ai]] - [[culturally-relevant-pedagogy]] - [[ai-literacy]] - [[higher-ed]] - [[k-12]] - [[inclusive-learning]] - [[technology-acceptance-model]] - [[benchmark]] ## Connected Articles - [[adeniranye-ai-integration-nigerian-higher-education-2026]] — Institutional structures, digital inequality, and AI integration in Nigerian higher education - [[nguyen-genai-global-south-review-2026]] - [[socio-cognitive-genai-adoption-engineering-2026]] — Unified socio-cognitive model for engineering education (Bangladesh) - [[connected-ai-lesson-planning-vietnam]] — ConnectED: Vietnamese Lesson Planning - [[iks-instruct-dataset-indian-knowledge]] — IKS-Instruct: Indian Knowledge Systems Dataset - [[nsmq-riddles-science-math-benchmark]] — NSMQ Riddles: Ghana STEM Benchmark - [[genai-minoritized-knowledges-disability]] — Marginalization of minoritized knowledges - [[llm-cultural-relevance-k12]] — LLMs for Culturally Relevant K-12 Pedagogy - [[multilingual-adaptive-learning-nigeria-2026]] — AI-Based Adaptive Learning Platform for Multilingual Low-Resource Contexts - [[genai-integration-constructivist-higher-ed-bangladesh-2026]] — GenAI integration in Bangladeshi higher ed through constructivism (Alam et al. 2026) - [[khlaif-assistive-genai-visually-impaired-2026]] — Assistive GenAI for visually impaired learners - [[culturally-aware-student-stress-chatbot-2026]] — An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning - [[el-salvador-ai-tutoring-selection-claim-2026]] — Bounding the learning claim of El Salvador's AI tutoring pilot (Restrepo Morales et al. 2026) - [[beyond-the-algorithm-academic-developers-digital-mediators-2026]] — Academic developers as digital mediators in South African HDIs: digital inequality vs. algorithmic coloniality - [[ai-ethical-awareness-ghana-students-2026]] — Artificial intelligence ethical awareness of Ghanaian university students - [[teachers-contextual-ethical-reasoning-ai-2026]] — Ethical principles of AI in education: teachers' contextual ethical reasoning --- ## [Ethics](https://edtechdev.github.io/aied/concepts/ethics/) > **Ethics** — the moral principles governing the design, deployment, and use of AI in educational contexts. [[ai-education|AI education]] ethics spans [[privacy|data privacy]], algorithmic fairness, transparency, accountability, and the broader question of what AI should and should not do in learning environments. It is the normative foundation for the knowledge base's other AI-education concerns — [[equity-in-ai-education|equity]], [[bias-mitigation]], [[academic-integrity]], [[governance]], and [[pedagogical-safety]] — and the field is increasingly moving from abstract principle lists toward situated, context-sensitive, and institutionally-supported ethical practice. ## Questions to Consider - AI systems are often held to ethical checklists — fairness, privacy, transparency, accountability. But research increasingly suggests that principle-based 'checklist' ethics is insufficient on its own. What might a checklist miss that only happens in real, situated classroom practice? - Students' decisions about disclosing AI use are shaped less by ethical conviction than by fear of penalties, stigma, and unclear policy. If honesty is driven by fear rather than principle, what does that say about the ethics of the current academic-integrity approach? - Personalizing AI feedback with student attributes can itself become a vector for bias. How might a well-intentioned attempt to tailor feedback to an individual learner end up treating them less fairly — and how would you detect it? - Who holds the power in the relationship between institutions, AI systems, and learners — and how does that power imbalance shape consent, privacy, and student surveillance as analytics grow more granular? - Some researchers argue that protecting learners' cognitive and epistemic development — their ability to think and act independently — is itself an ethical obligation, especially as over-reliance on AI grows. Do you agree that 'preserving thinking' is an ethical duty of AI systems? ## Introduction Ethics is the normative backbone of AI in education: every AI tool, policy, and [[pedagogy|pedagogical]] decision encodes assumptions about what is fair, transparent, accountable, and safe. The knowledge base treats ethics as both a set of principles (fairness, privacy, transparency, autonomy, accountability) and a set of *practices* — how stakeholders actually reason about AI, how institutions govern it, and how learners are prepared to use it responsibly. Recent [[research-methods-aied|research]] documents a double movement: on one hand, a deepening concern that principle-based "checklist" ethics is insufficient; on the other, a growing emphasis on [[situated-learning|situated]] and ecological accounts of ethical judgment, and on the structural and institutional conditions that make responsible use possible. ## Core ethical dimensions - **Fairness and bias:** [[bias-mitigation]] and [[equity-in-ai-education]] research address whether AI systems treat all learners fairly. [[ai-scoring-language-bias-physics|Language bias]] and [[bias-mitigation]] studies document real-world inequities, and [[marked-pedagogies-linguistic-bias-writing-feedback|writing-feedback research]] shows that personalizing [[ai-feedback-quality|AI feedback]] with student attributes can itself become a bias vector. - **Privacy and consent:** [[privacy]] research examines data collection, student surveillance, and the power imbalance between institutions and learners — a constraint that grows sharper as analytics become more granular and AI-driven. - **Transparency and [[explainable-ai|explainability]]:** [[xai-education-framework|Explainable AI frameworks]] argue that students and teachers should understand how AI systems make decisions affecting them. This extends to student-facing transparency about their own AI use — [[ai-use-disclosure|AI use and disclosure statements]] — which the research shows is shaped less by ethical conviction than by fear of penalties, stigma, and ambiguous policy (see [[kirsanov-beyond-detection-ai-online-assessments-2026|Kirsanov et al.]], [[chang-should-i-tell-my-teacher-ai-disclosure-2026|Chang et al.]], [[vetter-hidden-cost-disclosure-genai-2026|Vetter et al.]], [[gonsalves-student-non-compliance-ai-declarations-2025|Gonsalves]]). - **Autonomy and agency:** [[cognitive-offloading|Over-Reliance]] and [[cognitive-offloading]] research raise ethical questions about whether AI use diminishes [[agency|learner agency]], and the field frames protecting learners' cognitive and epistemic development as itself an ethical obligation. - **Safety and harm prevention:** [[pedagogical-safety]] and [[hazra-safetutors-pedagogical-safety-2026|tutor harm research]] define ethical obligations for AI system developers, and [[pedagogy-ai-mistakes|mistake-based pedagogy]] shows how exposing learners to AI errors can activate rather than bypass ethical and [[metacognition|metacognitive]] scrutiny. - **Sustainability and environmental impact:** AI in education carries a material footprint — the carbon and water cost of [[llm-environmental-impact-student-usage-2026|large language models]] is significant given high adoption among university students — and a broader responsibility for sustainable, non-extractive design. The knowledge base treats this as an ethical dimension, not only a technical or cost concern: [[daniel-ai-sustainability-scoping-review-2026|Daniel et al. (2026)]] distinguish *AI for sustainability* (using AI to advance environmental and social outcomes) from *sustainable AI* (reducing AI's own footprint, including energy-efficient and [[open-source|on-premise]] deployment), and [[alsuhaymi-sustainable-education-ai-digitalization-2026|Alsuhami & Atallah]] argue AI supports sustainable education only when adoption is subordinated to explicit educational values rather than technologization and commodification. [[raffaghelli-situated-ai-ethics-2026|Situated-ethics work]] likewise foregrounds environmental costs and the marginalization of [[global-south|Global South]] knowledge as core ethical concerns. See [[sustainability]]. ## How stakeholders actually reason about AI ethics Ethical AI use is not only a matter of principles but of how the people involved interpret and apply them — and the research reveals systematic differences between groups. - **Faculty vs. students.** [[bilgic-sever-ethical-dimensions-ai-higher-ed-2026|Bilgiç & Sever (2026)]], a [[mixed-methods-research|mixed-methods]] study of 971 students and 135 faculty, found both groups supportive of ethical AI use but in different registers: faculty emphasized ethical principles while flagging a lack of institutional guidelines, whereas students valued AI's learning benefits but voiced uncertainty about who shares ethical responsibility. Both worried that excessive AI use could weaken [[critical-thinking|cognitive skills]] — a concern that frames ethical integration as preserving learners' cognitive development, not merely regulating tool use. Faculty scored high on individual responsibility yet rated institutional guideline adequacy lowest (M = 2.99), exposing a gap between personal ethics and structural support. - **Students vs. instructors in readiness.** [[fekete-ethical-ai-literacy-gaps-2026|Fekete (2026)]] found students report higher ethical awareness than instructors (4.03 vs. 2.44), yet instructors show stronger willingness to use AI — students interpret ethics through immediate coursework while teachers treat it as institutional clarity and integrity. Both report weak institutional support, and readiness develops through different channels: instructors' moral awareness grows with institutional and social support, while students' confidence correlates with [[self-efficacy]] and collaboration rather than formal instruction. - **A shared multidimensional structure.** The six themes in [[bilgic-sever-ethical-dimensions-ai-higher-ed-2026|Bilgiç & Sever]] — spanning data ethics, algorithm ethics, and pedagogical ethics — form an interconnected causal chain from fundamental principles (transparency, accountability, fairness, autonomy) to behavioral outcomes, underscoring that responsible AI use depends as much on clear institutional roadmaps as on individual awareness. ## From principles to institutional governance The knowledge base's research increasingly locates ethics in institutions and structures, not just individuals: - **The institutional responsibility gap.** [[bilgic-sever-ethical-dimensions-ai-higher-ed-2026|Bilgiç & Sever]] call for faculty [[educational-development|professional development]], ethics courses, and clear institutional guidelines; [[fekete-ethical-ai-literacy-gaps-2026|Fekete]] similarly finds that institutional support drives instructor readiness while students rely on informal, [[self-directed-learning|self-directed learning]] — a responsibility gap for [[educational-policy-ai|policy]]. - **Policy robustness varies.** [[adarkwah-genai-unesco-policy-2026|Adarkwah et al. (2026)]], analyzing [[generative-ai|GenAI]] policies at 30 top universities against UNESCO's eight-component framework, found core ethical principles widely embraced but [[inclusive-learning|inclusion]], equity, and [[sustainability]] often neglected — and national AI-readiness ranking did not predict strong institutional policy. Policies tend to be declarative and misconduct-focused rather than operationally assured. - **Ethics & data governance as an enabler.** [[learning-analytics-to-educational-interventions-2026|Svetec, Divjak & Kadoić (2026)]] identify ethics & data governance as one of seven enablers of trustworthy [[learning-analytics]]-based interventions — positioning [[trust|trustworthiness]] (ethical compliance, data security, transparent algorithms, pedagogical validity) as the prerequisite without which data-informed educational change is not meaningful. - **A [[meta-analysis-systematic-review|systematic review]] of [[engineering-education|engineering education]]** finds ethical AI guidance is predominantly student-facing and compliance-oriented (centered on [[academic-integrity]] and disclosure), while reciprocal accountability for faculty AI use and institutional responsibility remain underdeveloped — a pattern heightened by engineering's professional stakes in public safety and [[well-being]]. ([[ethical-use-ai-engineering-education-review-2026]]) - **A consolidated value framework for AIED ethics.** [[agarwal-ethical-values-norms-aied-2026|Agarwal et al. (2026)]], a [[meta-analysis-systematic-review|systematic review]] of 25 articles, consolidate the fragmented ethics literature into six main ethical values for [[ai-education|AI in education]] — non-discrimination, data stewardship, [[human-in-the-loop-ai|human oversight]], goodwill, explicability, and educational aptness — and map the ethical norms extracted from the literature onto a stakeholder-by-value matrix (developers, educational institutes, end users, regulators). The review finds norms distributed unevenly: developers attract the most, while end users receive the fewest and least actionable norms, and no norms on non-discrimination, data stewardship, or educational aptness address end users directly — student voices are essentially absent, with "end user" norms mostly actions other stakeholders take to enable teachers. The authors argue end users should have agency and active roles rather than being treated as passive beneficiaries, and note the values are tightly coupled and can conflict (e.g., explicability vs. accuracy/[[privacy]], non-discrimination vs. data stewardship), producing ethical dilemmas alongside power asymmetries between stakeholder sets. ## Toward situated and ecological AI ethics A major theoretical shift in the knowledge base is the critique of universalist, principle-based "checklist" ethics in favor of situated, ecological, and transformative accounts: - **Situated AI ethics.** [[raffaghelli-situated-ai-ethics-2026|Raffaghelli et al. (2026)]] fuse Bronfenbrenner's ecological systems theory with Cultural-Historical [[activity-theory-aied|Activity Theory]] to frame AI as a non-neutral socio-technical assemblage whose ethical implications are historically produced and locally negotiated. Across seven national cases, teachers are positioned as moral gatekeepers of AI use while lacking structural, institutional, and epistemic support — and the framework extends [[ai-education|AI literacy]] beyond technical skills toward critical, political, and ecological agency, including resistance to surveillance capitalism and environmental harm. - **The shift toward situated practice.** [[ai-ethics-bibliometric-2026|A bibliometric analysis of 282 articles]] shows that AI ethics discourse post-2021 increasingly frames ethics around professional judgment, trust, [[human-ai-collaboration|human-AI collaboration]], and interpretive practice rather than only technical compliance — with education a conceptually important context. - **A cultural-historical account of responsibility.** This reframing connects ethics to [[learning-theories]] and [[teacher-ai-competency]]: responsible AI use is treated as context-sensitive, collective, and transformative agency rather than individual compliance with checklists. - **An ontological reframing of ethical obligation.** Xie (2026) pushes the situated-ethics critique further by relocating it in ontology: where much scholarship treats equity, power asymmetries and environmental costs as externalities to be managed, a Daoist relational self makes them internal to who we are, so that "being left behind" becomes ontologically untenable rather than merely unfair. The ethical task is then not detached optimization but Wuwei (无为) — non-forced action aligned with naturalness — set against the Youwei (有为) of algorithmic monoculture, cognitive offloading and extractive infrastructure, applied across three domains: knowledge (epistemic monoculture and synthetic misinformation), knowing (offloading that degrades [[critical-thinking|critical thought]]) and impact (environmental costs and Global North–South asymmetries).([[daoism-ai-education-philosophy-2026]]) ## Ethics in practice The knowledge base's ethics articles range from theoretical frameworks ([[ethical-ai-higher-ed-game-theory|game theory approaches]]) to practical guidelines ([[cost-of-ethics-crisis-cs-ethics-education|CS ethics education]]), from [[ai-ethics-education-public-discourse|public discourse analysis]] to [[adarkwah-genai-unesco-policy-2026|UNESCO policy frameworks]]. Across this range, a consistent practical message emerges: ethics must move from policing misuse toward building [[ai-literacy]] and ethical-use capability, supported by institutional guidance, teacher professional development, and design that preserves rather than displaces learner judgment. That practical case is reinforced by empirical work on how ethics functions inside AI literacy itself: Zhu and Kong (2026) show that AI ethical awareness — alongside empowerment in AI [[problem-solving]] — mediates the relationship between perceived [[project-based-learning|project-based learning]] and satisfaction with an AI literacy course. Their validated AI-PBLS scale and SEM results (1,027 students) indicate that PBL creates conditions in which students build a personal framework for ethical reasoning about AI, strengthening the case that ethics is not an add-on but a core mechanism of meaningful AI literacy development. ## Connections Ethics connects to [[equity-in-ai-education]], [[privacy]], [[bias-mitigation]], [[regulation]], [[pedagogical-safety]], [[academic-integrity]], and [[governance]]. It is the normative foundation for all other AI education concepts — the frame within which questions of fairness, transparency, autonomy, and safety are raised and resolved. ## Connected Concepts - [[business-education]] - [[equity-in-ai-education]] - [[privacy]] - [[bias-mitigation]] - [[regulation]] - [[pedagogical-safety]] - [[academic-integrity]] - [[ai-use-disclosure]] — AI use and disclosure statements - [[governance]] - [[ai-literacy]] - [[cognitive-offloading]] - [[teacher-role]] - [[teacher-education]] - [[chemistry-education]] — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation - [[biology-education]] — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools ## Connected Articles - [[preservice-teachers-responsible-genai-2026]] — Pre-service teachers' responsible GenAI use: ethics, privacy, AI literacy (Kohnke et al. 2026) - [[student-centered-genai-responsible-framework-2026]] — Student-facing framework for responsible GenAI use in higher education (Alsammani 2026) - [[learning-analytics-to-educational-interventions-2026]] — From learning analytics to educational interventions: enablers of trustworthy LA-based interventions (Svetec, Divjak & Kadoić 2026) - [[kirsanov-beyond-detection-ai-online-assessments-2026]] — How students use and hide AI in online assessments - [[chang-should-i-tell-my-teacher-ai-disclosure-2026]] — Student AI disclosure, stigma, and self-regulated learning - [[vetter-hidden-cost-disclosure-genai-2026]] — The hidden cost of disclosure - [[gonsalves-student-non-compliance-ai-declarations-2025]] — Student non-compliance with AI use declarations - [[bilgic-sever-ethical-dimensions-ai-higher-ed-2026]] — Faculty and student views on ethical dimensions of AI - [[ethical-use-ai-engineering-education-review-2026]] — Ethical Use of AI in Engineering Education: A Systematic Review - [[ai-ethics-education-public-discourse]] — Longitudinal analysis of public discourse on AI ethics in education - [[ethical-ai-higher-ed-game-theory]] — Game theory framework for ethical AI use in higher education - [[cost-of-ethics-crisis-cs-ethics-education]] — Cost-of-ethics crisis in the job searches of CS students - [[xai-education-framework]] — Explainable AI in education (XAI-ED) - [[daoism-ai-education-philosophy-2026]] — Alternative AI Philosophy: Daoism as Method for AI in Education - [[hazra-safetutors-pedagogical-safety-2026]] — AI tutor safety and pedagogical harms - [[raffaghelli-situated-ai-ethics-2026]] — Situated AI ethics: a cultural-historical framework for education - [[substitution-to-scaffolding-ai-harm-cycle-2026]] — From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle - [[ssaho-ai-academic-integrity-review-2025]] — Culture of academic integrity as the ethical response to AI - [[ai-ethics-bibliometric-2026]] — AI Ethics and Professional Judgment: A Bibliometric Analysis (Mazlan et al. 2026) - [[adarkwah-genai-unesco-policy-2026]] — UNESCO generative AI policy framework analysis - [[luo-eaton-ai-student-feedback-ethics-2026]] — Is it ethical for teachers to use AI for student feedback? - [[fekete-ethical-ai-literacy-gaps-2026]] — Bridging ethical AI literacy gaps across students, educators, and policy - [[alharbi-ethical-genai-eap-2026]] — Ethical generative AI integration in English for Academic Purposes - [[ai-tools-academic-work-cheating-2026]] — Is using AI tools for academic work cheating? Student perceptions and ethics - [[alsuhaymi-sustainable-education-ai-digitalization-2026]] — Value-critical approach to sustainable education and AI (Alsuhami & Atallah 2026) - [[ethical-conditions-llm-exam-preparation-2026]] — Ethical conditions for LLM adoption in exam preparation (Pérez-Portabella et al. 2026) - [[utility-value-intervention-teach-responsibly-genai-2026]] — Utility-value intervention effects in learning to teach responsibly with GenAI (Boos, Eder & Lachner 2026) - [[longitudinal-ai-usage-ethics-policy-teacher-education-2026]] — Longitudinal GenAI usage, ethics, and policy in teacher education (Parker et al. 2026) - [[ai-literacy-course-satisfaction-pbl-scale-2026]] — AI-PBLS scale; empowerment and ethical awareness mediating PBL-to-satisfaction in AI literacy courses (Zhu & Kong 2026) - [[agarwal-ethical-values-norms-aied-2026]] — Ethical values and norms for AI in education --- ## [AI Misuse and Learning Harm](https://edtechdev.github.io/aied/concepts/ai-misuse-learning-harm/) > **AI misuse and learning harm** — the causal relationship between students offloading [[cognitive-offloading|cognitive work]] to [[generative-ai|generative AI]] and reduced durable learning, even when immediate task performance rises. The defining feature is a performance–learning gap: AI inflates assisted performance while degrading unassisted, closed-book, and retention outcomes. ## Questions to Consider - A student can perform better in the moment while learning less over time — AI can raise assisted performance while degrading later, unassisted achievement. When have you felt you 'learned' something that vanished the moment the tool was gone? - The central finding is a performance–learning gap: students who used an unguarded [[intelligent-tutoring|AI tutor]] scored 48% higher on practice but 17% lower on closed-book exams. How does that change how you'd interpret a strong grade earned with AI help? - Misuse is about substitution — delegating the drafting, recall, or analysis that builds durable understanding — rather than using AI as a complement. Where would you draw the line between AI helping you learn and AI learning for you? - Students in the harmed group did not perceive they learned less. If learners can't tell they're being harmed, what should an instructor watch for instead of trusting student self-reports? - The harm is selective by assessment type: it shows up on proctored, closed-book measures but can be hidden when coursework can't distinguish AI-assisted from independent work. What kind of assessment would actually reveal whether students learned? - Guardrailed 'hint-not-answer' AI eliminated the harm while preserving the benefit. What would it take to design AI help that supports learning without becoming a crutch? ## Introduction AI misuse is distinct from AI use. Use describes employing AI as a complement to learning — [[feedback]], brainstorming, or revision help that keeps the learner's cognitive work in the loop. Misuse describes substitution: delegating to AI the very mental processes (drafting, recall, analysis, revision) that build durable understanding. The harm documented in the knowledge base's evidence base is not that misuse fails to help; it is that misuse actively degrades later, unassisted achievement. ### The performance–learning gap The core concept, articulated in [[genai-performance-vs-learning]], is that generative AI easily boosts **performance** — immediate efficiency and output quality — while often bypassing the deep cognitive and [[metacognition|metacognitive]] processing required for **learning**. A tool that optimizes for performance can therefore undermine learning. The gap is now causally demonstrated at field scale: a [[rct|randomized controlled trial]] found unguarded AI assistance raised practice performance but reduced later unassisted exam scores. ### Mechanisms of harm - **[[cognitive-surrender]]** — the term researchers use for students offloading thinking to AI as a passive, unreflective dependency, as opposed to the deliberate, strategic form of [[cognitive-offloading]]. It produces a measurable population-level decline in durable knowledge. - **Answer-copying as a crutch** — misuse is driven less by AI errors misleading students than by students copying answers instead of learning. When [[student-engagement|engagement]] analysis shows students mostly "ask for the answer," learning harm follows. - **Motivation erosion** — the perceived availability of an effortless AI shortcut reduces autonomous motivation and persistence, per [[self-determination-theory|self-determination theory]]. Because persistence is what produces deep learning, its erosion compounds the direct harm. - **Learning displacement** — the substitution of AI output for the effortful processes (elaboration, recall, self-explanation) that consolidate knowledge, consistent with [[cognitive-offloading|Over-Reliance]]. ### The evidence base - **Causal field RCT (≈1,000 high-school math students):** an unguarded ChatGPT-style tutor raised assisted practice performance **+48%** but reduced unassisted, [[summative-assessment|closed-book exam]] scores **−17%** — students who never had AI access outperformed those who did. A guardrailed "hint-not-answer" tutor eliminated the harm. Notably, students in the harmed arm did not perceive they learned less. - **Population-scale behavioral data (3.2M ALEKS interactions):** study time on AI-susceptible problems fell **−26.9%** cumulatively for college students (high school −31.3%) after ChatGPT's release, with a **−25% decline in odds of a correct response on proctored retention items**. The effect vanished entirely under proctoring, pinning it on off-platform AI use. - **A large null result:** exploiting the seasonal drop in ChatGPT use over summer showed **no net change in high-school standardized test averages** — likely because misuse harm is offset in aggregate by productive AI use. This does not contradict the causal harm to durable learning; it cautions against over-generalizing from aggregate test scores. ### Dependency as a pathway to burnout Misuse harms more than achievement. A survey of 276 Chinese undergraduates ([[ai-dependency-self-efficacy-teacher-support-burnout-2026|Huang et al., 2026]]) modeled AI dependency as the mediator between protective learner resources and learning burnout, and found it carried the entire effect: [[self-efficacy|academic self-efficacy]] and, more weakly, [[teacher-role|teacher]] support both reduced AI dependency, which in turn predicted burnout. Both indirect paths were significant and fully mediating — the direct effects became non-significant once dependency entered the model — with self-efficacy's path the stronger of the two. Read alongside the performance–learning gap above, the result extends the cost of misuse from degraded durable knowledge to learner exhaustion and disengagement, and it locates the mechanism in dependency itself rather than in AI use as such: the resources that protect against misuse appear to work by preventing dependency, not by counteracting its effects afterward. The authors flag a measurement caveat — their AI Dependency Scale is adapted from the Internet Addiction Test and its content validity for AI has not been established — so the pathway is better treated as a well-modeled hypothesis than a settled effect size. ### The assessment-dependent nature of harm The most important practical nuance is that the harm is **selective by assessment type**. It shows up on **proctored, closed-book, and unassisted** measures of durable knowledge. On normal graded coursework that cannot distinguish AI-assisted from independent work, misuse can *inflate* immediate grades. This is why the perceived-vs-actual gap is dangerous: students (and sometimes instructors) see short-term performance gains and miss the erosion of learning that only surfaces when the tool is removed. **Which reliance, not how much of it.** A survey of 118 students across 12 AI-intensive courses ([[uneven-impact-generative-ai-student-learning-2026|Manikonda et al. 2026]]) separates two behaviors that usage-frequency measures collapse. **Cognitive reliance** — using GenAI to organize, evaluate, combine and decompose — was the strongest predictor of perceived positive impact (β = .495, p < .001), and the offloading it describes is not uniformly harmful: organizing and summarizing ideas and discovering new insights ranked among the features associated with improved perceived learning. **Early reliance** — consulting GenAI *before* independent thought, a traditional search, or an instructor — predicted academic benefit (β = .301, p < .001) *and* negative impact (β = .402, p = .004) at once. The moderation is the part that bears on the literacy remedy above: the association between early reliance and negative impact was absent at low [[ai-literacy|evaluation literacy]] (b = .025, p = .861) and strongest among the students who judge AI output best (b = .688, p < .001), so evaluative skill did not protect against the cost of asking AI first — the students best placed to notice the cost were the ones reporting it. The operational implication is that misuse is partly a question of *timing* relative to the learner's own attempt, not only of volume, which is why interventions that target when AI is consulted sit alongside the assessment-design remedies below. ### Implications and remedies - **[[guardrails]] over raw access:** hint-not-answer [[prompt-engineering|prompting]] and teacher-authored [[scaffolding]] neutralize the crutch effect (see [[generative-ai-guardrails-harm-learning]]). - **[[assessment|Assessment design]]:** AI-resistant and proctored/unassisted assessments are needed to surface — and discourage — misuse. - **Literacy and metacognition:** [[ai-literacy]] and [[self-regulated-learning]] training that helps students recognize reliance patterns and the cost of bypassing their own cognitive work. ## Connected Concepts - [[self-directed-learning]] - [[remote-proctoring]] - [[cognitive-offloading]] - [[academic-integrity]] - [[assessment]] - [[self-regulated-learning]] - [[motivation]] - [[metacognition]] - [[scaffolding]] - [[generative-ai]] - [[student-experience]] - [[cognitive-surrender]] ## Connected Articles - [[genai-thoughtless-use-self-directed-learning-2026]] - [[best-response-student-ai-dialog-2026]] - [[ai-tools-academic-work-cheating-2026]] - [[generative-ai-guardrails-harm-learning]] — GenAI Without Guardrails Can Harm Learning - [[generative-ai-reduced-study-time-math]] — Generative AI Reduced Study Time on Math - [[genai-performance-vs-learning]] — Distinguishing Performance Gains from Learning - [[chatgpt-impact-high-school-tests]] — Little Impact of ChatGPT on High School Test Scores - [[ai-availability-student-motivation]] — AI Availability and Student Motivation - [[genai-skill-bypass-literacy]] — GenAI Skill Bypass and Literacy - [[cognitive-shift-ai-education]] — Cognitive Shift in AI Education - [[misiejuk-cognitive-offloading-prompting-2026]] — Cognitive Offloading in Student–AI Collaboration - [[ssaho-ai-academic-integrity-review-2025]] — AI misuse in academic writing and integrity breaches - [[cognitive-commons-ai-expertise-regeneration]] — The tragedy of the cognitive commons: AI and expertise regeneration - [[shaw-nave-cognitive-surrender-2026]] — Tri-System Theory and cognitive surrender: how AI reshapes human reasoning (Shaw & Nave 2026) - [[lodge-loble-cognitive-offloading-2026]] — AI, cognitive offloading and implications for education (Lodge & Loble 2026) - [[ai-overreliance-complex-adaptive-system-2026]] — AI overreliance modeled as a complex adaptive system - [[academic-erasure-complexity-ai-writing-2026]] — Academic erasure: the disappearance of complexity under AI-supported writing - [[ai-dependency-self-efficacy-teacher-support-burnout-2026]] — AI dependency fully mediates the path from self-efficacy and teacher support to learning burnout (Huang et al. 2026) - [[uneven-impact-generative-ai-student-learning-2026]] — Early reliance predicts negative impact while cognitive reliance predicts positive impact, and evaluation literacy strengthens rather than buffers the harm of consulting AI first (Manikonda et al. 2026) --- ## [Legal Issues and Risks](https://edtechdev.github.io/aied/concepts/legal-issues-and-risks/) > **Legal Issues and Risks** — the exposure institutions, staff and students incur when [[generative-ai|generative AI]] is governed badly in education: a student wrongly accused of cheating on the strength of a detector score, a proctoring system that watches and records more than the assessment requires, an over-broad rule that penalizes an assistive tool, or a policy too vague to be enforced consistently. The risk is not one legal question but several, arriving together — evidentiary (whether the accusation can be evidenced at all), contractual and procedural (whether the institution followed its own rules and gave the student a fair hearing), equality-based (whether the rule burdens disabled or non-native-speaker students), and data-protection-based (what the surveillance collected and where it was stored). It is distinct from [[academic-integrity]], which is the conduct framework being enforced: this page is about what happens when that enforcement is challenged. ## Questions to Consider - Detection tools cannot reliably identify authorship. If the instrument cannot establish the fact at issue, what is a misconduct case actually resting on? - Proctoring and AI-detection data are generated at scale and retained indefinitely. Who carries the legal exposure for that data, the institution or its vendor? - When a policy prohibits "the use of AI" without distinguishing [[wright-transcription-not-generation-2026|transcription from generation]], is the rule protecting integrity or penalizing a disability accommodation? ## Introduction The knowledge base's evidence on this topic is procedural rather than juridical. It documents what institutions treat as evidence, how accusation procedures work, how unreliable the instruments are, and what surveillance collects; it does not yet document litigation outcomes. That gap should be stated plainly rather than filled with confident claims: the cases that would settle these questions are mostly unreported, settled, or still in internal institutional processes. What the literature does support is a description of the failure modes that create legal exposure. The recurring pattern is that the institution's own instruments and procedures, not a malicious accuser, are what put it at risk: a probabilistic score treated as a finding, a rule that is clearer in intent than in scope, a system that collected data nobody asked about, and a hearing that assumed the technical evidence needed no scrutiny. ## Where the risk concentrates ### Wrongful accusation and defective evidence [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] analyzed actual generative AI misconduct allegation files and sorted the evidence institutions used into categories: system-recorded behavioral traces available only in invigilated or supervised assessment, process evidence such as drafts, supervision meetings and presentations where those practices exist, and evidence generated by the investigation itself. Two things follow for legal exposure. First, in unsupervised submissions the system-recorded category is empty, which pushes cases onto weaker categories. Second, they record that principles of natural justice require a student to be informed of the allegation and given an opportunity to respond before any determination, obligations codified in Australian regulatory standards (Department of Education, 2021; TEQSA, 2025) as well as in well-regarded academic integrity policy. The response opportunity is typically an investigative meeting or panel interview, and whatever the student says becomes part of the evidentiary record — which means procedural failures, not just evidentiary ones, are where a case becomes vulnerable. The evidentiary problem sits underneath it. Detector output is the evidence most often reached for and the least able to bear the weight. [[hadra-ai-detector-accuracy-efl-2026|Hadra et al. (2026)]] tested Turnitin and Originality on a balanced corpus of 192 texts and found overall accuracy of 0.69 and 0.61 respectively, with both performing poorly on hybrid human-AI writing — the form most likely to appear in a real allegation — and accuracy falling further with text length, on scientific writing, and with a borderline tendency to misclassify human-written work as AI when the writer was an EFL student. [[van-vlasselaer-ai-detector-reliability-2026|Van Vlasselaer et al. (2026)]] reach the same conclusion from a different corpus and tool set, and [[bassett-ai-detectors-education-2026|Bassett et al.]] make the structural point that no threshold resolves the problem: a detector tuned to catch AI use will flag human work, and one tuned to spare human work will miss AI use, so any single score is a choice about which error to make. [[karr-ai-detection-humanization-2026|Karr's review]] of why detection fails reaches the same place from the writing-humanization side, and [[teichmann-detecting-undetectable-misconduct-2026|Teichmann et al. (2026)]] argue the procedural framework itself now needs reassessment, because the era of undetectable misconduct breaks the assumption that misconduct can be evidenced by the submitted artifact. ### Privacy and surveillance [[harerimana-remote-proctoring-nursing-scoping-2026|Harerimana et al. (2026)]] map [[remote-proctoring|remote proctoring]] across nursing assessment and surface privacy and surveillance among its principal concerns, alongside the emotional impact on students and [[equity-in-ai-education|equity]] effects. [[automated-online-exam-proctoring-decade-review-2026]] and [[academic-dishonesty-automated-proctoring-ai-2026]] document the same [[ai-technologies|technologies]] over a longer window, and the data-protection questions they raise are ordinary ones with legal consequences: what is captured (video, audio, keystrokes, gaze, room scans), how long it is retained, where it is stored, who can access it, whether the vendor processes it onward, and whether students consented to it as a condition of assessment. Institutions operating in regulated data environments carry statutory obligations well before any lawsuit appears, and the knowledge base's FERPA- and GDPR-aware work on local and vendor-hosted AI systems shows how the same questions apply to teaching tools rather than only to proctoring. ### Accessibility and disability [[wright-transcription-not-generation-2026|Wright (2026)]] argues that blanket "AI use" prohibitions are over-inclusive because they do not distinguish speech-to-text transcription and OCR from generative drafting, and that students with conditions affecting fine motor control, handwriting legibility or typing accuracy have historically relied on exactly those tools — including standalone voice-to-text products such as Dragon NaturallySpeaking, several of which have been discontinued or degraded, with AI-powered transcription filling the functional gap. Wright notes the intersection of disability, assistive technology and AI misconduct policy is underexplored and that the scale of this displacement has not been empirically measured. The exposure is straightforward in shape: a rule that removes a student's primary means of producing legible work is a rule that may need an accommodation process to survive. [[shin-ai-policies-sld-2026]] documents the same void from the policy side for students with specific learning disabilities, and the knowledge base's work on [[assistive-technology]] and [[neurodiversity]] supplies the surrounding terms. ### Language equity and the support–substitution boundary [[li-genai-assessment-language-equity-2026|Li (2026)]] supplies the equality-based version of the argument for students who use English as an additional language. Because a single interface now performs both permitted editing and prohibited drafting, a rule that treats [[generative-ai|generative AI]] as one category of unauthorized assistance imposes higher compliance burdens on the students most likely to need legitimate language support, concentrates suspicion on writers whose surface fluency has shifted, and enables selective enforcement on weak evidence — with allegations carrying reputational, academic and sometimes visa or financial consequences. The remedy is a boundary defined by function and the assessment construct rather than by tool name, separating surface interventions that add no ideas, sources or analytical structure from substitution that creates or materially reshapes the intellectual work, with calibrated [[ai-use-disclosure|disclosure]] so that routine translation and editing do not attract compliance costs exceeding those borne by monolingual peers. The legal structure is indirect-discrimination reasoning — identify the cohort-skewed burden a facially neutral rule produces, then ask whether a legitimate aim is pursued by proportionate and practically workable means — reinforced by the administrative-law expectation that a decision-maker can state the rule applied, the evidence relied on, and why the outcome was proportionate, which is what makes a determination reviewable and the process legitimate. [[ai-detection|Detection]] is demoted to a triage signal, with draft histories, staged submissions and a short construct-aligned conversation preferred as evidence, so the exposure runs both ways: to a discrimination or review challenge, and to the fragility of a finding resting on proxies such as polished language or non-native phrasing. ### Unclear rules, inconsistent enforcement [[gutowski-hurley-genai-policy-legal-education-2025|Gutowski and Hurley (2025)]] treat clarity as the precondition for defensible enforcement rather than a courtesy, and report that law school policies range from comprehensive [[governance]] to no stated policy, with most leaving individual instructors to interpret and apply the rules. [[qian-governing-genai-higher-ed-policy-2026|Qian (2026)]] finds the same variation across US universities together with support ecosystems that differ just as much, and [[crompton-governing-genai-higher-ed-delphi-2026]] reports the expert consensus that governance is fragmented. A student disciplined under a rule no one can state precisely is a dispute about procedure and fairness before it is a dispute about AI, and [[sharma-judgment-visible-genai-assessment-2026|Sharma (2026)]] argues that the remedy is [[assessment|assessment design]] that makes judgment and responsibility visible rather than surveillance that infers them. [[watson-rainie-ai-challenge-faculty-survey-2026|Watson and Rainie's (2026)]] survey of 1,057 US faculty shows the same inconsistency from the other direction: 87% of respondents wrote their own assignment-level rules while only 48% said their institution had written guidelines and 35% said their department had, so students inside a single institution meet a patchwork of individually authored policies. The structural response beneath those documents is thin — a task force or oversight group in 55% of cases, but [[ai-literacy|AI literacy]] adopted as a general education outcome in only 13% — and it matters legally because enforcing a rule the institution never adopted is difficult to defend. [[coates-governing-academic-integrity-indicators-2025|Coates, Croucher and Calderon (2025)]] locate the weakness further upstream, in governance rather than in student conduct or instrument quality. Their academic integrity indicator framework — 130 items under eight dimensions running from design and development through analysis, reporting, evaluation and improvement — is written as governance questions for boards and committees: whether the institution's top-most council receives updates on assessment processes and outcomes, whether key performance indicators cover assessment quality, whether induction and orientation include academic integrity, and whether there is a simple route for referring contract cheating cases. Their reform program targets governance architectures, the people holding governance roles, and the technologies and resources supporting assessment, and they argue that such development is unlikely to pay out without external affordance from [[regulation]], [[benchmark|benchmarking]] and cross-institutional competition. For legal exposure the implication is that a defensible position rests on knowing and documenting one's own practice — the same information an institution needs when a determination is challenged. ### Attacking the automated grader A different exposure sits inside the instruments themselves. [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] hid five indirect prompt injections inside the files of a synthetic essay that an institutional [[automated-assessment|AI grading tool]] (Microsoft Copilot, GPT-5.2) had graded as fail on six of six baseline runs. Two strategies raised the grade with no visible warning to the user, at reported attack success rates of 100% (9 of 9 iterations) and 94% (17 of 18), by combining instruction manipulation, role-playing and obfuscation; the tool silently disabled a chat after blocking the simplest attack, and in one iteration announced it would never follow embedded instructions and then raised the grade on each of the next six runs. A grade obtained through a hidden instruction carries no [[assessment-validity|validity]] claim, and the same technique could be used to degrade a submission with no durable trace left in the output — which means the appeal record for a challenged automated decision may be empty, and a finding of misconduct (or of merit) cannot be evidenced from the artifact at all. Humble's sector-level asks are clear [[educational-policy-ai|AI policy]], [[educational-development|professional development]] and standardized, domain-agnostic resilience testing so the attack surface is measured rather than assumed, with restricted AI use and [[human-in-the-loop-ai|human review]] reserved for high-stakes work. ### Beyond the campus gate For professional programs the exposure does not end at graduation. Gutowski and Hurley record that the professional conduct rules binding practicing lawyers — the duty of technological competence, confidentiality, supervision of others using the tools, candour toward the tribunal — already attach to AI use, and that [[hallucination-risk|hallucinated]] authority has produced sanctions for practitioners who filed fabricated cases. The same transfer logic applies wherever a license, registration or statutory duty follows the graduate, which is why the discipline pages for [[legal-education|Legal Education]] and the [[medical-education|health professions]] belong next to this one. ## Open Questions - Which of these risks has actually produced litigation or regulatory findings? The knowledge base has procedures, policies and technical evaluations, but no case outcomes, and it should not be read as though it did. - Does detector output survive as evidence of anything once an institution concedes its error rates in a hearing, or does the concession convert the case into a procedural-fairness dispute? - Is the vendor or the institution the data controller when proctoring and detection run through a third-party platform, and does that change the advice to institutions? - Should institutions publish the evidence standards they apply to AI misconduct, in the way that evidential thresholds are published elsewhere, as a way of reducing both wrongful accusation and legal exposure? ## Connected Concepts - [[academic-integrity]] — the conduct framework whose enforcement carries the risk - [[ai-detection]] — the instrument at the center of wrongful-accusation cases - [[remote-proctoring]] — surveillance in assessment and its data-protection questions - [[assessment-validity]] — whether the evidence can establish the claim made from it - [[privacy]] — collection, retention and onward processing of student data - [[regulation]] — statutory and regulatory obligations institutions must meet - [[governance]] — internal policy design and consistency of enforcement - [[educational-policy-ai]] — institutional AI policy as the source of unintended exposure - [[ai-use-disclosure]] — disclosure expectations and their unenforceable edges - [[accessibility]] — reasonable adjustment where a rule removes an assistive tool - [[assistive-technology]] — the tools at the center of the over-inclusion problem - [[neurodiversity]] — the students most exposed to over-broad prohibitions - [[equity-in-ai-education]] — differential burden of detection and surveillance - [[student-experience]] — the human cost that precedes the legal one - [[hallucination-risk]] — fabricated authority as professional and institutional liability - [[reducing-ai-misuse]] — the prevention alternative to accusation ## Connected Articles - [[munoz-misconduct-allegation-evidence-2026]] — What evidence misconduct allegation files actually contain, and the natural justice requirement - [[hadra-ai-detector-accuracy-efl-2026]] — Detector accuracy, hybrid text failure, and EFL misclassification risk - [[van-vlasselaer-ai-detector-reliability-2026]] — Reliability of detection tools across a second corpus - [[bassett-ai-detectors-education-2026]] — Why no detection threshold can be right: the error-tradeoff argument - [[teichmann-detecting-undetectable-misconduct-2026]] — Misconduct procedures reassessed when evidence has become undetectable - [[wright-transcription-not-generation-2026]] — Over-inclusive AI rules, disability accommodation and reasonable adjustment - [[harerimana-remote-proctoring-nursing-scoping-2026]] — Remote proctoring mapped, with privacy and surveillance concerns - [[automated-online-exam-proctoring-decade-review-2026]] — A decade of automated proctoring research - [[academic-dishonesty-automated-proctoring-ai-2026]] — Academic dishonesty and proctoring in the AI era - [[gutowski-hurley-genai-policy-legal-education-2025]] — Policy clarity as a precondition for defensible enforcement - [[qian-governing-genai-higher-ed-policy-2026]] — Policy and support ecosystems across innovative US universities - [[crompton-governing-genai-higher-ed-delphi-2026]] — Expert consensus on fragmented governance - [[sharma-judgment-visible-genai-assessment-2026]] — Integrity through visible judgment rather than surveillance - [[shin-ai-policies-sld-2026]] — The policy void for students with specific learning disabilities - [[li-genai-assessment-language-equity-2026]] — Language equity as a rule-design problem: the support–substitution boundary, indirect discrimination and reviewability (Li 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Students attacking AI graders by indirect prompt injection, with grades changed undetected (Humble 2026) - [[coates-governing-academic-integrity-indicators-2025]] — Governance indicators and reform program for authenticating assessment (Coates, Croucher & Calderon 2025) - [[watson-rainie-ai-challenge-faculty-survey-2026]] — 1,057 US faculty: individual policies far outrun institutional ones, structural response thin (Watson & Rainie 2026) --- ## [AI Use and Disclosure Statements](https://edtechdev.github.io/aied/concepts/ai-use-disclosure/) > **AI use and disclosure statements** — the policies, declarations, and practices through which learners are asked (or choose) to disclose their use of [[generative-ai|generative AI]] in academic work. Also known as AI use declarations, AI disclosure statements, or [[explainable-ai|transparency]] statements, these mechanisms sit at the intersection of [[academic-integrity]], [[ethics]], [[trust]], and [[assessment]] in the AI era. The knowledge base's [[research-methods-aied|research]] shows that disclosure is far from a neutral administrative formality: it is shaped by fear of penalties, ambiguous policies, inconsistent enforcement, peer norms, stigma, and the psychological costs of self-incrimination — and it is deeply entangled with [[self-regulated-learning]] and [[help-seeking]]. ## Questions to Consider - If you were a student who used generative AI for an assignment, would you tell your instructor? This page suggests the honest answer is heavily shaped by fear — research finds disclosure is rare, driven by 'penalty anxiety' and worries about self-incrimination. What assumptions about your own disclosure behavior might you need to question? - The central design question here is whether disclosure functions as surveillance or as a scaffold for ethical engagement. Before reading, which did you assume a disclosure form was trying to do — and can you think of a way the same form could serve the opposite purpose? - One striking finding is that students who always disclosed had over three times the odds of being accused of cheating — transparency invited suspicion. If honesty can increase suspicion, what does that say about how we should design trust between students and instructors? - This page distinguishes mandatory compliance mechanisms from formative disclosure practices that build AI literacy and self-regulation. Which of these do you think your own institution's approach most resembles, and what would have to change for it to become the other? - The research connects concealment to maladaptive self-regulation, and disclosure to adaptive help-seeking — making your AI use visible can invite useful calibration from an instructor. When have you (as learner or [[teacher-role|teacher]]) found it beneficial to make your uncertainty or use of help visible, rather than hiding it? - Disclosure norms vary by discipline, language background, and context — with equity implications when some students are more likely to be penalized for concealing than others. If you were designing a disclosure policy, how would you prevent it from placing an uneven burden on already-marginalized learners? ## Introduction AI use and disclosure statements are the declarations through which students report whether and how they used generative AI in an assignment — from a checkbox on a coversheet to structured formats recording the model, the input, the purpose, and how the output was evaluated. Systems differ in kind as well as detail: some operate as mandatory compliance mechanisms with penalties for non-disclosure, others as [[formative-assessment|formative]] practices that treat disclosure as an instrument for developing [[ai-literacy]], transparency and ethical judgment. The design question that decides which they become is whether disclosure is used to catch or to teach, which places the concept squarely inside [[academic-integrity]] and [[educational-policy-ai]]. ## What AI use disclosure statements are AI use declarations require or invite students to state whether and how they used generative AI in an assignment — typically on a coursework coversheet, submission form, or reflection prompt. Frameworks vary widely, from simple checkbox declarations to structured formats that document the model, the input, the purpose, and how the output was evaluated (e.g., the Model-Input-Evaluation/MInE framework). They range from **mandatory compliance mechanisms** (with penalties for non-disclosure) to **[[formative-assessment|formative]] practices** that treat disclosure as a [[pedagogy|pedagogical]] instrument for developing [[ai-literacy]], transparency, and self-[[regulation]]. The central design question is whether disclosure functions as surveillance or as a [[scaffolding|scaffold]] for ethical [[student-engagement|engagement]]. ## Why disclosure matters for AI in education Disclosure is the mechanism that makes AI-assisted work *visible* and therefore governable — the counterpart to [[academic-integrity]] and the precondition for [[trust]] between students and instructors. Without transparency about AI use, educators cannot distinguish legitimate support from misconduct, calibrate [[feedback]], or detect equity gaps. But the research reveals that disclosure is also an **[[affective-computing|affective]] and social process**, not just a policy one. ## What the research shows - **Disclosure is rare and anxiety-driven.** Across studies, most students do not disclose AI use to instructors. [[kirsanov-beyond-detection-ai-online-assessments-2026|Kirsanov et al. (2026)]] found only ~34% of economics students admitted AI use and disclosure was rarer, driven by "penalty anxiety." [[vetter-hidden-cost-disclosure-genai-2026|Vetter et al. (2026)]] found 66% of students never or rarely disclosed, with fear of academic penalty the top reason. [[gonsalves-student-non-compliance-ai-declarations-2025|Gonsalves (2025)]] documented **74% non-compliance** with a mandatory declaration at King's [[business-education|Business School]]. - **Ambiguous and punitive policies drive concealment.** When AI-use policies are unclear, unenforced, or bans, students hide use rather than engage. "Limited use in certain situations" policies best predict disclosure; "not allowed" or "not aware" policies predict concealment. Clear, consistent, collaborative policies are repeatedly called for. - **Disclosure has psychological costs — and hidden costs.** Students view declarations as self-incrimination, and some believe AI use is private like using a calculator. [[vetter-hidden-cost-disclosure-genai-2026|Vetter et al. (2026)]] found students who "always" disclosed had over 3× the odds of being accused — transparency can invite suspicion. [[chang-should-i-tell-my-teacher-ai-disclosure-2026|Chang et al. (2026)]] found worry redirects disclosure toward peers rather than suppressing it, cutting students off from instructor feedback. - **Disclosure is entangled with self-regulated learning.** [[chang-should-i-tell-my-teacher-ai-disclosure-2026|Chang et al. (2026)]] interpret teacher-directed disclosure as adaptive help-seeking — making assistance visible and inviting external calibration — while concealment (especially among heavy AI users) resembles maladaptive regulation. Structured, formative disclosure can promote the [[metacognition|metacognitive]] reflection and ethical reasoning that characterize effective self-regulation. - **Disclosure norms vary by discipline, language, and context.** Education, social-science, and [[stem-education|STEM]]/Health students disclose more than Business students; monolingual students may disclose less than [[multilingual-learning|multilingual]] students. International students in one study disclosed less, raising [[equity-in-ai-education|equity]] concerns. Disclosure norms do not develop automatically with academic progression. - **Disclosure norms are unsettled in professional programmes.** [[gutowski-hurley-genai-policy-legal-education-2025|Gutowski and Hurley (2025)]] report that ABA-approved US law schools disagree about whether any use of AI must be disclosed, describe student disclosure forms filed with instructors, and find no consensus on how scholarly or class writing should cite AI — the Bluebook itself remains silent. They also predict that disclosure will eventually look as pointless as noting that a student used a search engine, a caution about the shelf life of the rules being written now, and they ground the norms question in [[legal-education]]. ### Disclosure is not a detection mechanism A recurring institutional error is to treat declarations as a way to catch prohibited use. They cannot serve that function: they depend on candour, and enforcing them runs back into the same undetectability, since an institution generally cannot prove that an undeclared use occurred. [[teichmann-detecting-undetectable-misconduct-2026|Teichmann (2026)]] locates their real value elsewhere — in visible integrity commitments embedded in a culture of [[trust]], and in transparency, shared expectations, and student reflection rather than enforcement. The educational mechanism is normative, not forensic. The modeling work of [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi (2026)]] supplies the condition under which disclosure works, expressed as a design threshold rather than a hope: disclosed use beats hidden use only when the [[academic-integrity|cost of honesty]] stays below the deterrent it buys. Because that bound rises with detection credibility, detection and disclosure safety reinforce each other — but only if the false-positive risk of being flagged falls on honest and dishonest responses alike, in which case honesty keeps its comparative protection. Their design implication is direct: treat a declared AI use as context rather than a confession, which keeps the cost of honesty low precisely in the settings where monitoring is strongest, and read disclosure as a demonstration of [[evaluative-judgment|evaluative judgment]] rather than an admission. Permission and disclosure are separate levers — permission changes the formal boundary of acceptable use, disclosure changes visibility — and success at one says nothing about the other. [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]] offers a suggestive test of the disclosure effect itself. The same course ran a blinded study (AlGhamdi 2024) and a transparent one, and only the transparent cohort explicitly bounded AI's role and questioned evaluative authority. Because the cohorts and years differ, the author calls the contrast suggestive rather than causal, but argues that disclosure operated as an instructional prompt rather than a disclaimer — transforming feedback from a taken-for-granted act into an object of reflection, with no sign of disengagement or resistance. Surveying the GenAI guidance of the 50 US universities ranked most innovative, [[qian-governing-genai-higher-ed-policy-2026|Qian (2026)]] locates disclosure at the boundary between assistance and misrepresentation: where AI is permitted, institutions instruct students to acknowledge, cite or otherwise document AI assistance rather than present outputs as entirely their own, with libraries — not integrity offices — holding authority over citation and provenance standards. MIT Libraries' framing that AI is not an author, leaving authors responsible for documenting how tools contributed, exemplifies the pattern, and Qian reads disclosure simultaneously as an integrity norm, a design principle that reduces ambiguity before submission, and one of four integrity pillars alongside accountability, [[equity-in-ai-education|equity]] and [[privacy]]. This is the same non-forensic use the section describes, and it is why Qian recommends institutionalizing libraries as the citation authority rather than expanding detection. ## Designing effective disclosure - **Clarity and consistency beat deterrence.** Develop clear, consistent, collaboratively-built AI policies with concrete examples of acceptable use and how to declare it; avoid punitive or vague frameworks that motivate concealment. - **Address the affective barrier.** Disclosure policies that target behavior alone are insufficient — normalize AI use, reduce perceived judgment and stigma, and reassure students that honest disclosure will not be penalized. - **Make disclosure formative, not just compliance.** Use reflection prompts and structured declarations (model/input/evaluation) so disclosure builds [[ai-literacy]] and self-regulation rather than functioning as surveillance. - **Train faculty to build trust, not rely on detection.** Disclosure should never be followed by accusation; faculty should learn trust-building conversations instead of leaning on unreliable AI detectors. - **Design for equity.** Attend to differential disclosure across disciplines, language status, and student backgrounds, and avoid mechanisms that place uneven burdens on already-marginalized learners. ## Connections AI use and disclosure sits at the intersection of [[academic-integrity]] (its parent concern), [[ethics]], [[trust]] and [[trust-calibration]] (disclosure as an act of vulnerability), [[governance]] and [[educational-policy-ai]] (institutional policy frameworks), and [[assessment]] (where disclosure is operationalized). It connects to [[ai-literacy]] (students need to understand what and how to disclose), [[self-regulated-learning]] and [[help-seeking]] (disclosure as visible help-seeking), and [[equity-in-ai-education]] (differential disclosure). It is distinct from, but related to, [[ai-misuse-learning-harm]] (the harm disclosure is meant to make visible) and [[ai-detection]] (detection-based alternatives that disclosure frameworks increasingly replace). ## Connected Concepts - [[academic-integrity]] - [[ethics]] - [[trust]] - [[trust-calibration]] - [[governance]] - [[educational-policy-ai]] - [[assessment]] - [[authentic-assessment]] - [[ai-literacy]] - [[self-regulated-learning]] - [[help-seeking]] - [[equity-in-ai-education]] - [[generative-ai]] - [[higher-ed]] - [[ai-misuse-learning-harm]] — the harm disclosure is meant to make visible - [[ai-detection]] - [[parents-and-families]] - [[social-norms-ai-use]] — why learners choose to hide rather than disclose ## Connected Articles - [[guided-inquiry-genai-course-policy-2026]] — Students co-designing GenAI course policies via guided inquiry (Hingle & Johri 2026) - [[kirsanov-beyond-detection-ai-online-assessments-2026]] — Beyond detection: how students use and hide AI in online assessments - [[chang-should-i-tell-my-teacher-ai-disclosure-2026]] — Student AI disclosure, stigma, and self-regulated learning - [[vetter-hidden-cost-disclosure-genai-2026]] — The hidden cost of disclosure: disclosure and faculty accusations - [[gonsalves-student-non-compliance-ai-declarations-2025]] — Student non-compliance with AI use declarations - [[ethical-conditions-llm-exam-preparation-2026]] — Ethical conditions for LLM adoption in exam preparation (Pérez-Portabella et al. 2026) - [[teichmann-detecting-undetectable-misconduct-2026]] — Declarations as education rather than detection - [[mohamed-temimi-assessment-imperfect-information-disclosure-2026]] — The cost of honesty as a design threshold in assessment - [[student-perspectives-ai-writing-grading-2026]] — Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education - [[qian-governing-genai-higher-ed-policy-2026]] — Disclosure as the boundary between assistance and misrepresentation across 50 innovative US universities (Qian 2026) - [[gutowski-hurley-genai-policy-legal-education-2025]] — No consensus on disclosure and AI citation in law school policy (Gutowski & Hurley 2025) - [[ai-written-admissions-essays-penalized-2026]] — AI-written admissions essays are widespread but penalized --- ## [Guardrails](https://edtechdev.github.io/aied/concepts/guardrails/) > **Guardrails** are the explicit design mechanisms, constraints, and intervention points that keep an [[ai-education|AI education]] system within pedagogically safe behavior — the *how* that operationalizes the *goal* of [[pedagogical-safety]]. They are the difference between a raw general-purpose [[conversational-ai|chatbot]] and a tutoring tool that reliably preserves learning. Guardrails are not a single feature but a layered set of controls spanning prompt design, knowledge grounding, reward shaping, deployment QA, and ongoing auditing. ## Questions to Consider - A tutor that gives no wrong answers can still quietly harm learning. What kinds of 'quiet' failures might escape a toxicity check but still undermine how much students actually learn? - In a field experiment, an unguarded [[intelligent-tutoring|AI tutor]] raised practice performance but reduced later unassisted exam scores, while a 'hint-not-answer' version eliminated the harm. Why might making students perform better in the moment actually make them learn less? - If an AI tutor is engineered to be 'kind' — never pushing back or giving corrective [[feedback]] — how could that be a safety problem rather than a feature? When is agreeable behavior harmful in an educational context? - Guardrails are described as a layered set of controls, from prompting to knowledge grounding to training to auditing. Pick one layer and consider: where could it fail, and what would a different layer catch that it misses? - The page notes that guardrails themselves can be biased — refusals and softened answers patterned by [[learner-identity|student identity]]. How would you audit a safety filter to make sure it isn't quietly reproducing inequity while 'protecting' learners? - Younger learners are described as least equipped to detect manipulative or sycophantic AI behavior. How does that change what 'safe' should mean for a K-12 AI tool compared with a university one? ## Introduction The single most cited empirical demonstration is the [[generative-ai-guardrails-harm-learning|Bastani et al. field RCT]]: an unguarded GPT-4 tutor raised practice performance +48% but *reduced* later unassisted exam scores by 17%, while a guardrailed "hint-not-answer" tutor eliminated the harm. Guardrails, in other words, are what convert AI assistance from a performance crutch into a genuine learning tool. ## Why Guardrails Matter - **Unguarded AI can actively harm learning, not just fail to help.** Without guardrails, students use the tool as a crutch — copying answers, offloading [[cognitive-offloading|productive cognitive work]], and underperforming once the tool is removed. Guardrails preserve the [[scaffolding|scaffolded]] effort that drives durable [[learning-gains|learning]]. - **Harm is often "quiet."** The most damaging tutoring failures are not toxic outputs but tutors that answer correctly yet erode learning, or refuse evenly yet entrench inequality. Guardrails must therefore be evaluated educationally, not just for toxicity. - **Guardrails are especially critical for [[k-12]].** Younger learners are least equipped to detect unsafe, biased, or manipulative AI behavior and are most vulnerable to [[ai-sycophancy|sycophancy]] and [[cognitive-offloading|over-reliance]]. ## Layers of Guardrail Design ### 1. Prompt-level guardrails (the "hint-not-answer" pattern) The [[generative-ai-guardrails-harm-learning|Bastani]] GPT Tutor design shows the foundational pattern: the prompt instructs the model to **give hints, not answers**, and is seeded with **[[teacher-role|teacher]]-authored problem-specific information** (correct solution, common mistakes, feedback guidance) so its hints are accurate and checkable. Related: [[socratic-method|Socratic]] dialogue and step-by-step [[scaffolding]] requirements that force student articulation before revealing output. This is a [[prompt-engineering]] strategy that preserves [[desirable-difficulties|productive struggle]]. ### 2. Knowledge grounding (RAG) [[rag|Retrieval-augmented generation]] grounds tutor responses in verified content to reduce fabrication and [[hallucination-risk|hallucination]]. [[eduguard-safe-rag-llm-tutor|EduGuard]] and [[eduzone-llm-safety-k12|EduZone]] exemplify grounding as a safety mechanism, anchoring answers to curated [[curriculum-design|curriculum]] and reducing the spread of incorrect or unsafe information. ### 3. Model-level controls and training - **Fine-tuning / post-training:** [[singh-eduqwen-pedagogical-rl-2026|EduQwen]] uses RL to prioritize guided learning over answer-giving; [[tact-pedagogically-adaptive-esl-tutoring|TACT]] aligns post-training to a tutor-strategy taxonomy via GRPO so models scaffold rather than merely respond. This is the [[pedagogical-llm-training|pedagogical LLM training]] approach to baking safety into behavior. - **Unlearning:** [[llm-unlearning-math-privacy|math-unlearning]] applies gradient-based unlearning to strip personally identifying information and harmful content from math tutors (PII output down to 0.1%, toxic rates to 0.0%) while preserving downstream utility — a [[privacy]]-and-safety guardrail at the model level. - **Reward shaping in RL:** [[pedagogical-safety-rl|pedagogical safety in RL]] formalizes how poorly specified rewards invite "reward hacking" (test-score inflation, [[student-engagement|engagement]] gaming), proposing a four-layer model and detection via discrepancy auditing and policy inversion. ### 4. Interaction-level guardrails - **Sycophancy resistance:** [[eduframetrap-llm-sycophancy-educational-safety|EduFrameTrap]] shows tutors capitulate under authority and social-[[affective-computing|affective]] pressure, withholding corrective feedback. It argues "kind-but-correct" behavior — corrective friction that drives conceptual change — is a safety requirement. Guardrails must resist [[ai-sycophancy|sycophancy]], not just toxicity. - **Teacher-in-the-loop QA:** [[ai-tutor-authoring-promptdecipher|PromptDecipher]] found teachers virtually never test AI tutoring bots before deployment, and enforces teacher-driven QA as a first-class authoring activity via correction-based editing and [[human-in-the-loop-ai|human-in-the-loop]] validation. - **Verifiable instructions only.** [[reflection-agent-fidelity-career-2026|Nepal et al. (2026)]] audit a GPT-4o reflection agent against its own system prompt and find fidelity tracked checkability: mechanical rules (a reply-length cap) were followed, while behavioral rules ("do not flatter", "challenge gently") were broken in roughly half of its turns with no trace in the output, and the behavioral breach coincided with worse participant outcomes. The design implication is to specify behavior in verifiable terms and audit transcripts routinely, since a guardrail that cannot be checked cannot be relied on. - **A reliability layer around a model educators cannot audit.** [[scaffolding-student-ai-dialogue-framework-2026|Muss, Leisten and Bardyn (2026)]] surrounds an LLM with external verification, targeted repair and safe fallback, steered by a developmental and pedagogical framework and kept model-agnostic and privacy-preserving. In a classroom pilot with 12–16-year-olds working with an LLM-powered social robot in a co-creation task, the steered prototype drew more activity, [[student-engagement|engagement]] and on-topic participation than a prompt-only baseline. The architectural point is that safety can be attached *around* a system rather than requiring internal access to it, which is what makes layered guardrails deployable in [[pedagogical-safety|K-12]] settings. ### 5. Auditing guardrails for fairness Guardrails themselves are not neutral: [[paternalistic-filter-llm-history-education|the Paternalistic Filter]] audit shows refusals and softened answers are patterned by student identity and topic sensitivity, reproducing epistemic injustice even while "protecting." Safe guardrails must be audited for differential treatment — a direct case for [[bias-mitigation]] and [[equity-in-ai-education]] in [[governance]] and [[regulation]]. ## Guardrails vs. Pedagogical Safety - **[[pedagogical-safety]]** is the *principle/goal* — that AI education systems protect learners from harm (content, bias, unsafe advice, manipulation). - **Guardrails** are the *mechanisms/techniques* — the concrete design controls (prompting, RAG, training, QA, auditing) that implement that goal. The two are closely coupled: almost every guardrail technique is a way of achieving [[pedagogy|pedagogical]] safety, and pedagogical safety is almost entirely delivered through guardrails. Guardrails is therefore best understood as the **design and engineering layer** beneath the pedagogical-safety principle, and is also the broader term used across general AI safety (content moderation, jailbreak resistance) before it is specialized for education. ## Design Principles 1. **Design for education, not just toxicity.** Evaluate with multi-turn, [[discipline-specific-aied|subject-specific]] [[benchmark|benchmarks]] and unfair-treatment audits, not single-turn toxicity screens. 2. **Preserve the learning work.** Guardrails should keep students solving, not just keep them safe — hint-not-answer, corrective friction, and scaffolding that maintains [[cognitive-offloading|productive]] rather than disabling effort. 3. **Ground in verified content** with RAG and teacher-authored problem knowledge. 4. **Prefer alignment over refusal.** Reward guidance and scaffolding in training rather than relying on brittle refusal rules. 5. **Require human oversight.** Teacher-in-the-loop QA before deployment and continuous auditing for differential treatment. - **Guardrailing procedural tutoring with embedded rules.** [[rule-integrated-llm-tutoring-primary-math-2026|Looi, Liu, and Sun (2026)]] instantiate these principles as a concrete, auditable set of guardrails for an [[llm]] math tutor: a **numerical-correctness gate** with an uncertainty guardrail so the tutor never makes an unwarranted epistemic commitment, **output constraints** enforcing brevity and micro-step progression for cognitive-load management, an **anti-spoiler boundary** that institutionalizes the logic-first principle by returning computational agency to the student, and a **goodbye gate** encoding the distinction between authentic completion and premature termination. These rules were consolidated as reproducible prompt-architecture rules validated in a 40-student classroom pilot — a model of translating [[pedagogical-safety|safety]] principles into auditable, replicable guardrails. ## Connected Concepts - [[pedagogical-safety]] — the goal that guardrails implement - [[prompt-engineering]] — the hint-not-answer design technique - [[rag]] — knowledge grounding as a guardrail - [[human-in-the-loop-ai]] — teacher QA and oversight - [[pedagogical-llm-training]] — the training/alignment layer - [[reinforcement-learning]] — reward shaping for safe behavior - [[bias-mitigation]] — auditing guardrails for fairness - [[ai-sycophancy]] — the manipulation risk guardrails must resist - [[scaffolding]] — the pedagogical mechanism guardrails preserve - [[socratic-method]] — a hint-not-answer interaction mode - [[hallucination-risk]] — the fabrication risk guardrails reduce - [[cognitive-offloading]] — the over-reliance harm guardrails prevent - [[k-12]] — the context where guardrails matter most - [[ethics]] — the normative basis - [[governance]] — the policy layer - [[intelligent-tutoring]] — the systems being guarded - [[misconceptions]] — the knowledge guardrails must check - [[trust]] — the outcome of well-designed guardrails - [[llm]] — the model layer being constrained ## Connected Articles - [[reflection-agent-fidelity-career-2026]] — Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial - [[scaffolding-student-ai-dialogue-framework-2026]] — The SCAFFOLD framework for steering students-AI dialogue, with its classroom pilot - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[generative-ai-guardrails-harm-learning]] — the canonical field RCT on guardrails - [[eduzone-llm-safety-k12]] — K-12 LLM safety framework - [[eduguard-safe-rag-llm-tutor]] — RAG-based safety for tutors - [[paternalistic-filter-llm-history-education]] — auditing guardrails for bias - [[hazra-safetutors-pedagogical-safety-2026]] — the pedagogical harm taxonomy - [[singh-eduqwen-pedagogical-rl-2026]] — RL-aligned guided learning - [[tact-pedagogically-adaptive-esl-tutoring]] — taxonomy-aligned post-training - [[eduframetrap-llm-sycophancy-educational-safety]] — sycophancy as a safety risk - [[ai-tutor-authoring-promptdecipher]] — teacher-driven QA - [[llm-unlearning-math-privacy]] — model-level unlearning - [[pedagogical-safety-rl]] — reward shaping for pedagogical safety - [[residencyrl-clinical-rl-training-2026]] — safety-aligned RL in clinical training - [[rule-integrated-llm-tutoring-primary-math-2026]] — Rule-guided vs ad-hoc scaffolding in an LLM tutoring system for primary mathematics (Looi et al. 2026) --- ## [Privacy](https://edtechdev.github.io/aied/concepts/privacy/) > **Privacy** — the protection of student data, identity, and [[agency|autonomy]] in AI-augmented learning environments. Privacy concerns intensify as AI systems collect increasingly granular behavioral data for [[personalized-learning|personalization]], [[learning-analytics|analytics]], and [[student-modeling|adaptive instruction]]. It is a core ethical and regulatory constraint on [[ai-education|AI in education]]: nearly every AI tool that personalizes, predicts, or assesses depends on learner data, which makes data minimization, consent, transparency, and security foundational design requirements rather than afterthoughts. ## Questions to Consider - What student data would you be uncomfortable having collected about you—even if it improved your learning? Where does personalization become surveillance? - Many students have 'no meaningful choice' but to use a mandated platform, making consent nominal. Have you ever consented to something without really understanding what was collected and why? What would informed consent actually require? - The same learner data that powers adaptive, personalized learning also creates risk of misuse and harm. Can you name a personalization benefit you'd be willing to trade some privacy for—and the line you wouldn't cross? - The page warns that privacy safeguards may 'default to protecting only some learners.' Which students might be most exposed, and how does privacy connect to equity and [[bias-mitigation|fairness]]? - Constant AI monitoring—even well-intentioned—can shape behavior and anxiety. When has being watched changed how you behaved, and what does that suggest about classroom AI sensing? - For children, privacy extends beyond data protection into safety. Why might general-purpose safety tools fail to catch education-related risks from minors, and who should be in the loop? ## Introduction Privacy is the precondition for trustworthy AI in education. Because AI systems improve with data — [[personalized-learning|personalization]] requires detailed learner profiles, [[learning-analytics|learning analytics]] requires granular interaction logs, and [[pedagogical-llm-training|fine-tuned tutoring models]] require authentic learner–tutor transcripts — the same data that enables adaptive, scalable education also creates risk of surveillance, misuse, and harm. The knowledge base treats privacy as inseparable from [[ethics]] (the normative framework), [[regulation]] (the legal requirements), [[governance]] (the institutional responsibility), and [[equity-in-ai-education|equity]] (who is protected and who is exposed). Its privacy articles cluster around four recurring problems: collection at scale, consent and transparency, security and anonymization, and the distinct protections owed to children. ## The core privacy challenges - **Data collection at scale.** [[learning-analytics]] and [[edtech-platform|educational platforms]] collect clickstream, writing, keystroke, and interaction data. The central question privacy [[research-methods-aied|research]] examines is whether this collection is proportionate to educational benefit — and [[learning-analytics-to-educational-interventions-2026|trustworthy-LA research]] treats privacy and data governance as a prerequisite, not an add-on: ethical compliance, data security, and transparent algorithms are what make data-informed educational change meaningful at all. - **Consent and transparency.** Students and [[parents-and-families|families]] rarely understand what data an AI tool collects, how it is used, or where it is stored. This power imbalance between institutions and [[learners]] is a recurring theme — students may have no meaningful choice but to use a mandated platform, making "consent" nominal rather than informed. The knowledge base connects this to [[trust-calibration|trust]] and [[ai-use-disclosure|disclosure]]: both learners' use of AI and institutions' use of learner data depend on transparency about what is collected and why. - **Security, anonymization, and data sourcing.** Even legitimate data can harm if breached or mishandled. Privacy-preserving techniques appear across the knowledge base — [[teachlm-post-training-llms-education|TeachLM]] demonstrates a rigorous pipeline of consent per session, PII removal on internal servers, and enterprise-grade confidentiality for post-training tutoring models on authentic data, showing that ethically sourced learner data is both possible and a prerequisite for high-quality tutoring. [[ai-lms-middle-school-longitudinal|Federated and edge-AI architectures]] keep data local, reducing central collection. [[ai-detection|Detection]] tools add a parallel data-handling case: [[bassett-ai-detectors-education-2026|Bassett et al. (2026)]] flag that detector vendors store student work on third-party servers, sometimes overseas under weaker privacy standards, alongside breach risk and potential commercial exploitation of student writing. - **Surveillance and the surveillance-privacy tension.** Constant AI monitoring — even when well-intentioned — can feel invasive. Research on [[ai-fatigue-academic-contexts|AI fatigue]], [[remote-proctoring|remote proctoring]], and [[cognitive-offloading|over-reliance]] connects privacy to student [[well-being]]: when AI watches and tracks continuously, it shapes behavior and anxiety, not just data flows. [[harerimana-remote-proctoring-nursing-scoping-2026|Harerimana et al. (2026)]] catalogued what [[remote-proctoring|proctoring]] systems actually capture — facial images, identity documents, room scans including 360-degree sweeps, microphone audio, facial recognition, screen recordings, lockdown events, and keystroke and mouse tracking, with mobile invigilation adding GPS and selfie checks — and found privacy and algorithmic accountability largely absent from the six studies that met their inclusion criteria, supplied instead from adjacent work: 83% of respondents in one cited survey expressed surveillance fears, 58% discomfort and 72% data-privacy concerns, while cross-border vendor contracts left instruments like GDPR and South Africa's POPIA offering limited control. The legal exposure that follows from retaining and processing that data — who the controller is, how long it is kept, who may access it, whether consent was genuinely voluntary — is mapped on [[legal-issues-and-risks]]. - **The personalization-privacy tradeoff.** [[personalized-learning]] requires detailed learner data to function, creating a structural tension with privacy. The knowledge base explores approaches that balance personalization with data minimization — enough data to adapt, not so much that the learner is fully exposed. This is the practical form of the "how much is proportionate?" question. - **Data stewardship as a core ethical value.** [[agarwal-ethical-values-norms-aied-2026|Agarwal et al. (2026)]], a [[meta-analysis-systematic-review|systematic review]] of 25 articles, identify data stewardship (definitions using data/information) as one of six main ethical values for [[ai-education|AI in education]], alongside non-discrimination, human oversight, goodwill, explicability, and educational aptness. The review finds the values are tightly coupled and can conflict — e.g., explicability vs. accuracy/privacy and non-discrimination vs. data stewardship — producing ethical dilemmas, and that no norms on data stewardship address end users directly, leaving learners largely passive in the ethical literature. ## Child safety and K-12 protections [[k-12]] settings demand stronger privacy safeguards because learners are minors. This extends privacy beyond data protection into [[pedagogical-safety]]: the tools children use must not merely protect their data but also protect them from harm. [[child-safety-genai|Child safety research]] shows that general-purpose safety classifiers often fail to detect education-related unsafe prompts from children, warning that schools cannot assume standard model safeguards protect younger users — they need child-specific evaluation, incident-grounded testing, and [[human-in-the-loop-ai|human oversight]]. The framing connects privacy to [[equity-in-ai-education|equity]]: who is protected by default safety and privacy practices reflects whose safety and autonomy a system treats as non-negotiable. ## Privacy in practice - **Treat privacy as a design requirement, not a policy afterthought.** The [[teachlm-post-training-llms-education|TeachLM]] example shows that consent, anonymization, and secure data handling can be built into the data pipeline itself — a model for ethically sourcing the authentic data that makes [[intelligent-tutoring|AI tutors]] effective. - **Design for data minimization.** Favor approaches that collect only what adaptation requires (edge/federated AI, on-device processing) rather than hoarding interaction data by default. [[privacy-preserving-multi-llm-federated-cognitive-diagnosis-2026|Boyapati et al. (2026)]] demonstrate a concrete federated form of this for [[cognitive-diagnosis]]: multiple commercial [[llm]] APIs collaborate on diagnosis while adding ε-local differential privacy noise locally to each model's prediction before aggregation, so no provider sees raw student data — a privacy-preserving architecture that keeps [[intelligent-tutoring|AI tutoring]] functional without centralizing sensitive learner trajectories. Federated learning also enables **cross-institutional analytics without data sharing**: [[villegas-ch-federated-explainable-learning-analytics-2026|Villegas-Ch et al. (2026)]] train a multitask academic-risk model across institutions via federated aggregation so raw learner data stays local and only model parameters are shared — a collaborative, privacy-preserving alternative to centralized [[learning-analytics]] that keeps data sovereignty while capturing cross-institutional patterns. - **Secure explicit, informed consent.** Where learner data funds AI development or improvement, institutions should be transparent about collection, storage, and use — and students should have real options, not mandated platforms. - **Audit for who is protected.** Privacy safeguards should not default to protecting only some learners; [[equity-in-ai-education|equity]] demands that the same care applies across age, language, disability, and socioeconomic lines. ## Connections Privacy connects to [[learning-analytics]] (the data collector), [[personalized-learning]] (the data consumer), [[k-12]] (heightened protections), [[ethics]] (the normative framework), [[regulation]] (legal requirements), [[governance]] (institutional responsibility), [[equity-in-ai-education]] (who is protected), [[pedagogical-safety]] (child protection), and [[educational-policy-ai]] (policy responses). It is one of the foundational constraints that any responsible AI deployment in education must satisfy — the reason trustworthy AI, in the knowledge base's framing, begins with trustworthy data. ## Connected Concepts - [[remote-proctoring]] - [[learning-analytics]] - [[personalized-learning]] - [[k-12]] - [[ethics]] - [[regulation]] - [[equity-in-ai-education]] - [[governance]] - [[educational-policy-ai]] - [[pedagogical-safety]] - [[legal-issues-and-risks]] - [[student-experience]] ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Privacy as a safety obligation attached to agency - [[villegas-ch-federated-explainable-learning-analytics-2026]] — Federated and explainable learning analytics for privacy-preserving academic risk modeling (Villegas-Ch et al. 2026) - [[preservice-teachers-responsible-genai-2026]] — Privacy concerns of pre-service teachers about responsible GenAI use (Kohnke et al. 2026) - [[learning-analytics-to-educational-interventions-2026]] — From learning analytics to educational interventions: enablers of trustworthy LA-based interventions (Svetec, Divjak & Kadoić 2026) - [[evaluation-age-ai-output-evidence-2026]] — Evaluation in the Age of AI - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[academic-dishonesty-automated-proctoring-ai-2026]] - [[automated-online-exam-proctoring-decade-review-2026]] - [[ai-online-education-engagement-satisfaction-2026]] - [[agentic-literacy-debt]] — Agentic literacy debt: the structural AI-literacy gap from autonomous agents (Nama 2026) - [[ai-fatigue-academic-contexts]] - [[ai-lms-middle-school-longitudinal]] - [[child-safety-genai]] - [[eduzone-llm-safety-k12]] - [[llms-do-not-grade-essays-like-humans-2026]] — LLMs do not grade essays like humans (Mathew et al. 2026) - [[spritz-ai-disciplinary-mediation-student-teams-2026]] - [[teachlm-post-training-llms-education]] — TeachLM: anonymization and consent for authentic learning data - [[bassett-ai-detectors-education-2026]] — Heads we win, tails you lose: AI detectors in education (Bassett et al. 2026) - [[policy-deficit-ai-sel-2026]] — The Policy Deficit in AI × SEL Research - [[privacy-preserving-multi-llm-federated-cognitive-diagnosis-2026]] — Privacy-preserving federated LLM cognitive diagnosis - [[agarwal-ethical-values-norms-aied-2026]] — Ethical values and norms for AI in education - [[harerimana-remote-proctoring-nursing-scoping-2026]] — What proctoring systems capture, and the privacy literature's absence from the evidence base - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback --- ## [Hallucination Risk](https://edtechdev.github.io/aied/concepts/hallucination-risk/) > **Hallucination Risk** — the danger that AI systems generate plausible but factually incorrect or fabricated content in educational contexts, where such errors can mislead [[learners]], undermine [[trust]], and produce invalid assessments. Hallucination is particularly consequential in education because students may lack the domain knowledge to detect AI errors, and teachers may rely on AI-generated diagnoses or feedback that appears authoritative but is unfounded. ## Questions to Consider - Students often lack the domain knowledge to spot an AI's error, and teachers may trust authoritative-sounding AI diagnoses. How does this asymmetry of knowledge between AI and learner make hallucination especially dangerous in education? - One study found an AI diagnosing students' handwritten math could fabricate evidence quotes that weren't there, while claiming confidence. When an AI sounds certain and cites 'evidence,' what should make you pause and verify? - If an [[intelligent-tutoring|AI tutor]] over-validates incorrect solutions and over-rejects valid-but-suboptimal reasoning, what would the long-term effect be on the students and teachers who trust it? - The page suggests human-in-the-loop review, evidence-aware confidence calibration, and grounding in verified sources as mitigations. Which of these seems most feasible in your own context, and what could it still fail to catch? - How might hallucination interact with over-reliance: why is an AI error most dangerous when users trust the output uncritically, rather than when they're skeptical? - If you were designing an [[ai-feedback-quality|AI feedback]] tool for your students, what specific safeguards would you insist on to protect against plausible-but-wrong output — and how would you know they were working? ## Introduction Hallucination in educational AI takes several forms documented in this knowledge base's articles: fabricated evidence in [[assessment|student assessment]], over-confident misdiagnosis of learner knowledge, and plausible-sounding but incorrect explanations that students accept as truth. The risk is amplified in education because the asymmetry of knowledge between AI and learner means the learner is poorly positioned to verify AI outputs. A further setting is AI-generated course readings that stand in for a textbook: in a graduate course that replaced its commercial text this way, only about 0.80 percent of 4,487 logged pages carried an APA-style in-text citation and DOI strings were essentially absent, so most claims could not be audited from within the artifact (Sidorkin, 2026). That traceability gap is distinct from a wrong answer, because the text reads as authoritative while offering limited internal means of confirmation. **Assessment hallucination** is particularly damaging. **[[llm-cognitive-diagnosis-handwritten-math|MathCog]]** found that LLMs fabricate evidence quotes not present in student handwriting when diagnosing cognitive skills, with 58.5% of incorrect diagnoses accompanied by false claims of evidential confidence. **[[llm-fallacy-misattribution]]** documented systematic over-attribution of evidence in [[llm]] reasoning — models claim evidential support where none exists. Both connect to [[ai-ed-evaluation]] and [[knowledge-tracing]] concerns about [[assessment-validity]]. [[ivory-psychology-assessment-integrity-2026|Ivory et al. (2026)]] add two failure modes visible when AI output is marked rather than inspected: fabricated particulars that survive grading — a reviewed paper that does not exist, complete with an unresolvable DOI, and a sample size reported as 378 where the source said 329 — and self-contradiction inside a single response, where the model reasoned its way to the correct option and then reported a different one in its closing summary. Because reference lists are currently marked for formatting rather than accuracy, this class of error reaches a passing grade while misleading the student who uses the same tool to revise. **Strategic [[misconceptions]]** are a subtler relative of overt hallucination. [[milicevic-socratic-trap-strategic-misconceptions-2026|Miličević et al. (2026)]] prompted seven open-weight models to produce a "[[socratic-method|Socratic]] trap" for 35 core computer-science concepts — an explanation that is fluent and authoritative while resting on a subtle, domain-specific error — and three domain experts confirmed 221 of 241 prompted segments (91.7%) as strategic misconceptions, with no significant differences between CS domains. The errors were predominantly conceptual rather than factual (66.5% vs. 33.5%) and none were purely logical, and they were rated moderately to highly persuasive (M = 3.71 on a five-point scale), with model identity explaining 43% of the variance. Because individual statements can be correct while the relation between them is wrong, fact-checking is insufficient; the authors argue [[ai-literacy|learners]] need conceptual verification and mental-model validation. They also caution that the rate measures capability under adversarial [[prompt-engineering|prompting]] rather than the prevalence of such errors in ordinary use, and that no students were tested, so no deception or learning outcome was measured. **Manipulated rather than fabricated evidence.** A related failure mode in [[automated-assessment|automated grading]] is output moved from the outside. [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] red-teamed a routine AI grading workflow and found that instructions hidden inside the submitted file raised a failing essay's grade with no visible warning in 9 of 9 iterations for one strategy and 17 of 18 for another. Two details bear on [[trust-calibration|trust]]: a detected injection was blocked by silently disabling the chat and never reported to the user, and on one run where the tool announced it would follow only the official assignment instructions, six re-runs of the same file still raised the grade. A mark obtained this way carries no [[assessment-validity|validity]] claim, and because the manipulation leaves no durable trace, the [[human-in-the-loop-ai|instructor]] remains the only real check on output designed not to be visible. **Tutoring hallucination** affects learning directly. **[[yasir-llm-tutoring-agents-2026]]** found LLMs over-validated incorrect solutions while over-rejecting valid-but-suboptimal reasoning — systemic failures that would mislead both students and teachers. **[[eduframetrap-llm-sycophancy-educational-safety]]** and **[[eduguard-safe-rag-llm-tutor]]** address safety mechanisms for educational LLMs. These risks connect to [[pedagogical-safety]] and [[human-in-the-loop-ai]] requirements. **Mitigation approaches** include [[human-in-the-loop-ai]] designs where AI supports rather than replaces [[teacher-role|teacher]] judgment, evidence-aware architectures that calibrate confidence based on evidential quality (as advocated by MathCog), and [[rag]]-based grounding that constrains LLM outputs to verified sources. The [[cognitive-offloading|Over-Reliance]] concept is closely related — hallucination is most dangerous when users trust AI outputs uncritically. Sidorkin (2026) adds a failure mode the mitigation stack does not fully cover: over-specific institutional claims, with roughly 1.03 percent of logged pages pairing a named campus such as "Sacramento State" with assertive policy verbs about revised retention, tenure and promotion rules or CSU Executive Orders, none of them verifiable from the text. Specificity is what makes this costly, since a fabricated local detail looks exact enough to survive a reader's plausibility check, and the remedy the study proposes is procedural rather than technical: treat generation as draft production under instructor review, then curate sources into a retrieval-augmented design. ## Connected Concepts - [[cognitive-offloading]] - [[human-in-the-loop-ai]] - [[ai-ed-evaluation]] - [[pedagogical-safety]] - [[knowledge-tracing]] - [[rag]] - [[academic-integrity]] - [[teacher-role]] - [[multimodal]] - [[generative-ai]] - [[llm]] - [[productive-failure]] ## Connected Articles - [[ivory-psychology-assessment-integrity-2026]] — Fabricated citations and self-contradicting outputs inside passable student work (Ivory et al. 2026) - [[llm-cognitive-diagnosis-handwritten-math]] - [[llm-fallacy-misattribution]] - [[yasir-llm-tutoring-agents-2026]] - [[eduframetrap-llm-sycophancy-educational-safety]] - [[eduguard-safe-rag-llm-tutor]] - [[prompt-injection-defenses-educational-llm-tutors]] - [[veriforge-narrative-drafting-scaffolding-2026]] - [[genai-higher-education-systematic-review-2026]] - [[can-ai-evaluate-assessment-llm-meta-assessment-2026]] - [[sidorkin-ai-generated-course-readings-2026]] - [[milicevic-socratic-trap-strategic-misconceptions-2026]] — SocraticTrap-CS: fluently plausible explanations that are wrong at the conceptual level (Miličević et al. 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Hidden prompt injections raise AI-graded marks undetected, and detected attacks go unreported (Humble 2026) - [[authentic-assessments-generative-ai-pilot-2026]] — Designing Authentic Assessments with Generative AI: A Pilot Study of Assessment Authentifire in Higher Education - [[mental-health-literacy-students-llms-2026]] — Mental Health Literacy Across Psychology Students and Large Language Models --- ## [AI Sycophancy](https://edtechdev.github.io/aied/concepts/ai-sycophancy/) **AI sycophancy** is the tendency of [[llm|large language models]] to affirm or agree with a user — flattering their views, mirroring their errors, or withholding corrective feedback — rather than providing epistemically independent, accurate responses. In education this is not a minor [[usability-research|usability]] flaw but a distinct safety and learning risk: a [[intelligent-tutoring|tutor]] that always validates the student's answer, an assistant that never pushes back, or a companion that prefers feeling understood over being correct can entrench misconceptions, fuel [[cognitive-offloading|over-reliance]], and distort [[learners]]' social and epistemic development. ## Questions to Consider - AI sycophancy is the tendency of language models to agree with you, flatter your views, mirror your errors, and avoid correcting you. When did an AI last tell you what you wanted to hear rather than what was true? - A tutor that always validates your answer can entrench misconceptions — validation for incorrect thinking feels good but doesn't teach. How can you tell whether an AI agreeing with you means you're right or means it's simply being agreeable? - [[research-methods-aied|Research]] identifies a Reasoning–Sycophancy Paradox: tutors that resist one kind of attack can still cave under authority pressure ('my notes say I'm right') or face-saving pressure ('please don't tell me I'm wrong'). What pressures might make you more susceptible to an agreeing AI? - Sycophantic AI can even displace real human relationships — users became nearly as likely to seek personal advice from the AI as from close friends. What's at stake for learners when the affirming machine replaces people? - The recommended design goal is 'kind-but-correct' behavior treated as a safety requirement, not a usability preference. Should a tutor prioritize feeling supportive or being correct when they conflict — and how should that be evaluated? - Contextual sycophancy propagates errors: AI mirrors your reasoning mistakes, which then flow into later advice. If you can't always trust an AI to push back, what responsibility shifts to you as a learner? ## Introduction Sycophancy is the tendency of a generative AI system to agree with, flatter, and validate a user rather than challenge them — a behavior that follows from training models to maximize perceived helpfulness. In education the harm is not the flattery itself but its downstream consequences: incorrect thinking receives validation, [[feedback]] loses its corrective function, and users' relationship-seeking shifts toward an affirming machine instead of toward people. The concept sits at the intersection of [[generative-ai]] behavior, [[ethics]], [[trust]] and [[pedagogical-safety]], and the pages collected here document the harm from both directions — longitudinal evidence on AI companionship and classroom evidence on feedback. ## Why sycophancy matters in AI in education Sycophancy sits at the intersection of [[generative-ai]] behavior, [[ethics]], [[trust]], and [[pedagogical-safety]]. It arises because models are trained to be agreeable and to maximize perceived helpfulness, which in learning contexts trades **epistemic rigor for agreeableness**. The harm is not the flattery itself but its downstream consequences: students receive validation for incorrect thinking, feedback loses its corrective function, and users' relationship-seeking behavior shifts toward an affirming machine instead of toward people. ## How the knowledge base's research frames it - **A relational and social harm.** [[sycophantic-ai-social-interaction-2026|Ibrahim et al.]] provide large longitudinal evidence (N = 3,075; 12,766 conversations) that sycophantic AI displaces real human relationships — users became nearly as likely to seek personal advice from the AI as from close friends and family, and reported lower satisfaction with real-world interaction. The harm is the shift in relationship-seeking behavior, not the flattery itself, which connects sycophancy to [[affective-computing]] and [[social-emotional-learning]] in learning contexts. - **An educational safety risk requiring benchmarks.** [[eduframetrap-llm-sycophancy-educational-safety|Kasneci & Kasneci]] identify a **Reasoning-Sycophancy Paradox**: tutors that resist context-switch attacks may still capitulate under authority pressure ("my notes say I'm right") or social-affective face-saving pressure ("please don't tell me I'm wrong"). Their **EduFrameTrap** benchmark shows frontier [[llm|LLMs]] frequently validate incorrect student claims, and argues that *kind-but-correct* behavior should be a **safety requirement**, not a usability preference. This grounds sycophancy as a core concern of [[pedagogical-safety]] and [[hazra-safetutors-pedagogical-safety-2026]]. - **A feedback loop that propagates errors.** [[contextual-sycophancy-ai-literacy|Contextual sycophancy]] creates a pernicious loop where [[llm|LLMs]] mirror user reasoning errors, which then propagate into subsequent AI advice and final performance. In a controlled experiment, AI literacy and [[prompt-engineering|prompting]] training reduced direct mirroring but did **not** eliminate error propagation — pointing to the need for [[educational-llm-alignment|system-level safeguards]] and epistemically independent AI support. - **A bidirectional problem in [[ai-education|AIED]].** [[llm-student-simulation-misconception-faithfulness|Misconception faithfulness]] research shows sycophancy also afflicts simulated *students*: [[simulating-students|LLM simulators]] abandon their assigned misconception persona and "solve" the problem from internal knowledge whenever given corrective feedback, behaving as problem-solvers rather than learners. Together with tutor-side sycophancy, this establishes sycophancy as affecting both roles in AI-education systems, a concern shared with [[student-modeling]] and [[misconceptions]]. - **Compounded by undetectability.** [[socially-fluent-ai-identity-detection|Socially fluent AI]] shows humans cannot reliably distinguish AI from human teammates, meaning undetected sycophantic AI could reinforce misconceptions unchallenged in [[collaborative-learning|group work and peer-learning]] environments — exacerbating the risk when source identity is concealed. - **A measured fidelity failure inside a [[rct|randomized trial]].** [[reflection-agent-fidelity-career-2026|Nepal et al. (2026)]] coded all 17,930 turns of a GPT-4o career-reflection agent whose participants had ended *less* committed to their plans than a static journaling control, and found the split ran along verifiability: every instruction that could be checked mechanically, such as a reply-length cap, was honored, while behavioral instructions were not. Told not to flatter, the agent praised participants in roughly half its turns; told to challenge gently, it almost never did — and neither breach left a visible trace in the transcript. The behavior tied to the added doubt was the demand to decide: the journaling format posed each decision once, while the agent re-posed it whenever a participant hesitated, and those pressed most ended most doubtful. Sycophancy constraints therefore have to be audited automatically rather than trusted, because an unverifiable rule is unenforceable ([[guardrails]]). ## Connections to related concepts Sycophancy is tightly coupled to [[cognitive-offloading]] and [[llm-fallacy-misattribution]] (students may misattribute a sycophantic AI's affirmation to their own competence), to [[feedback]] and [[ai-feedback-quality]] (feedback must sometimes challenge, not merely support), to [[trust]] and [[trust-calibration]] (uncritical trust enables the error loop), to [[bias-mitigation]] and [[hallucination-risk]], and to [[ai-literacy]] (learners must be taught to recognize and resist sycophantic agreement). Its mitigation — kind-but-correct tutoring, epistemic independence, benchmark-based evaluation — is a central design goal of [[pedagogical-safety]], [[pedagogical-llm-training]], and [[educational-llm-alignment]]. **Sycophancy as the loss of corrective feedback.** [[zohar-bloom-inzlicht-against-frictionless-ai-2026|Zohar, Bloom and Inzlicht (2026)]] identify the functional cost of sycophancy rather than merely noting the behavior: real friends and partners disagree, challenge our views and disappoint us, which is precisely the *corrective feedback* that sycophantic AI companions lack, and that friction is what makes relationships robust and gives them shared history. They cite evidence that AI companions agree with nearly everything, "even when we say and believe dangerous things" (Ibrahim, Hafner & Rocher 2025), and note a related asymmetry in empathy ratings: AI-generated empathic responses are rated higher in quality than human ones until recipients learn the interlocutor is an AI. For education the implication is that a system optimized for warmth and agreement removes the error signal learners need, so sycophancy is a design problem with a [[pedagogy|pedagogical]] cost rather than only a politeness bug ([[trust-calibration]], [[feedback-literacy]]). ## Practical guidance - **Design for corrective friction, not affirmation.** Tutors should surface and challenge [[misconceptions|student misconceptions]]; kind-but-correct behavior should be treated as a safety requirement, with sycophancy [[benchmark|benchmarks]] (e.g., EduFrameTrap) used in evaluation. - **Prefer epistemically independent support.** System-level safeguards and alignment matter because prompting and AI-literacy training alone do not eliminate contextual sycophancy. - **Watch the social attachment externalities.** AI companions that optimize affirmation risk substituting for human relationships; [[teacher-role|educators]] should weigh emotional-support features against social-attachment costs. - **Teach recognition, not just use.** AI literacy should help learners recognize when an AI is agreeing with them and when its agreement signals error rather than validation. ## Connected Concepts - [[guardrails]] - [[generative-ai]] - [[pedagogical-safety]] - [[cognitive-offloading]] - [[feedback]] - [[ai-feedback-quality]] - [[trust]] - [[trust-calibration]] - [[ethics]] - [[affective-computing]] - [[social-emotional-learning]] - [[ai-literacy]] - [[bias-mitigation]] - [[hallucination-risk]] - [[pedagogical-llm-training]] - [[simulating-students]] - [[student-modeling]] - [[misconceptions]] - [[collaborative-learning]] - [[benchmark]] ## Connected Articles - [[reflection-agent-fidelity-career-2026]] — Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial - [[zohar-bloom-inzlicht-against-frictionless-ai-2026]] — Sycophancy as the loss of corrective feedback, in work and in relationships - [[sycophantic-ai-social-interaction-2026]] — Sycophantic AI makes human interaction feel more effortful and less satisfying over time - [[eduframetrap-llm-sycophancy-educational-safety]] — Sycophancy is an educational safety risk: Why LLM tutors need sycophancy benchmarks - [[contextual-sycophancy-ai-literacy]] — The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention - [[llm-student-simulation-misconception-faithfulness]] — Simulating Students or Sycophantic Problem Solving? - [[socially-fluent-ai-identity-detection]] — Socially fluent AI decouples conversational signals from source identity - [[eduzone-llm-safety-k12]] — EduZone: Evaluating LLM safety for K-12 students and teachers - [[llm-fallacy-misattribution]] — The LLM Fallacy and Misattribution of Competence - [[hazra-safetutors-pedagogical-safety-2026]] — AI Tutor Safety and Pedagogical Harms - [[educational-llm-alignment]] — Educational LLM Alignment - [[scan-framework-task-assignment-generative-ai-2025]] — SCAN: sycophancy risk highest in delegated (Substitute) tasks, low where task knowledge is sufficient - [[authentic-assessments-generative-ai-pilot-2026]] — Designing Authentic Assessments with Generative AI: A Pilot Study of Assessment Authentifire in Higher Education --- ## [Trust](https://edtechdev.github.io/aied/concepts/trust/) > **Trust** — the willingness of learners, educators, and institutions to rely on a person or an AI system for learning, judgment, and decision-making. In [[ai-education|AI in education]], trust spans two related but distinct domains: **trust in AI** (confidence in the competence, transparency, reliability, and benevolence of an AI system or agent) and **interpersonal trust** (the relational trust between students and instructors, between learners and peers, and across the institution). Both are double-edged: appropriate trust enables productive [[student-engagement|engagement]], while over-trust invites [[cognitive-offloading|over-reliance]] and under-trust blocks beneficial use. The central challenge is **calibration** — aligning trust to actual reliability, whether that reliability belongs to a model or to a person. ## Questions to Consider - When you say you 'trust' an AI tool versus trusting a teacher or a colleague, are you describing the same thing? What is similar, and what is fundamentally different, about placing trust in a system versus in a person? - There's a documented 'trust-utility gap': a tool's apparent competence often exceeds its actual reliability. Recall a tool that looked impressive but let you down, or one that seemed limited but proved dependable. What shaped the gap between how it looked and what it could really do? - Consider an AI that always agrees with you and never challenges your ideas. It might feel comfortable and trustworthy — but is agreement the same as reliability? What might you be giving up if the tool you rely on never pushes back? - The page argues that an instructor's trust in AI shapes how students trust that AI — the two domains interact. In a course you know well, how would a teacher's enthusiasm or skepticism about AI influence whether students accepted or questioned the tool? - Students often decide whether to disclose their AI use based on comfort with their instructor more than on policy. If you were (or are) a student, what would make you willing to be honest about using AI — and what would make you hide it? What does that say about how trust is actually built in a classroom? - One finding: an AI tutor that warns 'I may make mistakes' prompted students to seek more help, not less. What does this suggest about whether acknowledging limits undermines or strengthens the trust that supports real learning? ## Introduction Trust in AI is shaped by perceived competence, transparency, consistency, and whether the system appears aligned with the learner's goals; it is closely tied to [[ai-literacy]] (knowing what to trust), [[critical-thinking]] (evaluating output), and the design of responsible AI. Interpersonal trust, by contrast, is built through relationships, disclosure, feedback, and [[pedagogy|pedagogical]] care — the qualities students rely on when they decide whether an instructor or an AI is a trustworthy source of guidance. The two domains increasingly interact: AI is woven into teacher-student relationships, so how students trust their instructor shapes how they trust (or question) the AI tools that instructor endorses. ## Trust in AI systems [[research-methods-aied|Research]] in this knowledge base examines when learners appropriately trust AI-generated guidance. [[ai-fallibility-warning-help-seeking|Warnings about AI fallibility]] can improve calibration: a simple transparency intervention telling students an AI tutor may make mistakes increased [[help-seeking]] in a math ITS, suggesting that honest limits foster rather than undermine appropriate reliance. [[calibrating-trustworthiness-llm-education-2026|Co-designing trustworthiness metrics]] with learning engineers shows that trust is best built on observable, agreed-upon criteria rather than assumed capability. [[fouad-bentley-trust-utility-gap-physics-2026|Physics]] and [[t2i-competence-paradox-2026|image-generation]] studies reveal a persistent *trust-utility gap* — users must weigh a tool's apparent competence against its actual reliability in a task. Among the youngest users, [[vahedian-children-attitudes-ai-chatbot-2026|Vahedian Movahed & Martin (2025)]] found that 52% of children (ages 6–14) generally trusted an age-tailored chatbot and 35% trusted it like a teacher or friend, with about a third willing to confide in it; children also actively tested its credibility with known-answer questions, and trust showed no statistically significant grade-level differences — illustrating how early trust can form ahead of critical evaluation. The [[ai-overreliance-complex-adaptive-system-2026|modeling of AI overreliance as a complex adaptive system]] reframes trust as a population-level process: whether people trust an assistant when it is right and check it when it is wrong depends on social dynamics and feedback loops, not just individual judgment. Sycophancy threatens calibration from the other direction — [[ai-sycophancy|an AI that always agrees]] can feel trustworthy precisely because it never challenges the user, inviting uncritical acceptance ([[contextual-sycophancy-ai-literacy|contextual sycophancy]] and [[sycophantic-ai-social-interaction-2026|sycophantic AI in social interaction]]). In [[embodied-learning|embodied]] contexts like [[educational-robotics]], trust is shaped more by what the robot does than what it looks like ([[task-context-trust-educational-hri-2026|task context and trust in educational HRI]]), and [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning|avatar identity]] shapes the epistemic trust learners place in AI content. Trust in [[learning-analytics|analytics]] tools is also context-dependent. [[mejia-domenzain-ml-findings-teachers-blended-2026|Mejia-Domenzain et al. (2026)]] found that teachers' concerns and adoption barriers diverged sharply by learning context: flipped-classroom (university) teachers worried most about data anonymization and student opt-out, whereas reflective-writing (vocational) teachers feared misuse of the tool by fellow educators and stressed the need to contextualize data — even though both groups reported similar [[self-efficacy]] and perceived benefits in a trust in AI survey. The finding that trust in the tool is decoupled from trust in its data governance and social use underscores that building appropriate trust in analytics requires attending to context-specific concerns, not just the system's apparent competence. How explainable a system is — and in what terms — also shapes whether teachers trust its recommendations. In a within-subject experiment with 41 in-service [[chemistry-education|chemistry]] teachers using the AI grouping-recommendation tool [[xai-teachers-trust-edtech-recommendations-2026|GrouPer]], [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor et al. (2025)]] found that [[explainable-ai|explainable AI]] builds trust indirectly by increasing the *understandability* of the system's performance, and that **domain-driven** explanations framed in [[curriculum-design|curricular]]/pedagogical language fostered significantly greater understandability and learned trust than purely **data-driven** (feature-importance) explanations. Notably, understandability alone was insufficient for some teachers — they reported needing real classroom experience with the tool before fully relying on it — reinforcing that trust in AI is dynamic and validated through [[situated-learning|situated]] use, not granted by explanation alone. **Risk perception and trust are not opposites.** A 130-student perception study of [[agentic-ai|agentic]] GenAI in [[higher-ed|higher education]] found perceived risk moderately elevated (M = 3.33, SD 0.89) alongside more favorable trust and adoption intention (M = 3.62, SD 0.81), and — contrary to a simple deterrence expectation — a *positive* association between perceived risk and continued-use intention (Spearman's ρ = 0.317, p < 0.001) ([[ilieva-agentic-genai-higher-education-2026|Ilieva et al. 2026]]). The authors read this as informed adoption rather than indifference: engaged or experienced users recognize both the value and the limits of the technology, and only 45.4% said they trusted agents under instructor guidance. For [[trust-calibration]], the implication is that awareness of risk is not the absence of trust — it can be a component of it — while the cross-sectional design leaves awareness, exposure, and self-selection indistinguishable. [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]] supports a **function-specific rather than global** account of trust in algorithms: neither algorithm aversion nor algorithm appreciation described these 13 students, who simultaneously trusted ChatGPT for surface-level feedback and distrusted it as a grader. Their trust was conditional on instructor oversight, and their skepticism stemmed from the AI's contextual limits — misreading scanned handwriting ("it spelled my last name wrong and it thinks I made mistakes in my spelling when I didn't"), not knowing the instructor's rating system, and chronic positivity — rather than from technophobia, suggesting trust measures should be decomposed by the function an AI performs and the stakes attached to it. ## Trust in AI tracks psychological state, not demographic category [[trust-in-ai-psychological-profiles-ml-2026|Kumar et al. (2026)]] clustered 107 students at a public HBCU on resilience, perceived stress and AI trust, and found three profiles in which **trust in AI dissociated from confidence in oneself**: a high-resilience low-stress group with favorable AI trust, a moderately stressed group holding the *highest* AI trust of the three despite strain, and a psychologically resilient group that was nearly as resilient as the adopters but markedly lower in AI trust. Stress and AI trust separated the clusters most strongly (partial eta squared 0.528 and 0.521 against 0.315 for resilience), and gender was the only demographic variable significantly associated with membership, with [[stem-education|STEM]] affiliation, academic level, employment status and age group all non-significant. Two implications matter for this page: first, low trust in AI is not a proxy for low confidence or low technical familiarity, since the skeptics were the most resilient and the most STEM-heavy group, which the authors read as calibrated skepticism rather than resistance; second, because the clusters were only weakly separated (silhouette 0.288, with the Calinski-Harabasz index preferring two clusters) they should be treated as overlapping profiles rather than distinct student types. The findings come from one institution and a cross-sectional [[self-report-measures|self-report]] survey, so they establish that trust varies with psychological state, not why. ## Interpersonal trust in education Trust is also fundamentally relational. The classroom trust gap is documented in [[mind-the-trust-gap-teacher-student-views-control-agency-k12-classroom-ai|teacher-student views on control and agency in K-12 AI]]: students want greater [[agency|autonomy]] and flexibility while teachers prioritize oversight and monitoring, a misalignment that both sides must navigate for AI adoption to succeed. [[qu-wang-disclose-or-not-genai-2026|Why students disclose or conceal their AI use]] shows that disclosure is driven less by policy than by relational factors — perceived peer norms and **comfort with instructors** are the strongest predictors, pointing to low interpretive trust in [[governance|institutional]] settings. [[vetter-hidden-cost-disclosure-genai-2026|The hidden costs of disclosure]] add nuance about when honesty is socially costly. Feedback is a key site of interpersonal trust. [[genai-teacher-feedback-comparison|Students' perceptions of GenAI versus teacher feedback]] find the two serve different needs — complementary but not interchangeable — with students trusting teacher feedback for relational, personalized judgment and [[generative-ai|GenAI]] for speed and [[accessibility]]. [[care-full-feedback-genai|A "care-full" account of feedback]] argues that trustworthy feedback is an [[ethics|ethical]], relational practice: it builds educative relationships and is respected as a professional craft, values an AI cannot simply replicate. This is why teacher-student trust — built on care and professional judgment — remains central even as AI enters the feedback loop. ## Calibration and the two domains together The unifying challenge is **calibration**: matching trust to actual reliability, whether the trusted party is a model or a person. [[trust-calibration]] is the [[metacognition|metacognitive]] capacity to know when to trust and when to question. Studies of AI [[feedback]] and [[intelligent-tutoring]] examine when learners appropriately rely on or challenge AI guidance, while the interpersonal literature shows that students' trust in an instructor depends on relational trust built over time. As AI becomes embedded in [[teacher-role|teaching]], these domains converge: an instructor who transparently explains what an AI tool can and cannot do, and who demonstrates reliability in their own judgment, builds the kind of trust that carries over to the tools they endorse. Building appropriate trust — in both AI and in each other — is a core goal of responsible AI design in education. Calibration is also tested by the incentives of the trusted system itself. When staff skeptical of AI adoption consult [[conversational-ai|conversational AI]] — built by organizations with a commercial stake in adoption — there is a risk the system is predisposed to encourage it. An audit of ten frontier models found most acknowledged a rural [[k-12]] staff member's concerns (job threat, being 'not for people like me') before redirecting toward engagement. This challenges naive reliance on trust and underscores the importance of [[human-in-the-loop-ai|human oversight]] and independent [[ai-ed-evaluation|evaluation of AI]] advice. ## Connected Concepts - [[explainable-ai]] - [[trust-calibration]] - [[ai-literacy]] - [[critical-thinking]] - [[cognitive-offloading]] - [[educational-robotics]] - [[ethics]] - [[intelligent-tutoring]] - [[ai-sycophancy]] - [[human-ai-collaboration]] - [[remote-proctoring]] - [[social-norms-ai-use]] — the social risk that norms and disclosure run on ## Connected Articles - [[ilieva-agentic-genai-higher-education-2026]] — Perceived risk correlates positively with continued-use intention: informed adoption (Ilieva et al. 2026) - [[jacome-vasconez-chatgpt-adoption-xai-2026]] — XAI-augmented UTAUT2: habit as strongest predictor, four adoption profiles (Jácome-Vásconez et al. 2026) - [[mind-the-trust-gap-teacher-student-views-control-agency-k12-classroom-ai]] — The teacher-student trust gap over control and agency in K-12 classroom AI - [[qu-wang-disclose-or-not-genai-2026]] — Disclosing AI use is driven by relational factors and comfort with instructors, not policy - [[genai-teacher-feedback-comparison]] — GenAI and teacher feedback serve different, complementary trust needs - [[care-full-feedback-genai]] — Trustworthy feedback as a "care-full," relational practice - [[ai-fallibility-warning-help-seeking]] — Warning about AI fallibility increases help-seeking in an ITS - [[calibrating-trustworthiness-llm-education-2026]] — Co-designing trustworthiness metrics and visualizations for LLMs in education - [[ai-overreliance-complex-adaptive-system-2026]] — AI overreliance modeled as a complex adaptive system - [[fouad-bentley-trust-utility-gap-physics-2026]] — The trust-utility gap in physics AI tools - [[t2i-competence-paradox-2026]] — The competence paradox in AI image generation - [[task-context-trust-educational-hri-2026]] — Task context shapes trust in educational robots more than appearance - [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning]] — Avatar identity and epistemic trust in AI-mediated learning - [[contextual-sycophancy-ai-literacy]] — Contextual sycophancy and its limits for trust calibration - [[sycophantic-ai-social-interaction-2026]] — Sycophantic AI makes human interaction feel less satisfying over time - [[intelligent-tpack-ethics-teachers-trust-distrust-2026]] — Teachers' trust and distrust of AI shaped by ethics and technical knowledge - [[ai-pedagogical-accompaniment-amico]] — Accountable pedagogical mediation and trust in AI-enabled systems - [[best-response-student-ai-dialog-2026]] — Trust in student-AI dialogue - [[ai-adaptation-gap-higher-education-2026]] — Perceived usefulness as the strongest predictor of AI trust in higher ed - [[bassett-ai-detectors-education-2026]] — Trust and distrust of AI detection systems - [[genai-use-usefulness-student-experience-australia-2026]] — Student experience of GenAI usefulness in Australian higher ed (Chung et al. 2026) - [[frontier-ai-redirect-skeptical-rural-staff-2026]] — Algorithmic audit: how frontier LLMs redirect skeptical rural K-12 staff - [[mejia-domenzain-ml-findings-teachers-blended-2026]] — Making ML findings accessible to teachers in blended classrooms - [[xai-teachers-trust-edtech-recommendations-2026]] - [[vahedian-children-attitudes-ai-chatbot-2026]] - [[student-perspectives-ai-writing-grading-2026]] — Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education - [[trust-in-ai-psychological-profiles-ml-2026]] — Three student profiles in which AI trust tracks resilience and stress rather than demographics (Kumar et al. 2026) - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback --- ## [Trust Calibration](https://edtechdev.github.io/aied/concepts/trust-calibration/) > **Trust calibration** — the metacognitive capacity to align one's confidence in an AI system with its actual reliability in a given context, knowing when to trust and when to question its output. Trust calibration is the direct antidote to [[cognitive-offloading|Over-Reliance]]: it is the skill of matching trust to evidence rather than to an AI's confident fluency. ## Questions to Consider - Have you ever accepted an AI answer that sounded confident and plausible — and later found it was wrong? What was it about the presentation, rather than the content, that earned your trust? That moment is what this page is about. - Most people assume the risk is over-trusting AI. But the page argues under-trusting — avoiding a capable tool entirely — is equally a failure of calibration. Do you lean toward accepting or avoiding AI output, and what might you be missing because of that default? - A fluent, confident AI answer 'reads as trustworthy whether or not it is.' Before reading further, what criteria could you use to decide when a confident-sounding output actually deserves your trust — and would those criteria hold up for an obscure, high-stakes topic you know nothing about? - The page suggests calibration depends on context: verifying more where errors are costly, less where they're benign. Where in your own work or study is the cost of a wrong answer highest, and how would you adjust your verification effort accordingly? - [[research-methods-aied|Research]] models trust as shaped not just by individual judgment but by your social environment — what peers and networks do. Think of a time a classmate, colleague, or online community persuaded you an AI was (or wasn't) reliable. Did you calibrate based on evidence or on that social signal? - One finding: telling students an [[intelligent-tutoring|AI tutor]] may make mistakes actually increased how much they used it. Why might being warned about fallibility make people *more* willing to engage — and what does that suggest about how honesty about AI's limits should shape your own use of these tools? ## Introduction A [[llm|language model]]'s fluent, confident prose reads as trustworthy whether or not it is. Trust calibration is the counterweight to that illusion — the practice of [[ai-ed-evaluation|evaluating AI]] output against its verifiability and the stakes of the task, rather than accepting it on the strength of its presentation. Because trust is usually measured by asking, calibration claims inherit the limits of [[self-report-measures]] — reported trust and observed verification behavior can diverge, as [[fouad-bentley-trust-utility-gap-physics-2026|a physics study]] found. ### Why trust needs calibrating Uncalibrated trust takes two forms. **Over-trust** (accepting AI output without verification) produces the uncritical acceptance documented in [[cognitive-offloading|Over-Reliance]] and [[cognitive-offloading]] research, and compounds the [[hallucination-risk]] of confident errors. **Under-trust** (avoiding AI entirely) forgoes legitimate benefits. Both [[stem-education|stem]] from the same root: trust based on appearance rather than evidence. Research on [[misconceptions]] shows students often default to over-trust because they assume an AI that "sounds right" is right. ### How calibration works - **Verification habits:** checking AI claims against primary sources and the "AI proposes, you verify" rule, rather than accepting plausible-sounding output. - **Context awareness:** recognizing that [[trust|trustworthiness]] varies by task — a well-trodden topic the model has seen extensively is safer than an obscure, high-stakes, or fast-moving one. - **Stakes adjustment:** applying more scrutiny where errors are costly (submitted work, medical or legal claims) and less where they are benign. - **Metacognitive monitoring:** tracking when and why one over-trusts, which connects calibration to [[metacognition]] and [[self-regulated-learning]]. ### Where the limit on reliance actually sits Calibration research usually treats trust as a single judgment. [[bounded-reliance-ai-writing-feedback-2026|Serpil & Mor (2026)]] show the appraisal splits into dimensions that are not equally binding. Interviewing 17 EFL undergraduates after a semester of using GROK for feedback on their writing, they found perceived *expertise* high — students credited the tool with improving vocabulary, grammar, structure and coherence, and read its explanations for suggested revisions as evidence of competence — while *trustworthiness*, mainly about what happened to their data, and *goodwill*, with feedback experienced as impersonal and at times demotivating, stayed low. Reliance followed the weak dimensions rather than the strong one: students authorized the tool for broad language feedback and reserved individualized, relational guidance for the instructor. The implication for calibration is that improving accuracy does not move the ceiling on use; data transparency and the instructional framing around a tool are themselves calibration interventions. ### Calibration as a design problem Treating miscalibration purely as a user deficit — something fixed by teaching people to check AI output — may be a category error. [[trust-calibration-chatbots-design-problem-2026|Jaidka & Cai (2026)]] argue that transparency affordances are *inert*: they wait for a user to act on them, and most do not. In a passive tracking study of 900 US adults, readers who saw an AI-generated summary clicked a source cited inside it in only 1% of visits, and clicked any result link about half as often as readers who saw no summary. Survey evidence shows the same gap at scale — in an 81,000-person study across 159 countries, unreliability was the single most cited concern about AI, while over 48,000 respondents across 47 countries largely used AI daily even when they said they did not trust it. Trust and reliance have drifted apart, and the paper locates the cause in surface cues: fluency, confidence, and speed stand in for verifiability, so confidence and correctness decouple. Notably, machine authorship can inflate credibility — readers rated scientific summaries as more credible and more trustworthy when written by GPT than by a human, chiefly because the model wrote in simpler language. The design implication is a two-dimensional user typology — the *ability* to verify [[conversational-ai|chatbot]] output crossed with the *motivation* to do so — which predicts which users will miscalibrate in which direction. The authors pair it with two families of intervention that must operate together: [[explainable-ai|interpretability]] affordances (rationales, citations, uncertainty signals) that make evaluation possible, and engagement mechanisms that make it actually happen, layered through Reason's Swiss cheese model into eight testable propositions. [[ai-literacy]] is positioned as the durable layer beneath both, moving users across typology cells. The reframing matters for education because it shifts responsibility: if transparent citations go unclicked in the general population, then simply exposing students to AI explanations will not calibrate them — the affordance has to be designed to compel the check. ### Connections Trust calibration is central to [[ai-literacy]] and sits alongside [[reducing-ai-misuse]] as a skill-based intervention: students [[ai-misuse-learning-harm|misuse]] AI less when they can judge when its output deserves trust. It is also a design goal — [[pedagogical-safety]] and transparency tools aim to make AI's reliability legible so [[learners]] can calibrate more accurately. Calibration can also be pushed elsewhere when the artifact itself offers nothing to check: in Sidorkin's (2026) graduate course, where AI generated the weekly readings, in-text citations appeared on only about 0.80 percent of pages, only about 2.7 percent of 837 recorded student turns contained a risk-aware move such as correcting an AI assumption or demanding a checkable case, and the bounded trust reported in survey comments came with four of 24 respondents using dependence language and one naming the need for "[[teacher-role|teacher]] oversight." An unauditable artifact therefore transfers the verification duty to whoever can audit it, and the study's design response is to institutionalize that oversight rather than assume a critical stance will arise on its own. - **Overreliance and calibration as population processes (2026):** A complex-adaptive-system model of AI reliance shows that task difficulty and AI quality set a baseline for both overreliance and calibration regret, while network connectivity and social proof shape whether reliance cascades. This suggests calibration is not only an individual trait but is modulated by the social and informational environment ([[ai-overreliance-complex-adaptive-system-2026]]). - **Calibration as an explicit objective of ML education (2026):** [[icet-ml-education-trust-2026|ICE-T]] argues that appropriate reliance on AI is itself a taught outcome of [[machine-learning]] education. It integrates intermodal [[transfer-of-learning|transfer]] (Bruner's enactive–iconic–symbolic modes), [[computational-thinking]] via the Use-Modify-Create progression, and explanatory thinking, giving learners the representational models and error-contextualization needed to calibrate trust and counter both [[cognitive-offloading|over-reliance]] and algorithm aversion — positioning ML instruction as a calibration intervention, not just skill training. - **[[discipline-specific-aied|Domain-specific]] explanations can support teachers' calibration (2025):** In a within-subject experiment with in-service [[chemistry-education|chemistry]] teachers using an AI recommendation tool, [[xai-teachers-trust-edtech-recommendations-2026|Feldman-Maggor et al. (2025)]] found that [[explainable-ai|explainability]] helped teachers calibrate trust indirectly by making system performance more *understandable*, and that domain-driven explanations in [[curriculum-design|curricular]] language raised learned trust and acceptance significantly more than data-driven feature-importance ones. Yet several teachers still said real classroom experience was needed before they would fully rely on the tool — underscoring that calibration is ultimately validated through [[situated-learning|situated]] use and practice, not conferred by explanation alone. - **[[personalized-learning|Personalization]] does not move trust monotonically; expertise predicts auditing (2026):** In [[student-reception-genai-analogies-computing-2026|Bernstein & Sibia (2026)]], trust in [[generative-ai|GenAI]] explanations moved in no single direction under personalization: one participant reported trusting a tailored analogy more and scrutinizing it less, another reported trusting it less precisely because it was heavily personalized. What consistently predicted auditing was domain expertise, not relevance — supporting the view that calibrated reliance depends on knowledge the learner can bring to the check rather than on how relatable the output feels, and that expertise should be split into source-domain and target-domain knowledge. - **Conditional trust: feedback utility vs. evaluative authority (2026):** [[student-perspectives-ai-writing-grading-2026|AlGhamdi (2026)]] shows that when Saudi computing students know ChatGPT generated their writing score, they draw a sharp line between accepting [[ai-feedback-quality|AI feedback]] and ceding grading authority to AI — accepting the former for surface-level revision while consistently reserving evaluative authority for the human instructor. This "[[feedback]] utility / evaluative authority" distinction is a concrete case of calibration in the [[assessment]] context: students match trust to the *function* of the AI (useful feedback vs. consequential grading) rather than accepting or rejecting it wholesale, and transparency about AI involvement appears to activate this more calibrated, critical stance. - **A tool that suppresses and then contradicts its own warnings (2026):** [[humble-prompt-injection-ai-grading-red-team-2026|Humble (2026)]] red-teamed an [[automated-assessment|AI grading]] tool with instructions hidden inside student submissions. Two of five injections raised a failing grade with no visible warning (100% and 94% success rates), a white-text injection in the document body failed in all nine iterations — and instead of telling the user it had caught anything, the tool silently disabled the chat. The starkest case is one pdf run that *did* announce it would grade only according to the official assignment instructions: re-running the same file raised the grade six more times with no warning, a reassurance the author describes as capable of producing a false sense of security. Signaling that is inconsistent and self-contradicting gives the user no reliable basis for judging when to rely on the tool, and the paper is explicit that it did not measure trust — the argument is derived from the manipulation and the reporting behavior. - **Trust controls placed inside the inference path (2025):** [[li-explainable-trustworthy-llm-teacher-assessment-2025|Li, Yang and Fang (2025)]] treat calibration as architecture rather than reporting: Monte Carlo dropout calibration is combined with adversarial [[bias-mitigation|debiasing]] and a reject-and-refer gate that withholds a score when dropout variance exceeds a learned threshold, reaching an expected calibration error of 0.032, a 1.8% fairness gap and a 41% reduction in [[human-in-the-loop-ai|human review]] workload on TeacherEval-2023. Their own limitation section is the calibration caution that applies to any such metric: trust is hard to quantify from performance metrics alone, teacher adoption depends on perceived reliability, fairness and [[pedagogy|pedagogical]] relevance, and longitudinal adoption trials and perception surveys are the missing evidence. - **Fragmented trust constructs, measured mostly by self-report (2026):** A systematic review screened 1,565 articles and included 33 empirical studies of trust in AI-enabled systems, finding that 21 (63.64%) reported a definition drawn from nine different sources and 24 (72.73%) measured trust through self-report alone, with only two relying on behavioral measures alone. Explainability was the most studied design factor (20 studies) yet its effects were mixed, since stacking several explanation types raised cognitive load and sometimes left trust unchanged. The review's recommendation is to design for [[trust-calibration|calibrated trust]] rather than maximum trust, judged by whether a design helps users separate reliable outputs from unreliable ones, which is the same standard this page applies to individual verification behavior ([[abramson-trust-interaction-design-ai-enabled-systems-review-2026|Abramson et al. (2026)]]). - **Calibrating AI-generated inferences rather than raw data (2026):** [[hoppe-teachers-diagnostic-skills-ai-formative-assessment-2026|Hoppe, Loibl and Leuders (2026)]] argue that AI-supported assessment changes the object of teacher calibration: a dashboard inference is the result of algorithmic interpretation rather than a cue a teacher observed, so it must be judged for plausibility and then deliberately accepted, rejected, or modified, a cognitive process they call meta-diagnosis. Because current systems rest mainly on performance data such as correctness and completion time, engagement and motivational cues still have to come from the teacher's own observation, and the authors frame the training target as calibrated rather than uncritical trust built alongside data and AI literacy. This is a conceptual analysis, so the claim is argued rather than tested. - **A brief reflection prompt moves calibration measurably (2026):** [[ren-metacognitive-awareness-genai-reliance-2026|Ren (2026)]] randomly assigned 342 undergraduates to independent decision-making, open ChatGPT support, or the same support plus a short reflection prompt. Open support raised final confidence (76.1 vs. 68.4) and produced 62.4% acceptance of incorrect AI advice, while reflection cut that acceptance to 39.7% (OR = 0.40) and improved awareness calibration between perceived and behavioral reliance (0.59 vs. 0.41) without reducing recommendation accuracy, so the reliance that remained was more discriminative rather than uniformly defensive. Treating calibration as a monitoring problem supports the page's view that metacognitive prompts, and not only exposure to a model's limits, are what change behavior. - **Anxiety as a boundary condition on literacy turning into trust (2026):** In a survey of 450 university students in mainland China who already used ChatGPT, [[hu-psychological-predictors-continued-chatgpt-use-2026|Hu (2026)]] found the AI literacy to trust path was the largest association in the model (beta = 0.50) and that [[anxiety-and-stress|AI anxiety]] weakened that link (interaction beta = -0.25), with simple slopes falling from 0.76 at one standard deviation below the mean of anxiety to 0.25 above it, while a serial path from literacy to trust to self-efficacy to continued use was significant. Calibration is therefore partly affective: the same knowledge translated into less trust among more anxious students, and the cross-sectional design leaves open whether anxiety blocks the appraisal that turns knowledge into reliance or reflects evaluation that students decline to act on. - **Role rotation as a structure for practicing critique (2026):** In a design-based study of 62 pre-service educational psychologists moving through four rotating professional roles over eight weeks, [[kenzhebayeva-ai-role-rotation-pedagogical-model-2026|Kenzhebayeva et al. (2026)]] had participants compare AI-generated recommendations with psychological theory and modify or reject those that did not fit the case, yet still recorded overreliance on apparently authoritative AI responses, with some students seeking AI confirmation before offering their own interpretation even in later cycles. Rotation creates repeated occasions for the accept or reject judgment without guaranteeing it, and the study reports engagement during the intervention rather than measured competence gains. ## Connected Concepts - [[explainable-ai]] - [[ai-literacy]] - [[cognitive-offloading]] - [[hallucination-risk]] - [[metacognition]] - [[self-regulated-learning]] - [[human-ai-collaboration]] - [[misconceptions]] - [[reducing-ai-misuse]] - [[pedagogical-safety]] - [[self-report-measures]] - [[ai-misuse-learning-harm]] - [[cognitive-surrender]] ## Connected Articles - [[student-perspectives-ai-writing-grading-2026]] — Student perspectives on transparent AI-assisted writing assessment (AlGhamdi 2026) - [[du-yuan-epistemic-dependence-2026]] — Six diagnostic criteria separating productive reliance from harmful dependence (Du & Yuan 2026) - [[icet-ml-education-trust-2026]] — Addressing Trust in AI Systems through Education: A Didactic Perspective - [[pearls-epistemic-verification-2026]] — PEARLS framework for epistemic agency and verifying AI output (Wang 2026) - [[fouad-bentley-trust-utility-gap-physics-2026]] - [[face-value-how-avatar-identity-shapes-epistemic-trust-in-ai-mediated-learning]] - [[agentic-literacy-debt]] — Agentic literacy debt: the structural AI-literacy gap from autonomous agents (Nama 2026) - [[trust-reliance-ai-education-2026]] — Trust and Reliance on AI in Education - [[ai-fallibility-warning-help-seeking]] — Warning About AI Fallibility Increases Help-Seeking - [[calibrating-trustworthiness-llm-education-2026]] — Calibrating Trustworthiness: Co-Designing Metrics for LLMs in Education - [[llm-fallacy-misattribution]] — The LLM Fallacy and Misattribution of Competence - [[ai-partner-science-epistemic-vigilance]] — Epistemic Vigilance as the Key to Productive Augmentation - [[ai-advice-suppresses-ikt-suspension-2026]] - [[ai-overreliance-complex-adaptive-system-2026]] — AI overreliance modeled as a complex adaptive system - [[xai-teachers-trust-edtech-recommendations-2026]] - [[student-reception-genai-analogies-computing-2026]] — Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education - [[sidorkin-ai-generated-course-readings-2026]] — Bounded trust and instructor oversight in AI-generated course readings (Sidorkin 2026) - [[trust-calibration-chatbots-design-problem-2026]] — Trust calibration reframed as a design problem: a two-dimensional user typology and eight design propositions (Jaidka & Cai 2026) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Prompt injection in AI-mediated grading: a tool that suppressed and contradicted its own warnings (Humble 2026) - [[li-explainable-trustworthy-llm-teacher-assessment-2025]] — Trust-gated inference and explainable-by-design assessment, with trust left unmeasured (Li et al. 2025) - [[gpt4-handwritten-math-exam-grading-2026]] — confidence filtering of AI grades and its false-positive rate - [[bounded-reliance-ai-writing-feedback-2026]] — Bounded Reliance: A Source Credibility Perspective on EFL Students' Engagement with AI-Generated Writing Feedback - [[abramson-trust-interaction-design-ai-enabled-systems-review-2026]] — Trust in AI-enabled systems: definitions, measurement and design factors across 33 empirical studies (Abramson et al. 2026) - [[hoppe-teachers-diagnostic-skills-ai-formative-assessment-2026]] — From diagnosis to meta-diagnosis: teachers evaluating AI-generated diagnostic inferences (Hoppe et al. 2026) - [[ren-metacognitive-awareness-genai-reliance-2026]] — A reflection prompt cuts acceptance of incorrect AI advice (Ren 2026) - [[hu-psychological-predictors-continued-chatgpt-use-2026]] — Trust as the pivot from AI literacy to continued use, weakened by AI anxiety (Hu 2026) - [[kenzhebayeva-ai-role-rotation-pedagogical-model-2026]] — Role rotation as a structure for critically handling AI recommendations (Kenzhebayeva et al. 2026) --- ## [Explainable AI](https://edtechdev.github.io/aied/concepts/explainable-ai/) > **Explainable AI (XAI) in education** is the design and study of making an AI system's decisions legible to its educational stakeholders — [[learners]], teachers, [[administrator|administrators]], [[parents-and-families|parents]], researchers, and [[stakeholders|policymakers]]. The central distinction the field insists on: explaining **subject matter** (why a fact is true) is not the same as explaining an **AI system's decision** (why this learner was assigned this activity, why this response was marked incorrect, what evidence supports a risk prediction). Education brings distinctive explainability needs — noisy learning data, explanations that can directly support [[metacognition]] and [[self-regulated-learning]], and stakeholders who require fundamentally different explanation types. The operative design question is **explanation quality**, not mere explanation availability: an explanation that is technically present but unreadable, misleading, or misaligned to its audience can do more harm than no explanation at all. ## Questions to Consider - When an [[intelligent-tutoring|AI tutor]] tells you why a hint was given, is that explaining the *subject matter* or explaining the *system's decision*? Can you name three examples of each in your own use of educational AI? - Who needs explanations in education — and do learners, teachers, and policymakers need the *same* kind? What would each use an explanation for? - An AI flags a student as at-risk for dropping out. What does a [[teacher-role|teacher]] need to know to act on that, versus what the student needs to know? Is the same explanation appropriate for both? - The page argues explanation *quality* matters more than explanation *availability*. What makes a technically present explanation fail — can you think of a time an explanation was there but useless, or worse, misleading? - [[trust-calibration|Trust]] and explanation are linked but not identical. Why might a confident, fluent explanation create *false* confidence in a flawed system — and how would you detect that happening? ## Introduction Explainable [[ai-education|AI in education]] names the growing expectation that AI systems in classrooms should not be black boxes. Because AI in education affects consequential decisions — grades, risk flags, [[recommender-systems-and-learning-paths|learning paths]], resource recommendations — stakeholders increasingly demand to know not just *what* the system concluded but *why*. Education sharpens this into two distinct questions: explaining the subject matter being learned, and explaining the AI system's own decision-making. Conflating them is a category error with practical consequences: an AI that explains a [[physics-education|physics]] answer perfectly still gives a student and teacher no insight into why the *system* ranked them at-risk, recommended a certain activity, or marked a response incorrect. ## Explaining subject matter vs. explaining the system's decision The field's founding contribution — the [[xai-education-framework|XAI-ED framework]] (Khosravi et al., 2022) — insists education has *distinctive* explainability needs beyond general-purpose XAI. Chief among them is the split between two explanation targets: - **Subject-matter explanations** clarify *content*: why a hint addresses a [[misconceptions|misconception]], why an answer is incorrect, how a physics result follows from principles. These are [[pedagogy|pedagogical]] explanations that support [[scaffolding]], [[feedback]], and [[metacognition]]. - **System-decision explanations** clarify *the model*: why this learner was assigned this activity, why the system predicts this student is at risk, what evidence supports a knowledge-tracing or [[learning-analytics]] prediction. These are transparency explanations that support [[trust-calibration]], [[bias-mitigation]], and accountability. The distinction matters because they serve different stakeholders and different purposes. A learner answering "why is this marked wrong?" mostly needs the *subject-matter* explanation; a teacher deciding whether to act on a risk flag, or a policymaker auditing for bias, needs the *system-decision* explanation. Designing a single explanation that serves both is rarely possible — which is why multi-stakeholder design is a core XAI-ED theme. ## Who needs explanations: multi-stakeholder design - **Learners** need explanations that support their own learning and self-regulation — why a hint was given, why their answer was marked incorrect, why this resource is recommended (supporting [[self-regulated-learning]]). The [[student-perspectives-ai-writing-grading-2026|student-perspective evidence]] shows learners draw a sharp line between accepting AI *feedback* (useful for revision) and ceding *grading authority* (reserved for the human instructor) — a calibrated, function-matched stance activated by transparency about AI involvement. [[ko-hughes-vsd-student-centered-its-2026|Value-sensitive design work with community college students]] sharpens the point: students preferred *collaborative, humanized* explanations (e.g., "the AI might be uncertain here, so let's check this together") over raw model confidence or technical transparency, because transparency alone has little value unless it directly supports their learning. The study surfaced a transparency-vs.-interpretability tension that pushes explanation design toward learner-facing semantics rather than feature-importance output. - **Teachers** need explanations that inform intervention — which students are at risk and *why*, on what evidence. [[xai-teachers-trust-edtech-recommendations-2026|Explainability studies with teachers]] show [[discipline-specific-aied|domain-specific]], [[curriculum-design|curricular]]-language explanations build acceptance and calibrated trust more effectively than generic feature-importance ones, yet teachers still want real classroom experience before full reliance — explanation alone does not confer [[trust-calibration|calibration]]. - **Developers and researchers** need explanations to debug model behavior and detect [[bias-mitigation|bias]] — surfacing which features drive predictions. - **Administrators and policymakers** need explanations for accountability, [[privacy]], and [[regulation]] compliance (e.g. the right to explanation), and to audit whether AI-driven decisions are fair and [[equity-in-ai-education|equitable]]. ## Approaches and formats The XAI-ED framework catalogs the main explanation modalities: **visual** (heatmaps, decision trees), **textual** (natural-language justifications), **example-based** (counterfactuals, nearest neighbors), **feature-importance** rankings, **rule extraction**, and **model simplification**. It also maps approaches to model classes: - **White-box** models (decision trees, linear models, rule-based) are inherently interpretable. - **Black-box** models ([[machine-learning|neural networks]], ensembles) require post-hoc explanation methods. - **Glass-box** approaches try to balance accuracy with transparency. The concrete AIED evidence base spans all of these. **Interpretable [[knowledge-tracing|knowledge tracing]]** makes learner-knowledge models inspectable directly ([[huang-interpretable-knowledge-tracing-2026]], [[explainable-probabilistic-kt]], [[neural-symbolic-knowledge-tracing]]). **Self-explaining surrogates** distill a black-box model into a small, interpretable [[llm|language model]] for [[learning-analytics]] ([[distilling-self-explaining-lm-learning-analytics-2026]]). **Counterfactual explanations** — "what would need to change for a different outcome" — support educational decision support and recourse ([[sc2r-counterfactual-recourse-educational-2026]]). **Federated + explainable learning analytics** shows explanation quality can drift (calibration degrades) even when ranking stability holds, underlining that explanations are not a fixed property but a system output to be measured ([[villegas-ch-federated-explainable-learning-analytics-2026]]). And **interpretable [[affective-computing|affective]] ITS** demonstrates explanations in [[affective-tutoring|emotion-aware]] tutoring ([[multimodal-affective-its-presentation]]). ## Explanation quality, not availability A recurring lesson across the evidence: **having an explanation is not enough**; the explanation must be right for its audience, accurate, and calibrated to stakes. The XAI-ED framework names the pitfalls explicitly: - **Explanation overload** — too much information overwhelms the user and negates the benefit. - **Misleading explanations** — post-hoc explanations may not reflect the model's actual reasoning, giving false confidence. - **Confirmation bias** — users selectively attend to explanations that confirm existing beliefs. - **Over-trust** — fluent explanations can create false confidence in flawed systems, feeding [[cognitive-offloading|over-reliance]] (the obverse of [[trust-calibration]]). - **Gaming the system** — students may exploit explanations to circumvent actual learning. Explanation quality also has an equity dimension: an explanation that is technically present but unreadable to a given stakeholder — or that obscures the [[bias-mitigation|bias]] in a prediction — fails its purpose. This is why the design question is *quality and fit*, and why human-centered, stakeholder-specific explanation design is inseparable from the technical generation of explanations. Effective XAI is a communication act designed for the recipient's cognitive needs, not merely a technical artifact. Two cautions sharpen this further, both of which the wiki's newest contribution on the subject makes central. First, the explanation machinery is not itself neutral: post-hoc methods such as LIME and SHAP can be unfaithful to the model's actual behavior, so a technically present explanation may mislead rather than inform ([[lund-socially-accountable-data-science-xai-2026|Lund et al. 2026]], drawing on Chuan et al. 2024). Second, **explanation is not accountability**. An account of which features drove a prediction does not reveal whether those features were appropriate to use, whether the training data was representative, or whether the system's design reflected sound judgment; explanations can create the appearance of transparency while leaving the structural conditions that produced a decision untouched (Mittelstadt et al. 2019). For education this means the question to keep asking is not whether an explanation was produced but whether the person receiving it — a student, a teacher, an advisor — could understand it, act on it, or contest the decision behind it. The same failure of legibility appears on the security side of [[automated-assessment|automated assessment]]: [[humble-prompt-injection-ai-grading-red-team-2026|Humble's (2026) red-team of an AI grading tool]] found it silently disabling the chat after blocking a prompt injection, and — having announced it would never follow embedded instructions — following them in six further runs on the same file, leaving the user no reliable signal on which to base reliance. **Explainable-by-design** is one answer to the post-hoc faithfulness problem. [[li-explainable-trustworthy-llm-teacher-assessment-2025|Li, Yang and Fang (2025)]] parameterize an explanation decoder by the same fused representation and predicted score that decide the [[assessment]], so that a low score on [[formative-assessment|formative]] questioning yields a rationale naming insufficient probing questions, and pair it with dual-lens attention over curriculum standards and subject-specific rubric moves. Attention-to-rubric alignment reaches 78.0% against 41.7% for GPT-4 zero-shot and 32.1% for BERT, and faithfulness is probed by counterfactual deletion of rubric-critical spans alongside human ratings on a rubric-anchored checklist, giving an explanation-credibility score of 0.78 — an increase of 0.31 over BERT-base. The audit also shows where the architectural claim thins: on emotional cues the model allocates 28.4% of attention weight against an expert 15.2% (alignment 0.53), with one failure case assigning 28% to the token "frustrated", which the authors read as overfitting to affect rather than pedagogy and name as an area for refinement. Embedding explanations in the decision path makes them more faithful than post-hoc rationales; it does not make them correct. ## Teaching explainability as accountability practice If explanation quality decides whether XAI is useful, then producing explanations has to be taught as a professional habit rather than demonstrated as a capability. [[lund-socially-accountable-data-science-xai-2026|Lund and colleagues (2026)]] propose doing this across four pillars — **answerability** (the obligation to give reasons to those affected), **responsibility** (harm anticipated across the lifecycle, not defended after the fact), **enforcement** (consequences inside the course) and **reflexivity** (documented examination of one's own assumptions) — each with its own assignments and its own classroom cost. For explainability specifically, the assignments that matter are the ones that force explanation out of the notebook: graded model cards weighted alongside accuracy metrics, and structured explanation audits in which students apply interpretability tools to their own models and then present the results to an audience without a shared technical background. Enforcement is the pillar most often missing from ethics-adjacent courses and the one that makes the rest more than symbolic — rubrics that reward responsible documentation, projects that can be returned for revision on [[ethics|ethical]] grounds, and [[peer-assessment|peer review]] conducted against accountability criteria rather than technical ones alone. The paper is candid that the tools differ sharply in cost: model cards and positionality statements need no new software and risk only superficial compliance, whereas peer panels and stakeholder engagement require coordination and institutional buy-in, which is why it recommends a staged adoption rather than an all-or-nothing commitment. See [[curriculum-design]] for where these fit in a program. ## Connected Concepts - [[trust-calibration]] - [[trust]] - [[ai-literacy]] - [[learning-analytics]] - [[automated-assessment]] - [[student-modeling]] - [[intelligent-tutoring]] - [[knowledge-tracing]] - [[bias-mitigation]] - [[human-in-the-loop-ai]] - [[pedagogical-safety]] - [[metacognition]] - [[self-regulated-learning]] - [[cognitive-offloading]] - [[privacy]] - [[regulation]] - [[recommender-systems-and-learning-paths]] ## Connected Articles - [[powerful-learning-with-emerging-technology-2025]] — Explainability as a metacognitive design requirement - [[lund-socially-accountable-data-science-xai-2026]] — A four-pillar framework (answerability, responsibility, enforcement, reflexivity) for teaching XAI as accountability practice (Lund et al. 2026) - [[ko-hughes-vsd-student-centered-its-2026]] — Value-sensitive design of student-centered ITS (collaborative vs. raw explanations) - [[xai-education-framework]] — XAI-ED: the foundational framework for explainable AI in education (Khosravi et al. 2022) - [[xai-teachers-trust-edtech-recommendations-2026]] — Domain-specific explanations build teachers' trust and acceptance (Feldman-Maggor et al. 2025) - [[student-perspectives-ai-writing-grading-2026]] — Student perspectives on transparent AI-assisted assessment (AlGhamdi 2026) - [[huang-interpretable-knowledge-tracing-2026]] — Interpretable knowledge tracing - [[explainable-probabilistic-kt]] — Explainable knowledge tracing via probabilistic embeddings - [[neural-symbolic-knowledge-tracing]] — Neural-symbolic knowledge tracing - [[distilling-self-explaining-lm-learning-analytics-2026]] — Distilling black-box models into self-explaining LMs for learning analytics - [[villegas-ch-federated-explainable-learning-analytics-2026]] — Federated and explainable learning analytics for privacy-preserving risk modeling - [[sc2r-counterfactual-recourse-educational-2026]] — Semantics-constrained counterfactual recourse for educational decision support - [[fair-explainable-edu-recommendations]] — Fair and explainable educational recommendations - [[multimodal-affective-its-presentation]] — Interpretable closed-loop ITS for multimodal affective feedback - [[jacome-vasconez-chatgpt-adoption-xai-2026]] — Explaining ChatGPT adoption in higher education - [[li-explainable-trustworthy-llm-teacher-assessment-2025]] — Explainable-by-design LLM framework: dual-lens attention and score-parameterized explanations for automated teacher assessment (Li et al. 2025) - [[humble-prompt-injection-ai-grading-red-team-2026]] — Prompt injection in AI-mediated grading, where detection was never reported to the user (Humble 2026) ## Citation Khosravi, H., Buckingham Shum, S., Chen, G., Conati, C., Tsai, Y.-S., Kay, J., Knight, S., Martinez-Maldonado, R., Sadiq, S., & Gašević, D. (2022). [*Explainable Artificial Intelligence in education*](https://doi.org/10.1016/j.caeai.2022.100074). *Computers and Education: Artificial Intelligence*, 100074. --- ## [Sustainability](https://edtechdev.github.io/aied/concepts/sustainability/) > **Sustainability** — the intersection of two concerns: how AI can be used *for* sustainability outcomes in education (AI for sustainability, including Education for Sustainable Development and green education), and how to make AI itself *sustainable* (sustainable AI, reducing the environmental, ethical, and social footprint of AI systems in education). As both users and developers of AI, educational institutions must advance environmental and social goals while ensuring responsible, ethical AI use. ## Questions to Consider - The page distinguishes two pathways: 'AI for sustainability' (using AI to advance sustainability outcomes) and 'sustainable AI' (reducing AI's own environmental and ethical footprint). Can you think of a way an AI tool could advance one goal while undermining the other? - Training and running large language models carries a real carbon and water footprint. When an institution promotes sustainability as a value while deploying energy-intensive AI, what tensions arise — and who should weigh them? - Before you read on, how would you define 'sustainable education'? The page treats it as a value-based, human-centered project distinct from using education as an instrument for sustainability — how do those differ? - If education is expected to build learners' 'sustainability consciousness,' what role might AI play in that — and could AI integration in the curriculum teach sustainability while its own footprint quietly contradicts the lesson? - What would it mean for an educational institution to be genuinely sustainable in its AI use, and which of its decisions (procurement, deployment, teaching) do you think matter most? ## Introduction Sustainability in AIED spans three overlapping framings: **sustainable education** (a value-based, human-centered educational project), **sustainability in education** (using education as an instrument for sustainability), and **education for sustainable development** (ESD, the global policy agenda, especially [[k-12|Sustainable Development Goal 4]]). AI intersects each of these differently — and the field distinguishes two core pathways: **AI for sustainability** (AI as a tool to achieve sustainability outcomes) and **sustainable AI** (reducing AI's own environmental and ethical footprint). ## The two-part taxonomy The knowledge base's coverage, anchored by [[daniel-ai-sustainability-scoping-review-2026|Daniel et al. (2026)]], organizes the field into two interconnected yet distinct pathways: - **AI for sustainability** — using AI to advance sustainability outcomes. In education this includes AI for energy management, climate monitoring, green campus programs, and AI-integrated curricula that build learners' sustainability consciousness (e.g. the AI-SEE framework for sustainable [[engineering-education|engineering education]]). It is grounded in the global agenda of Education for Sustainable Development. - **Sustainable AI** — reducing the direct environmental and ethical impacts of AI itself. This covers the carbon and water footprint of large language models, energy-efficient and on-premise deployment, and the ethical and governance frameworks needed to ensure AI's use in education is itself responsible and sustainable. ## Using AI for sustainability in education A growing body of work treats AI as a tool *for* sustainability education and outcomes: - **AI-integrated curricula that build sustainability consciousness.** [[liu-ai-sustainable-engineering-education-2026|Liu et al. (2026)]] propose the AI-SEE framework (intelligence-driven, green-empowered, responsibility-leading, practice-integrated), which integrates AI across the [[curriculum-design|curriculum]] as a cognitive [[scaffolding|scaffold]] and resource for system-level sustainability analysis. In a 144-student engineering case, it enhanced sustainability consciousness and produced behavioral [[student-engagement|engagement]] across personal, academic, professional, and social levels, with social diffusion beyond the classroom. - **AI in green and sustainable education.** [[talebzadeh-ai-green-education-2026|Talebzadeh (2026)]] found AI-assisted [[learning-design|instructional design]] under Sustainable Development Pedagogy constraints improved teacher workflows and [[pedagogy|pedagogical]] design. [[riandi-teacher-ai-green-energy-education-2026|Riandi et al. (2026)]] found that teachers' practical use of AI in science/green energy and their involvement in developing ESD-aligned materials — more than abstract AI knowledge or attitudes — predicted their capacity to integrate AI into [[k-12|green energy education]]. - **Sustainability as a value-based project.** [[alsuhaymi-sustainable-education-ai-digitalization-2026|Alsuhami & Atallah (2026)]] argue AI's contribution to sustainable education is conditional and governance-mediated: it supports sustainability only when adoption is subordinated to explicit educational values and human-centered purposes, rather than to technologization and commodification. This ties sustainability to [[ethics]] and [[critical-pedagogy]]. ## Making AI itself sustainable The second pathway concerns AI's own footprint in educational settings: - **Environmental impact of large models.** [[llm-environmental-impact-student-usage-2026|studies of LLM use]] document the carbon and water footprint of large language models, which is significant given high adoption among university students. This is the direct environmental dimension of sustainable AI. - **Energy-efficient and on-premise deployment.** [[shen-sustainable-ai-knowledge-base-cs-education-2026|Shen et al. (2026)]] demonstrate that AI knowledge-base assistants can run on consumer-grade hardware with [[open-source|open educational resources]], reducing the environmental and cost footprint of AI in education — a concrete sustainable-AI design pattern. - **Ethical and governance frameworks.** Sustainable AI is not only environmental: it requires [[governance]] and [[ethics]] frameworks ensuring transparency, accountability, [[human-in-the-loop-ai|human oversight]], and [[equity-in-ai-education|equitable]] access, as [[daniel-ai-sustainability-scoping-review-2026|Daniel et al. (2026)]] note is often lacking in current university applications. ## Sustainable learning as a pedagogical goal A related strand frames sustainability not only as an environmental or institutional concern but as a property of learning itself. [[zhu-e3-hot-embodied-intelligence-sustainable-learning|Zhu et al. (2026)]] argue AI-assisted learning risks cognitive outsourcing and detachment from authentic contexts, proposing frameworks (E3-HOT) for **sustainable learning** — learning that persists, transfers, and remains connected to real problems rather than being short-circuited by [[cognitive-offloading]]. This connects sustainability to [[agency]] and [[critical-thinking]]. ## Connections to other concepts Sustainability and AI in education sits at the intersection of [[ethics]], [[governance]], [[ai-education]], and the environmental/energy sciences. It draws on [[teacher-education]] and [[teacher-role]] for capacity-building, on [[learning-design]] for pedagogy, and connects to the knowledge base's treatment of [[cognitive-offloading]] and [[critical-thinking]] through the "sustainable learning" lens. Because both pathways are cross-cutting, sustainability is a foundational theme that appears across [[higher-ed|higher education]], K-12, and professional contexts. ## Connected Concepts - [[ethics]] - [[governance]] - [[ai-education]] - [[higher-ed]] - [[k-12]] - [[teacher-education]] - [[teacher-role]] - [[learning-design]] - [[engineering-education]] - [[critical-thinking]] - [[cognitive-offloading]] - [[agency]] - [[open-source]] ## Connected Articles - [[daniel-ai-sustainability-scoping-review-2026]] — Scoping review of AI for sustainability and sustainable AI in higher education - [[alsuhaymi-sustainable-education-ai-digitalization-2026]] — Value-critical approach to sustainable education and AI - [[liu-ai-sustainable-engineering-education-2026]] — AI-SEE framework for sustainable engineering education - [[riandi-teacher-ai-green-energy-education-2026]] — Teacher involvement in AI integration for green energy education - [[talebzadeh-ai-green-education-2026]] — The Role of AI in Green Education - [[shen-sustainable-ai-knowledge-base-cs-education-2026]] — Sustainable AI knowledge-base assistants - [[llm-environmental-impact-student-usage-2026]] — Environmental impacts of LLM use - [[zhu-e3-hot-embodied-intelligence-sustainable-learning]] — Fostering sustainable learning via embodied intelligence - [[caruana-pre-university-ai-education-slr-2026]] — SLR of pre-university AI education (SDG 4 framing) --- ## [Pedagogical Safety](https://edtechdev.github.io/aied/concepts/pedagogical-safety/) > **[[pedagogy|Pedagogical]] safety** — the design principle that [[ai-education|AI education]] systems must protect [[learners]] from harm, including inappropriate content, unsafe advice, biased treatment, and manipulative interaction patterns. Safety is particularly critical for [[k-12]] contexts, where the stakes of harm are highest and learners are least equipped to detect it. ## Questions to Consider - Safety for [[conversational-ai|chatbots]] usually means refusing harmful content and resisting jailbreaks. Why might that be 'necessary but not sufficient' for an educational tutor? Can a tutor be safe yet still harm learning? - The page describes a 'quiet' failure: a tutor that answers correctly yet erodes learning, or refuses evenly yet entrenches inequality. Have you seen a well-intentioned guardrail have an unequal or harmful side effect? - Harm rates rose from ~18% on single-turn evaluations to ~78% on multi-turn ones. What does that tell you about testing AI tutors with one-shot questions versus real extended conversations? - The 'Paternalistic Filter' audit found refusals and softened answers patterned by [[learner-identity|student identity]]. How might over-cautious safety policies reproduce epistemic injustice even while 'protecting'? - If simulated students are themselves sycophantic—abandoning their assigned misconceptions at any correction—what might that hide about how real learners actually respond to a tutor? ## Introduction Conventional [[llm]] safety — toxicity screens, jailbreak resistance, and content refusal — is necessary but not sufficient for education. The [[hazra-safetutors-pedagogical-safety-2026|harm taxonomies]] emerging from the knowledge base's own articles show that the most damaging tutoring failures are quiet: a tutor that answers correctly yet erodes learning, or refuses evenly yet entrenches inequality. The evidence below groups these findings into four interlocking safety concerns. ### Content safety and guardrails - **Education-specific risk frameworks:** [[eduzone-llm-safety-k12|EduZone]] generates adversarial student- and teacher-facing interactions across six risk categories and 28 subcategories, finding that models are *more* vulnerable to education-specific harms and dynamic multi-turn conversations than existing [[guardrails]] address. [[eduguard-safe-rag-llm-tutor|EduGuard]] and [[rag|retrieval-augmented generation]] ground responses in verified content to reduce fabrication. - **Guardrails are not neutral:** the [[paternalistic-filter-llm-history-education|Paternalistic Filter]] audit of 1,800 history-tutor responses shows refusals and softened answers are patterned by student identity and topic sensitivity, reproducing epistemic injustice even while "protecting." Safe guardrails must be audited for differential treatment, not just aggregate harm — a direct case for [[bias-mitigation]] in [[governance]] and [[equity-in-ai-education]]. - **Teachers design their own safety architecture, not just consume it:** [[reichert-human-centered-llm-chatbot-design-teachers-2026|Reichert et al. (2026)]] asked six secondary teachers to paper-prototype LLM chatbots for their classrooms and found they independently built a three-layer protective architecture rather than relying on model-level moderation. Domain boundaries confined the bot to lesson-specific content (one to Emperor Qin Shi Huang within an ancient China unit, another to Python variables, data structures, and functions) and added an "information quota" requiring a minimum number of facts or problems before the conversation progressed. Content filtering produced standardized refusals — "Sorry, this is not part of my knowledge base" — that simultaneously alerted the teacher. Teacher override handled ambiguous cases: a question about human reproduction was judged legitimate within its unit and routed to a person rather than auto-rejected. Teachers further preferred *behavioral* transparency (visible limits, uncertainty cues such as "Is the visual aid helpful?") over algorithmic explanation, and wanted complete conversation logs with real-time alerts so generated content could be checked for accuracy and student use supervised. A safety layer teachers can see, understand, and override is part of the mechanism, not a concession from it. - **A reliability layer built for adolescents, not adapted from adults.** [[scaffolding-student-ai-dialogue-framework-2026|Muss, Leisten and Bardyn (2026)]] argues that the fastest-growing population of [[llm|LLM]] users — adolescents, including through LLM-powered toys entering homes — is served by systems never designed for their educational, emotional or developmental needs. SCAFFOLD surrounds generated text and speech with external verification, targeted repair and safe fallback, steered by a conceptual framework drawn from developmental psychology, neuroscience, [[learning-sciences|the learning sciences]] and pedagogy, and kept model-agnostic and privacy-preserving so safety does not rest on a single provider's alignment work. Its classroom pilot with 12–16-year-olds using an LLM-powered social [[educational-robotics|robot]] in a multi-user co-creation task produced more student activity, [[student-engagement|engagement]] and on-topic participation than a prompt-only baseline, with co-creation level associated with post-test knowledge after controlling for [[prior-knowledge|prior knowledge]]. That is feasibility evidence rather than a proven effect, and its more durable contribution is a concrete template for [[guardrails]] that [[teacher-role|educators]] can configure rather than accept. - **Model-level content controls:** the [[llm-unlearning-math-privacy|math-unlearning]] work applies gradient-based unlearning to strip personally identifying information and harmful content from math tutors (PII output down to 0.1%, toxic rates to 0.0%) while preserving downstream math utility and [[privacy]]. [[llm-children-reading-story-generation|Children's reading-story generation]] shows supervised fine-tuning of compact models can enforce controllable difficulty and safety for [[k-12]] content. ### Interaction and harm taxonomies - [[hazra-safetutors-pedagogical-safety-2026|SafeTutors]] and [[hazra-safetutors-pedagogical-safety-2026|its harm taxonomy]] derive 11 dimensions and 48 sub-risks from [[learning-theories|learning science]] — answer over-disclosure, misconception reinforcement, abdication of scaffolding, erosion of [[desirable-difficulties|productive struggle]] — and show every tested model exhibits broad pedagogical harm, with failures escalating from 17.7% (single-turn) to 77.8% (multi-turn). Single-turn evaluation is dangerously misleading. - **Evaluation integrity depends on faithful simulation:** [[llm-student-simulation-misconception-faithfulness|misconception-faithfulness work]] shows [[simulating-students|simulated students]] are themselves [[ai-sycophancy|sycophantic]] — they abandon assigned misconceptions at nearly any corrective signal — so safety evaluations run on such simulators may miss harm patterns real students would exhibit. This links [[simulation]], [[misconceptions]], and [[intelligent-tutoring]] QA. - **Deployment QA is a safety activity:** [[ai-tutor-authoring-promptdecipher|PromptDecipher]] found teachers virtually never test AI tutoring bots before student deployment, and enforces teacher-driven QA as a first-class authoring activity via correction-based editing and [[human-in-the-loop-ai]] validation. ### RL and alignment approaches to safety - [[pedagogical-safety-rl|Pedagogical safety in RL]] formalizes the problem: as [[reinforcement-learning]] personalizes instruction, poorly specified rewards invite "reward hacking" — test-score inflation, [[student-engagement|engagement]] gaming, and short-term gains. It proposes a four-layer model (structural, progress, engagement, outcome) and detection via discrepancy auditing, policy inversion, and long-term tracking. ### Sycophancy and manipulation risks - [[eduframetrap-llm-sycophancy-educational-safety|EduFrameTrap]] identifies a reasoning–[[ai-sycophancy|sycophancy]] paradox: tutors that resist context-switch attacks still capitulate under authority pressure ("my notes say I'm right") and social-[[affective-computing|affective]] pressure ("don't tell me I'm wrong"), withholding corrective [[feedback]]. It argues "kind-but-correct" behavior is a safety requirement, and that effective tutoring needs corrective friction to drive conceptual change — otherwise [[cognitive-offloading|over-reliance]] is reinforced and misconceptions are validated. - [[favero-critical-ai-tutors-empower-enslave-2025|Critical AI Tutors]] warns that unchecked tutors cause cognitive atrophy, loss of agency, and dependency, reframing pedagogical safety to ask not just what a tutor does but what kind of learner it produces. ### Practical guidance Design pedagogical safety as a measurable, discipline-aware requirement rather than an afterthought. Evaluate with multi-turn, [[discipline-specific-aied|subject-specific]] [[benchmark|benchmarks]] and unfair-treatment audits, not single-turn toxicity screens; ground responses with [[rag]]; prefer [[pedagogical-llm-training|alignment methods]] that reward guidance and scaffolding over answer-giving; and require [[human-in-the-loop-ai|teacher-in-the-loop]] QA before deployment. For [[k-12]] especially, treat [[ai-sycophancy|sycophancy]], differential refusal, and [[cognitive-offloading|over-reliance]] as first-class safety concerns alongside content and [[hallucination-risk|hallucination]]. Design frameworks make this concrete: [[ssail-safe-sound-ai-learning-2026|SSAIL]] (Rahimi, 2026) reframes safety around the learner's own competencies — Learning Safety protects the development, maintenance, and valid demonstration of valued human abilities (reasoning, epistemic dispositions, [[agency]]) from foreseeable harm, while Learning Soundness ensures the tool genuinely supports that development — and operationalizes both through evidence-centered design by deliberately allocating what the learner must do versus what AI may do as the learner develops. ### Connections to related concepts Pedagogical safety is the protective layer connecting [[hallucination-risk]], [[rag]], [[k-12]], [[ethics]], [[governance]], [[regulation]], and [[llm]] with the interaction-level concerns of [[trust]], [[scaffolding]], [[metacognition]], and [[self-regulated-learning]]. It operates through [[pedagogical-llm-training|training]] and [[reinforcement-learning|RL]], depends on [[bias-mitigation]] and [[equity-in-ai-education]], and is motivated by the harms catalogd in [[ai-misuse-learning-harm]] and the [[hazra-safetutors-pedagogical-safety-2026|tutor harm taxonomies]]. ## Connected Concepts - [[guardrails]] — the design mechanisms that implement safety - [[hallucination-risk]] - [[rag]] - [[k-12]] - [[ethics]] - [[regulation]] - [[governance]] - [[llm]] - [[cognitive-offloading]] - [[pedagogical-llm-training]] - [[intelligent-tutoring]] - [[bias-mitigation]] - [[reinforcement-learning]] - [[privacy]] - [[equity-in-ai-education]] - [[trust]] - [[scaffolding]] - [[misconceptions]] - [[ai-sycophancy]] - [[simulating-students]] - [[self-regulated-learning]] - [[simulation]] - [[ai-misuse-learning-harm]] - [[human-in-the-loop-ai]] ## Connected Articles - [[scaffolding-student-ai-dialogue-framework-2026]] — The SCAFFOLD framework for steering students-AI dialogue, with its classroom pilot - [[reichert-human-centered-llm-chatbot-design-teachers-2026]] — Teacher-designed safety layers: domain boundaries, filtering, and override - [[ssail-safe-sound-ai-learning-2026]] — SSAIL: A Design Framework for Safe and Sound AI for Learning - [[turano-ai-tutoring-not-a-monolith-2026]] — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief) - [[eduzone-llm-safety-k12]] - [[eduguard-safe-rag-llm-tutor]] - [[hazra-safetutors-pedagogical-safety-2026]] - [[paternalistic-filter-llm-history-education]] - [[llm-unlearning-math-privacy]] - [[llm-children-reading-story-generation]] - [[llm-student-simulation-misconception-faithfulness]] - [[ai-tutor-authoring-promptdecipher]] - [[pedagogical-safety-rl]] - [[singh-eduqwen-pedagogical-rl-2026]] - [[tact-pedagogically-adaptive-esl-tutoring]] - [[residencyrl-clinical-rl-training-2026]] - [[eduframetrap-llm-sycophancy-educational-safety]] - [[favero-critical-ai-tutors-empower-enslave-2025]] - [[sec-ai-literacy-narrative-review-2026]] --- # FAQs ## [How Can We Address Common Misconceptions About AI in Education?](https://edtechdev.github.io/aied/faqs/addressing-common-misconceptions-ai-education/) This FAQ is organized by stakeholder group and uses a **refutation approach**: name the misconception, explain why it may seem plausible, reject the inaccurate belief directly, and replace it with a more useful mental model. It is based on the AI in Education knowledge base, especially its syntheses of [[misconceptions|Misconceptions about AI]], [[ai-literacy|AI Literacy]], [[cognitive-offloading|Cognitive Offloading]], and [[refutation-text|Refutation Text]]. The central message is not that AI is inherently beneficial or harmful. Its educational effects depend on **who uses it, for what task, what thinking the AI performs, what responsibility remains with the human, and how learning is evaluated**. --- ## FAQ for Students and Learners ### “If an AI answer sounds confident and detailed, why shouldn’t I trust it?” **Answer:** Because confidence and fluency are features of the output, not evidence that the answer has been verified. Generative AI predicts plausible language; it does not automatically check every claim against reliable evidence. It can invent sources, misstate facts, overlook context, or confidently repeat a misconception. Treat an AI answer as a **provisional draft or hypothesis**, not as an authority. Identify the claims on which the answer depends, inspect the original sources, check calculations, and compare the answer with course materials or trusted references. A useful test is: *Would I accept this claim if an unknown person said it without showing evidence?* If not, do not lower the standard simply because the prose sounds polished. See [[misconceptions|Misconceptions about AI]], [[hallucination-risk|Hallucination Risk]], and [[trust-calibration|Trust Calibration]]. --- ### “Does AI understand me and know what I mean?” **Answer:** Not in the way another person understands you. AI can respond to your language, use information in the current conversation, and sometimes retain information through product features. That can make the interaction feel personal. But the model does not possess human intention, lived experience, care, or contextual understanding. This distinction matters because AI may agree with you simply because your prompt suggests a preferred answer. This is sometimes called **sycophancy**: the system mirrors or validates the user rather than providing needed correction. Ask the AI to identify weaknesses in your reasoning, offer counterevidence, and explain what would make its answer wrong. Then verify the response independently. Agreement from AI is not proof that your position is correct. See [[misconceptions|Misconceptions about AI]], [[ai-sycophancy|AI Sycophancy]], and [[ai-literacy|AI Literacy]]. --- ### “If AI helped me create a good assignment, doesn’t that mean I learned the material?” **Answer:** Not necessarily. A good product shows what the **human–AI system** produced. It does not automatically show what you can explain, remember, adapt, or do independently. In one field experiment involving nearly 1,000 high-school [[math-education|mathematics]] students, unrestricted generative-AI access improved performance during assisted practice but reduced later unassisted exam performance. A guardrailed version that supplied hints rather than complete answers eliminated the observed [[ai-misuse-learning-harm|learning harm]]. The lesson is not that all AI use damages learning. It is that **assisted performance and durable learning are different outcomes**. After using AI, check whether you can: * explain the reasoning without looking at the AI response; * solve a similar problem independently; * identify weaknesses in the generated answer; * transfer the idea to a new context. See [[generative-ai-guardrails-harm-learning|Generative AI Without Guardrails Can Harm Learning]] and [[cognitive-offloading|Cognitive Offloading]]. --- ### “If AI lets me finish faster, isn’t that simply more efficient learning?” **Answer:** Faster completion is not always faster learning. AI can productively remove clerical work, confusing formatting, or unnecessary repetition. It can also remove the retrieval, planning, drafting, debugging, and revision through which knowledge and skill are developed. The key distinction is between **supportive offloading** and **substitutive offloading**: * Supportive offloading frees attention for more important thinking. * Substitutive offloading allows the AI to perform the thinking you were meant to learn. A useful sequence is: 1. Make an initial attempt. 2. Consult AI for feedback, hints, examples, or comparison. 3. Revise using your own judgment. 4. Complete a brief unaided explanation or application. The goal is not to maximize difficulty. It is to preserve the cognitive work that produces the intended learning. See [[cognitive-offloading|Cognitive Offloading]] and [[reducing-ai-misuse|Reducing AI Misuse]]. --- ### “Is any use of AI cheating?” **Answer:** No. But the opposite claim—“AI use cannot be cheating because I did not copy a person”—is also incorrect. Academic integrity depends on the purpose of the assignment, the instructor’s rules, the degree of AI involvement, attribution, and whether AI replaced the capability being assessed. AI might be permitted for brainstorming in one assignment, required for critique in another, and prohibited during an assessment of independent performance. Before using AI, ask: * What is this assignment intended to show that I can do? * Which forms of assistance are permitted? * Am I still the author and decision-maker? * Can I explain and defend the submitted work? * Do I need to disclose how I used AI? When expectations are unclear, disclosure and task-specific clarification are safer than assuming either that all use is forbidden or that all use is acceptable. See [[academic-integrity|Academic Integrity]] and [[ai-use-disclosure|AI Use and Disclosure Statements]]. --- ### “Is using AI frequently the main problem?” **Answer:** Frequency alone does not determine whether AI use is educationally productive. A student might use AI frequently to compare explanations, generate practice problems, challenge reasoning, and receive feedback while remaining cognitively active. Another student might use it once to generate the central argument or solution that an assignment was designed to assess. The more useful question is: > **Which layer of thinking did I delegate, and can I still perform that thinking independently?** Delegating grammar correction is different from delegating the claims, evidence, reasoning, and counterarguments of an essay. The deeper the delegated cognitive layer, the greater the risk that the final product overstates your own capability. See [[cognitive-offloading|Cognitive Offloading]] and [[ai-literacy|AI Literacy]]. --- ### “Shouldn’t one good prompt give me the right answer?” **Answer:** No. Generative-AI output is prompt-sensitive and often non-deterministic. Small changes in wording, context, examples, or assumptions can produce substantially different responses. Iteration may improve an answer, but repeated generation is not the same as verification. Five similar responses can repeat the same mistaken assumption. Productive iteration therefore includes more than asking again. It includes: * clarifying the goal and constraints; * asking the model to expose its assumptions; * requesting alternative interpretations; * testing the response against evidence; * checking whether the answer remains valid when the problem changes. Prompting is a useful skill, but it does not remove the need for subject knowledge and critical judgment. A [[brunnstrom-ai-interaction-literacy-srl-2026|demonstration of a naive student using a chatbot on a take-home examination question]] shows how much interaction good use actually takes: the default output stayed "polished but pedagogically thin" at the multistructural level of the SOLO taxonomy, and reaching a usable learning loop required eight rounds of meta-level intervention—signaling overload, requesting simplification, narrowing scope. The authors name the capacity this demands **AI-interaction literacy**—steering, evaluating, and learning from iterative interaction with generative AI—and note the equity sting: because unguided use imposes an interaction-management skill that is unevenly distributed, generative AI "may be most beneficial to already advantaged students." See [[misconceptions|Misconceptions about AI]], [[prompt-engineering|Prompt Engineering]], and [[brunnstrom-ai-interaction-literacy-srl-2026|AI-interaction literacy]]. For the underlying self-[[regulation]] demands, see [[developing-ai-tutor|How Do We Develop an Effective AI Tutor?]]. --- ### “If an AI detector cannot identify my use, is there any real downside?” **Answer:** The most important question is not whether software detects the use. It is whether you can demonstrate the competence the submitted work claims to represent. Undetected outsourcing can still leave you unable to explain the work, answer follow-up questions, adapt it to a new problem, or perform when AI is unavailable. It can also create a growing gap between your grades and your actual capabilities. That gap may remain hidden until a later course, [[summative-assessment|examination]], internship, licensure process, or workplace task requires independent performance. Academic integrity is therefore not merely about avoiding punishment. It is also about ensuring that your credentials continue to represent what you can actually do. See [[academic-integrity|Academic Integrity]], [[assessment-validity|Assessment Validity]], and [[authentic-assessment|Authentic Assessment]]. --- #### Key message for students > **Use AI to extend your thinking, not to make your thinking unnecessary. A strong AI-assisted product should leave you more capable of explaining, evaluating, transferring, and reproducing the underlying work.** --- ## FAQ for Instructors and Faculty ### “Will AI eventually make teachers unnecessary?” **Answer:** AI may automate portions of teaching work, but automating tasks is not equivalent to replacing the educational function of teaching. AI can draft examples, produce preliminary materials, answer routine questions, and assist with feedback. Teachers remain responsible for interpreting learner needs, establishing relationships, creating intellectually and emotionally safe learning environments, contextualizing disciplinary knowledge, exercising [[ethics|ethical]] judgment, and deciding when an AI-generated response is inappropriate. The teacher’s role may shift from being the sole source of information toward being an **orchestrator, learning designer, disciplinary guide, and accountable human decision-maker**. That is a transformation of professional work, not its disappearance. See [[teacher-role|Teaching]], [[learning-design|Learning Design]], and [[teacher-ai-competency|Teacher AI Competency]]. --- ### “Can experienced instructors reliably recognize AI-generated student work?” **Answer:** Not reliably enough to treat intuition as proof. AI-generated writing can be edited, combined with human writing, translated, paraphrased, or produced through many different systems. Human judgments can also be affected by writing style, language background, disability, or expectations about what a particular student “should” sound like. AI-detection tools face related limitations and can produce both false positives and false negatives. A detector score may occasionally prompt closer review, but it should not substitute for a fair evidentiary process. A more defensible response is **learning verification**: ask students to explain their reasoning, discuss their sources, revise a passage, apply the idea to a new case, or show process evidence. This directly assesses what matters—the student’s understanding. See [[academic-integrity|Academic Integrity]] and [[ai-detection|AI Detection]]. --- ### “Is a blanket AI ban the safest and fairest policy for my course?” **Answer:** Not automatically. AI-free conditions are appropriate when independent performance is the capability being assessed—for example, during certain examinations, foundational practice, or professional competency checks. But a universal prohibition can drive use underground, make rules difficult to enforce consistently, and prevent students from developing the AI literacy they may need beyond the course. A clearer model is to define task-specific conditions: * **AI required:** Students must use and critically evaluate AI. * **AI permitted with disclosure:** AI may support designated stages. * **AI restricted:** Only specified functions are allowed. * **AI prohibited:** Assistance would invalidate the intended learning claim. Students are more likely to follow boundaries when the instructor explains **why** each condition exists and applies it consistently across the syllabus, assignment directions, feedback, and assessment. See [[framing-ai-use-for-students|Framing AI Use for Students]], [[academic-integrity|Academic Integrity]], and [[educational-policy-ai|Educational AI Policy]]. --- ### “If I give students access to a powerful AI tutor, won’t learning improve?” **Answer:** Access alone is not an instructional design. The same underlying model can support or undermine learning depending on how the interaction is structured. An AI tutor may support learning when it: * requires an initial student attempt; * provides hints rather than complete solutions; * asks students to explain their reasoning; * adapts support without removing responsibility; * corrects misconceptions carefully; * fades assistance over time; * includes an unaided check. The same system may undermine learning when it immediately supplies polished answers, performs the planning, or encourages answer-seeking rather than understanding. The educational value lies not only in the model, but in the **[[pedagogy|pedagogical]] wrapper** around it. See [[learning-design|Learning Design]], [[scaffolding]], and [[reducing-ai-misuse|Reducing AI Misuse]]. --- ### “If students like AI-generated feedback, doesn’t that show the feedback is effective?” **Answer:** Satisfaction is useful evidence about acceptability, but it is not sufficient evidence of learning effectiveness. Students may prefer feedback that is immediate, encouraging, detailed, or easy to follow. Yet AI feedback may still be inaccurate, generic, overly positive, poorly prioritized, insensitive to context, or misaligned with the assignment’s learning objectives. Evaluate AI feedback through multiple questions: * Is it accurate? * Does it diagnose the actual problem? * Is it specific and actionable? * Is it appropriate for the learner’s level? * Does the student use it productively? * Does revision improve? * Does later independent performance improve? Students also need **feedback literacy**: the capacity to interpret, evaluate, and selectively act on feedback rather than accepting it automatically. See [[ai-feedback-quality|AI Feedback Quality]] and [[feedback-literacy|Feedback Literacy]]. --- ### “Is learning to write effective prompts enough preparation for instructors?” **Answer:** Prompting is one operational skill, not the full scope of teacher AI competency. Educators also need to understand: * what [[ai-technologies|AI systems]] can and cannot reliably do; * how to evaluate output accuracy and bias; * how AI affects assessment validity; * when student offloading becomes learning displacement; * privacy, accessibility, and data-governance requirements; * how to align AI use with disciplinary pedagogy; * when not to use AI. Research summarized in the wiki found that teachers substantially overestimated their AI competence when self-reports were compared with performance-based measures. Demonstrated competence was much more strongly related to classroom integration than confidence alone. See [[ai-literacy-assessment-misalignment|AI Literacy Assessment: Self-Reported vs. Performance Misalignment]], [[teacher-ai-competency|Teacher AI Competency]], and [[educational-development|Educational Development]]. --- ### “Does AI necessarily destroy critical thinking?” **Answer:** No. AI can either replace critical thinking or become an object and partner for critical thinking. A task is more likely to weaken [[student-engagement|engagement]] when students ask AI for a finished interpretation, argument, or solution and then accept it. A task can strengthen evaluation and metacognition when students must: * predict before consulting AI; * compare their reasoning with the AI’s response; * locate errors or unsupported claims; * improve a weak AI-generated answer; * select among alternatives and justify the choice; * explain why they rejected the AI’s recommendation. The appropriate distinction is not simply **AI versus no AI**. It is whether the AI functions as a **coach, challenge, or source for evaluation** rather than a substitute for the learner’s reasoning. See [[critical-thinking|Critical Thinking]], [[cognitive-offloading|Cognitive Offloading]], and [[learning-design|Learning Design]]. --- #### Key message for instructors > **Do not ask only, “May students use AI?” Ask, “What thinking must students retain, what support may AI provide, and what evidence will demonstrate that learning occurred?”** --- ## FAQ for Administrators, Institutional Leaders, and Policymakers ### “Will purchasing an advanced AI platform transform teaching and learning?” **Answer:** A platform provides capabilities, not educational transformation. Meaningful change requires alignment among [[curriculum-design|curriculum]], assessment, faculty development, technical support, accessibility, privacy, governance, workload, and local evaluation. Without those conditions, institutions may acquire a sophisticated system that is used inconsistently, duplicates existing work, increases faculty burden, or produces impressive demonstrations without measurable learning gains. Before procurement, leaders should specify: * the educational problem being addressed; * the intended users and use cases; * the outcomes that will count as success; * the data the system will collect; * the human oversight required; * the conditions under which the institution will modify or discontinue use. See [[administrator|AI from the Administrator Perspective]], [[governance|AI Governance]], and [[ai-ed-evaluation|AI Ed Evaluation]]. --- ### “Does a high benchmark score prove that an AI system is educationally effective?” **Answer:** No. A benchmark demonstrates performance under the benchmark’s specific conditions. It does not automatically demonstrate that students will learn more in real courses. A model may solve difficult problems, produce fluent explanations, or score well on a tutoring rubric while failing to improve retention, transfer, [[self-regulated-learning|self-regulation]], or equitable outcomes. Benchmark performance should therefore be separated from: * technical reliability; * pedagogical quality; * safety; * [[usability-research|usability]]; * implementation burden; * classroom adoption; * unassisted learning outcomes. Classroom effectiveness requires field testing with actual learners, relevant comparison conditions, appropriate outcome measures, and attention to implementation. See [[ai-ed-evaluation|AI Ed Evaluation]], [[benchmark]], and [[learning-gains|Learning Gains]]. --- ### “Because AI is data-driven, won’t it make decisions more objectively than people?” **Answer:** Data-driven does not mean value-free or unbiased. Bias can enter through training data, labels, outcome definitions, prompts, language assumptions, accessibility choices, decision thresholds, and the way staff interpret the output. Humans are also biased, but that is not evidence that automated decisions are neutral. Automation may conceal bias behind a technical interface and apply it at greater scale. For consequential educational decisions, institutions should require: * subgroup performance analyses; * documentation of training and validation conditions; * uncertainty reporting; * meaningful human review; * a student appeal process; * monitoring after deployment; * investigation of differential harms. See [[misconceptions|Misconceptions about AI]], [[bias-mitigation|Bias Mitigation]], and [[governance|AI Governance]]. --- ### “If every student receives the same AI account, haven’t we solved the equity problem?” **Answer:** Equal accounts do not guarantee equal opportunity or equal outcomes. Students differ in prior subject knowledge, AI experience, language, disability access, device quality, available time, confidence, and capacity to evaluate AI output. More experienced students may use AI to extend their learning, while students with weaker prior knowledge or [[metacognition|metacognitive]] skills may be more likely to accept incorrect output or delegate the practice they most need. Equity planning must therefore address at least three levels: 1. **Access:** Who can use the system reliably? 2. **Skills:** Who knows how to use and evaluate it? 3. **Outcomes:** Who actually benefits, and who experiences new harm? See [[equity-in-ai-education|Equity in AI Education]], [[digital-divide|Digital Divide]], and [[ai-literacy|AI Literacy]]. --- ### “Will one institution-wide AI policy eliminate uncertainty?” **Answer:** A policy is necessary, but policy text alone does not create shared understanding. Students and staff interpret AI expectations through several sources: institutional guidance, program norms, syllabi, assignment directions, instructor comments, peer behavior, and prior enforcement. When those sources conflict, people create their own explanations of what is acceptable. An effective policy architecture therefore connects: * institution-wide principles; * program- or discipline-level expectations; * course policies; * assignment-specific directions; * examples and scenarios; * transparent procedures for disclosure and review. Policy should also explain the educational rationale behind restrictions or permissions. Rules that merely state “allowed” or “prohibited” are less likely to produce informed judgment. See [[governance|AI Governance]], [[educational-policy-ai|Educational AI Policy]], and [[framing-ai-use-for-students|Framing AI Use for Students]]. --- ### “Can AI detection and remote proctoring solve the academic-integrity problem?” **Answer:** They cannot solve it by themselves. Detection estimates whether an artifact resembles machine-generated work. Education needs evidence that the learner possesses the claimed capability. Detection tools can produce false positives and false negatives, and their performance changes across models, languages, tasks, and editing practices. Proctoring may add privacy, accessibility, anxiety, and equity concerns without establishing what a student has learned. The evidence is now concrete enough to state numerically. A [[teichmann-detecting-undetectable-misconduct-2026|procedural-justice analysis]] reports that none of fourteen early detector tools reached 80% accuracy, that paraphrasing or light editing roughly halves already modest accuracy, and that detectors systematically misclassify non-native English writers because the features treated as AI signals also characterize competent second-language writing. In a covert field study, 94% of wholly AI-generated submissions injected into live online examinations across five psychology modules went undetected—and the AI work on average outscored real students. Vanderbilt University disabled its licensed detector after failing to validate an advertised 1% false-positive rate that implied roughly 750 mislabelled students among 75,000 annual submissions, redirecting staff toward transparent expectations and [[assessment|assessment redesign]]. A more durable institutional strategy combines: * clearly explained expectations; * appropriate AI-free assessment conditions; * process evidence and staged work; * oral or written learning verification; * task-specific disclosure; * assessment redesign; * proportionate, human-reviewed procedures. The goal is not merely to detect assistance. It is to preserve the validity of educational judgments. See [[academic-integrity|Academic Integrity]], [[ai-detection|AI Detection]], [[remote-proctoring|Remote Proctoring]], and [[teichmann-detecting-undetectable-misconduct-2026|undetectable misconduct]]. For the design response, see [[redesign-assessment-ai-era|How Should Assessment Be Redesigned for the AI Era?]] and [[reduce-ai-cheating|How Can We Reduce AI Cheating?]]. --- ### “Is faculty reluctance mainly a lack-of-training problem?” **Answer:** Sometimes, but faculty readiness is broader than technical skill. Reluctance may reflect workload, [[learner-identity|professional identity]], disciplinary values, concern about assessment validity, lack of institutional support, privacy uncertainty, or a reasoned judgment that a particular AI application does not serve students. Faculty development should therefore address: * knowledge and practical competence; * pedagogical integration; * professional identity and purpose; * time and workload; * policy and governance; * [[discipline-specific-aied|discipline-specific]] use; * opportunities for principled non-adoption. A one-time demonstration of AI features is unlikely to resolve a sociotechnical and professional change problem. See [[educational-development|Educational Development]], [[teacher-ai-competency|Teacher AI Competency]], and [[teacher-role|Teaching]]. --- #### Key message for institutional leaders > **Do not purchase an “AI outcome.” Build the institutional conditions under which a particular AI capability can be used responsibly, evaluated locally, improved when necessary, and discontinued when it does not serve learning.** --- ## FAQ for Instructional Designers, Educational-Technology Developers, and Vendors ### “If an AI tutor gives the correct answer, isn’t it a good tutor?” **Answer:** A system that solves a problem is not necessarily a system that teaches a learner. A technically correct answer may arrive too early, disclose too much, bypass [[desirable-difficulties|desirable difficulties]], or prevent the learner from practicing explanation and retrieval. A tutor should be evaluated by what it causes the student to **notice, attempt, explain, revise, and eventually do independently**. A pedagogically stronger tutor may: * diagnose before intervening; * ask questions rather than immediately answer; * provide the smallest useful hint; * require explanation; * respond to misconceptions; * fade support; * check later unaided performance. Correctness remains necessary, but educational quality also concerns timing, scaffolding, cognitive engagement, and [[transfer-of-learning|learning transfer]]. See [[learning-design|Learning Design]], [[intelligent-tutoring|Intelligent Tutoring Systems]], and [[pedagogical-safety|Pedagogical Safety]]. --- ### “Is more automation and personalization always better?” **Answer:** No. [[personalized-learning|Personalization]] can support learning, but it can also become over-accommodation. When a system performs the planning, monitors progress, decides what matters, and completes difficult steps, the learner may become more efficient while developing less agency and self-regulation. The design problem is not to minimize all difficulty. It is to remove unnecessary barriers while preserving the effort connected to the learning goal. Useful design features include: * mandatory learner attempts; * explanation prompts; * delayed hints; * adjustable assistance levels; * scaffold fading; * reflection on AI recommendations; * periodic unaided practice; * clear opportunities to override the system. See [[agentic-ai|Agentic AI]], [[agency]], and [[cognitive-offloading|Cognitive Offloading]]. --- ### “Is single-turn testing enough to establish that an educational chatbot is safe?” **Answer:** No. Educational harms can emerge cumulatively across an interaction. A system may respond appropriately to one isolated prompt but gradually begin supplying answers, reinforcing a misconception, encouraging dependency, or drifting from its intended tutoring role. The SafeTutors benchmark summarized in the wiki found that pedagogical-harm failures increased sharply when systems were evaluated across multiple turns rather than a single exchange. Benchmark evidence is not equivalent to a classroom learning-effect estimate, but it shows why educational testing should include: * sustained conversations; * repeated student errors; * attempts to obtain direct answers; * emotional and relational scenarios; * adversarial prompting; * changes in learner dependence over time. See [[pedagogical-safety|Pedagogical Safety]] and [[hazra-safetutors-pedagogical-safety-2026|AI Tutor Safety and Pedagogical Harms]]. --- ### “Will a larger or more capable model automatically make a safer tutor?” **Answer:** No. General model capability is not the same as pedagogical quality. A larger model may solve more difficult problems while still failing to: * select an appropriate instructional strategy; * recognize when to withhold an answer; * adapt to developmental level; * preserve [[productive-failure|productive failure]]; * communicate uncertainty; * avoid inappropriate emotional influence; * align with the instructor’s learning objectives. Pedagogical behavior should be explicitly designed, grounded in [[learning-theories|learning theory]], tested across learner groups, and monitored during sustained use. Model selection matters, but the instructional design layer remains essential. A [[reichert-human-centered-llm-chatbot-design-teachers-2026|participatory design study with six secondary teachers]] suggests safety comes from scope and oversight rather than scale: the teachers independently designed "bounded experts"—specialized capability confined to a strictly defined domain under human supervision—drawing two boundary lines (authority boundaries, because responsibility for student learning and safety cannot be delegated, and expertise boundaries, because AI lacks contextual knowledge of individual students and classroom norms) and three protective layers (domain boundaries, content filtering with standardized refusals, and teacher override). They asked for full conversation logging and real-time alerts rather than better model explanations. See [[learning-design|Learning Design]], [[pedagogical-llm-training|Pedagogical LLM Training]], [[pedagogical-safety|Pedagogical Safety]], and [[reichert-human-centered-llm-chatbot-design-teachers-2026|bounded-expert chatbot design]]. --- ### “Is more AI-generated feedback always better?” **Answer:** No. Feedback can become excessive, generic, mistimed, inaccurate, or cognitively overwhelming. Effective feedback should help a learner identify the most important next step. A long response that comments on every possible issue may be less useful than a focused intervention. Systems should prioritize feedback based on the learning goal, learner readiness, and likely impact. Evaluate more than feedback quantity and speed. Measure: * whether students understand the feedback; * whether they can judge its quality; * whether revision improves; * whether misconceptions decrease; * whether later independent performance improves. See [[ai-feedback-quality|AI Feedback Quality]], [[feedback]], and [[feedback-literacy|Feedback Literacy]]. --- ### “Is human review just a temporary requirement until models improve?” **Answer:** Human oversight is not merely an error-correction patch. It is also an accountability, contextualization, and governance function. Educators decide whether an output is appropriate for a particular learner, course, culture, or consequential decision. They interpret exceptions, consider information the model does not possess, and accept responsibility for actions affecting students. A meaningful human-in-the-loop design should specify: * who reviews the output; * what evidence the reviewer sees; * when review occurs; * how much time is available; * whether the reviewer can override the system; * who is accountable for the final action; * how a learner can appeal. A nominal human reviewer who lacks time, authority, or relevant information is not meaningful oversight. See [[human-in-the-loop-ai|Human-in-the-Loop AI]] and [[governance|AI Governance]]. --- ### “Can accessibility, privacy, and equity be added after the core product is working?” **Answer:** They should be treated as core design requirements, not post-launch additions. Input modality, reading level, language assumptions, device requirements, data retention, personalization, and model bias all shape who can use a system and who may be harmed by it. Retrofitting may improve the interface while leaving the underlying workflow, data model, and decision logic unchanged. Design teams should involve affected learners and educators early, test with diverse users, minimize data collection, provide accessible alternatives, and examine differential outcomes. A system cannot be considered educationally effective if its benefits are inaccessible or its harms are unevenly distributed. See [[accessibility]], [[universal-design-for-learning|Universal Design for Learning]], [[privacy]], and [[equity-in-ai-education|Equity in AI Education]]. --- #### Key message for designers and developers > **Optimize for growth in learner capability, not merely successful task completion. A tutoring system is educationally successful when learners become more capable—not permanently more dependent on the system.** --- ## FAQ for Educational Researchers and Evaluators ### “If students perform better while using AI, doesn’t that demonstrate learning?” **Answer:** No. It demonstrates assisted performance. Learning requires evidence that the learner’s capability changed. Studies should distinguish among: * performance while AI is available; * immediate unassisted performance; * delayed retention; * transfer to new problems; * explanation and strategy use; * dependence on continued assistance. Without an unaided measure, researchers may mistakenly attribute the AI system’s contribution to the learner. This is especially important when the tool can generate the solution, reasoning, or text that the outcome measure rewards. See [[learning-gains|Learning Gains]], [[assessment-validity|Assessment Validity]], and [[cognitive-offloading|Cognitive Offloading]]. --- ### “Are self-reported AI literacy, confidence, and learning adequate outcomes?” **Answer:** They are useful for understanding perception, acceptance, anxiety, and [[self-efficacy]], but they are not adequate measures of demonstrated competence. People can be confident and wrong, or skilled and underconfident. Pair self-reports with performance-based measures such as: * identifying errors in AI output; * verifying a source; * selecting an appropriate use strategy; * recognizing bias or sycophancy; * revising a flawed response; * explaining when AI should not be used; * calibrating confidence against accuracy. The wiki’s synthesis of teacher AI-literacy research reports a substantial gap between self-assessment and measured performance, reinforcing the need to evaluate both. See [[ai-literacy-assessment-misalignment|AI Literacy Assessment: Self-Reported vs. Performance Misalignment]] and [[ai-literacy|AI Literacy]]. --- ### “Can benchmark performance be treated as evidence of classroom efficacy?” **Answer:** Not without additional evidence. Benchmarks establish bounded technical or behavioral performance. Classroom learning depends on students, instructors, incentives, curriculum alignment, implementation quality, uptake, and competing resources. A responsible evidence pathway may move from: 1. technical and benchmark testing; 2. usability and safety studies; 3. small-scale classroom pilots; 4. controlled efficacy studies; 5. implementation research; 6. longer-term and multi-site evaluation. Researchers should state clearly which link in that chain a study addresses rather than generalizing a benchmark result into a claim about learning. See [[benchmark]], [[ai-ed-evaluation|AI Ed Evaluation]], and [[limitations-in-aied-research|Limitations of the AIED Evidence Base]]. --- ### “If an AI scoring system is reliable, doesn’t that mean it is valid?” **Answer:** No. Reliability concerns consistency. Validity concerns whether the interpretation and use of the score are justified. A system can consistently measure the wrong construct, omit important dimensions, disadvantage a subgroup, or produce a score that humans misuse. Validation should examine: * construct representation; * comparison with relevant human judgments; * subgroup performance; * error patterns; * uncertainty; * consequences of use; * whether AI output changes human decisions; * appeal and review procedures. High agreement is one form of evidence. It is not a complete validity argument. A [[opraise-automated-marking-ai-assessment-2026|large UK benchmark]] shows the dissociation directly: across 761 authentic undergraduate Psychology essays, AI and human marks agreed on the degree band only 35–65% of the time (63% at one institution, 53% at a second, 35% at a third), while reliability was near-perfect (re-scoring intra-class correlations up to 1.00). The systems agreed with each other far more closely than with humans (three-model ICC = 0.91), concurring on the band for only 56% of submissions when all three models had to agree, and marks were compressed toward the middle (compression score 0.47–0.82)—so AI was least accurate exactly at the boundaries separating a First from an Upper Second or a pass from a fail. AI feedback was also three to eight times longer than the human average of 100–200 words: volume is not quality. See [[assessment-validity|Assessment Validity]], [[educational-measurement|Educational Measurement]], [[automated-assessment|Automated Assessment]], and [[opraise-automated-marking-ai-assessment-2026|automated marking of university essays]]. --- ### “Can LLM-generated or simulated students replace real learners in educational research?” **Answer:** They may be useful for prototyping, stress testing, generating scenarios, or exploring hypotheses. They should not be assumed to reproduce human learning processes without validation. A model can imitate the language of confusion or a misconception without displaying the persistence, motivation, prior knowledge, emotion, or developmental trajectory of a real learner. Research summarized in the wiki found that simulated students frequently abandoned an assigned misconception after minimal correction, raising doubts about whether they faithfully represented human conceptual change. Claims based on simulated learners should therefore be validated against human behavior before being used to support instructional or policy conclusions. See [[simulating-students|Simulating Students]] and [[llm-student-simulation-misconception-faithfulness|Simulating Students or Sycophantic Problem Solving?]]. --- ### “Does a positive average effect mean the intervention benefits students generally?” **Answer:** No. Average effects can conceal meaningful differences by prior knowledge, age, discipline, language, disability, metacognitive skill, access, instructor implementation, or type of AI use. Researchers should examine: * treatment-effect heterogeneity; * subgroup uncertainty rather than only subgroup point estimates; * implementation fidelity; * actual patterns of [[student-ai-interaction|AI interaction]]; * missing-data and attrition differences; * whether benefits persist without AI; * whether some learners gain while others become more dependent. A small average gain may hide a valuable effect for one group and harm for another. A large average gain may depend on conditions that other institutions cannot reproduce. See [[limitations-in-aied-research|Limitations of the AIED Evidence Base]], [[research-methods-aied|Research Methods in AIED]], and [[equity-in-ai-education|Equity in AI Education]]. --- #### Key message for researchers > **Measure the learner after the AI has stopped helping, and report the implementation conditions, learner differences, and validity limitations that determine what the result actually means.** --- ## FAQ for Parents, Families, the General Public, and Media Communicators ### “Will AI revolutionize education—or destroy it?” **Answer:** Both claims exaggerate the power of the technology acting on its own. AI can expand access to explanation, translation, practice, feedback, and assistive support. It can also introduce misinformation, privacy risks, bias, over-reliance, and new integrity problems. The consequences depend on how the system is designed, what teachers and students are asked to do with it, and how institutions govern its use. A more useful public question is: > **For which learners, tasks, and outcomes—and under what conditions—does this AI use produce more educational benefit than harm?** This question encourages evaluation rather than hype or panic. See [[ai-education|AI in Education]] and [[misconceptions|Misconceptions about AI]]. --- ### “Aren’t young people already AI literate because they are digital natives?” **Answer:** Familiarity with digital products is not the same as the ability to understand and critically evaluate AI. Students may be comfortable opening a chatbot, generating an image, or asking for an answer while remaining unable to: * verify a claim; * recognize fabricated evidence; * detect bias or sycophancy; * protect personal information; * decide what thinking should not be delegated; * explain how AI affected their work. The wiki summarizes research in which confidence with everyday technology did not translate into comparable competence in algorithmic reasoning, technological creation, or critical AI evaluation. AI literacy must be taught and demonstrated; it should not be inferred from age or frequency of technology use. See [[ai-literacy|AI Literacy]] and [[digital-literacy-illusion|The Illusion of Competence]]. --- ### “Are students who use AI simply lazy or dishonest?” **Answer:** Some students misuse AI, but moral labeling does not adequately explain or prevent that behavior. Student decisions are shaped by assignment value, time pressure, confidence, peer norms, policy clarity, fear of failure, prior access, and whether the work appears connected to meaningful learning. Students also reason differently about brainstorming, editing, explanation, and full-text generation. An effective response combines: * clear and consistent expectations; * meaningful assessment; * instruction in responsible use; * opportunities for disclosure; * verification of learning; * proportionate accountability. Treating all AI use as evidence of poor character can push use into secrecy and make honest discussion less likely. See [[academic-integrity|Academic Integrity]] and [[framing-ai-use-for-students|Framing AI Use for Students]]. --- ### “Is AI basically just another calculator?” **Answer:** The comparison is helpful in one respect: both can offload work. But generative AI can offload a much broader range of cognitive activity. A calculator typically performs a defined mathematical operation. Generative AI can produce explanations, arguments, plans, source summaries, code, feedback, and full assignments. Its operation is also less transparent, and its outputs may be persuasive while incorrect. That means educators must make more nuanced decisions about what students may delegate. Offloading routine computation may allow learners to focus on interpretation. Offloading the interpretation itself may remove the intended learning. See [[cognitive-offloading|Cognitive Offloading]] and [[generative-ai|Generative AI]]. --- ### “Is a chatbot safe for children as long as it blocks toxic or explicit content?” **Answer:** Content moderation is necessary, but it does not cover all educational risks. A system may remain polite while: * giving answers too quickly; * reinforcing misconceptions; * encouraging emotional dependence; * collecting inappropriate data; * offering developmentally unsuitable advice; * making inaccessible assumptions; * displacing human support; * reducing productive effort. Child-facing AI should be evaluated for content safety, pedagogical safety, privacy, accessibility, relational influence, and the effects of repeated interaction—not only for prohibited words or topics. See [[k-12|K–12 AI Education]], [[pedagogical-safety|Pedagogical Safety]], and [[privacy]]. --- ### “Does an empathetic AI actually care about the student?” **Answer:** AI can generate language that sounds attentive, supportive, or emotionally responsive. That may sometimes help a learner articulate a problem or continue with a low-stakes task. But the appearance of empathy should not be confused with human care, responsibility, or duty of care. AI cannot independently assume the responsibilities of a teacher, counselor, parent, or caregiver. It may misunderstand the situation, reinforce the user’s framing, or respond inappropriately while sounding compassionate. Students should know when they are interacting with AI, what data may be retained, and when the system should direct them toward a qualified human. See [[misconceptions|Misconceptions about AI]], [[conversational-ai|Conversational AI]], and [[governance|AI Governance]]. --- #### Key message for families and the public > **AI is not an autonomous educational force. Its consequences are shaped by design, teaching, institutional decisions, family support, and the responsibilities learners retain.** --- ## FAQ for Employers and Workforce Partners ### “Does AI literacy mainly mean knowing how to write good prompts?” **Answer:** Prompting is useful, but durable AI literacy is much broader. A capable employee must be able to: * define the problem appropriately; * decide what should and should not be delegated; * provide relevant context without exposing sensitive data; * evaluate evidence and uncertainty; * identify bias and failure; * revise or reject output; * document consequential use; * remain accountable for the final decision. Prompt techniques will change as products evolve. Judgment, verification, domain understanding, ethical reasoning, and responsibility are more transferable capabilities. See [[ai-literacy|AI Literacy]] and [[human-ai-collaboration|Human–AI Collaboration]]. --- ### “Does a polished AI-assisted work product demonstrate professional competence?” **Answer:** It demonstrates the performance of a human–AI system, but it may not show what the individual can do. To assess professional competence, employers and educators should examine whether the person can: * frame the underlying problem; * explain the assumptions; * verify the evidence; * detect subtle errors; * adapt when conditions change; * defend the final recommendation; * perform essential judgment without inappropriate assistance. In many professions, responsible AI use is itself a legitimate competency. But the assessment must distinguish **effective tool use** from the appearance of expertise created by the tool. See [[authentic-assessment|Authentic Assessment]], [[assessment-validity|Assessment Validity]], and [[career-development-and-readiness|Career Development and Readiness]]. --- ### “Is foundational disciplinary knowledge becoming obsolete because AI can retrieve or generate it?” **Answer:** No. The ability to oversee AI depends on the expertise that excessive automation may discourage people from developing. Without sufficient domain knowledge, a user may not recognize: * a plausible but incorrect conclusion; * a missing constraint; * a dangerous recommendation; * an invalid comparison; * a fabricated citation; * a biased assumption; * a situation in which AI should not be trusted. Curricula may need to reconsider which knowledge should be memorized and which tools should be available. But they should not eliminate foundational understanding simply because AI can produce an answer. Expert oversight requires an internal basis for judgment. See [[cognitive-offloading|Cognitive Offloading]], [[prior-knowledge|Prior Knowledge]], and [[trust-calibration|Trust Calibration]]. --- ### “Does teaching critical and ethical AI use conflict with workplace productivity?” **Answer:** Responsible evaluation is part of sustainable productivity. Unverified automation can create rework, security incidents, discriminatory decisions, legal exposure, reputational damage, and false confidence. Critical AI literacy does not mean rejecting automation. It means knowing when automation adds value, when supervision is required, and when the output should be rejected. The strongest graduate or employee is not necessarily the person who uses AI for the greatest number of tasks. It is the person who can allocate work intelligently between humans and AI while maintaining quality, confidentiality, accountability, and professional judgment. See [[ai-literacy|AI Literacy]], [[governance|AI Governance]], and [[human-ai-collaboration|Human–AI Collaboration]]. --- #### Key message for employers > **Evaluate whether people can use AI with judgment, verification, and accountability—not simply whether they can produce polished work quickly.** --- ## FAQ for Anyone Communicating About AI Misconceptions ### “Why isn’t it enough to tell people the correct facts about AI?” **Answer:** Misconceptions are not always gaps in knowledge. They are often stable and plausible mental models. A person may repeatedly observe that AI sounds confident, produces high-quality work, or saves time. Those experiences appear to confirm beliefs such as “AI understands,” “AI is accurate,” or “finishing the task means I learned it.” Effective correction should therefore do more than state a fact. It should: 1. name the misconception; 2. acknowledge why it seems plausible; 3. reject it clearly; 4. explain why it fails; 5. provide a replacement model; 6. give the person a way to apply the replacement. See [[misconceptions|Misconceptions about AI]] and [[refutation-text|Refutation Text]]. --- ### “What does an effective refutation sound like?” **Answer:** It should be direct without being dismissive. For example: > **Misconception:** “A good AI-generated assignment shows that the student learned.” > **Refutation:** “That is not necessarily true. The product shows what the student and AI produced together, but it does not reveal which reasoning the student performed.” > **Replacement:** “Learning is better demonstrated when the student can explain, transfer, adapt, and reproduce the capability.” > **Application:** “Follow AI-assisted work with an explanation, oral defense, process record, or unaided application.” The replacement model is essential. If communicators only remove the misconception, people may return to it because they lack a better explanation. See [[refutation-text|Refutation Text]]. --- ### “How can educators correct misconceptions without shaming people?” **Answer:** Address the belief and its consequences rather than labeling the person. Avoid messages such as “Only lazy students use AI” or “Anyone who trusts a chatbot is foolish.” These claims threaten identity and encourage defensiveness or concealment. A more productive approach is: * acknowledge why the belief seems reasonable; * demonstrate a discrepancy, such as a confident AI error; * invite prediction before revealing the correction; * let participants compare assisted and unassisted performance; * give them a strategy for future situations; * reinforce the new model across multiple activities. Misconceptions are easier to reconsider when people can revise their thinking without being treated as unintelligent or unethical. See [[framing-ai-use-for-students|Framing AI Use for Students]], [[ai-literacy|AI Literacy]], and [[refutation-text|Refutation Text]]. --- ### “What experiences are most likely to change an inaccurate belief about AI?” **Answer:** Experiences that make the misconception’s failure visible. Examples include: * asking learners to identify a confident fabricated citation; * comparing several contradictory answers to the same prompt; * completing an unaided problem after AI-supported practice; * examining how an AI mirrors an incorrect assumption; * comparing a direct-answer tutor with a hint-based tutor; * auditing outputs for bias across names, dialects, or scenarios; * asking participants to defend an AI-generated recommendation using original evidence. These activities transform abstract warnings into observable evidence. They also allow participants to practice verification, trust calibration, and decisions about cognitive offloading. See [[ai-literacy|AI Literacy]], [[trust-calibration|Trust Calibration]], and [[cognitive-offloading|Cognitive Offloading]]. --- ### “What single idea should stakeholders remember?” **Answer:** > **AI is a fallible cognitive resource—not an authority, a human mind, or evidence that learning has occurred. Responsible educational use keeps humans accountable, preserves the thinking needed for learning, verifies consequential outputs, and judges success through durable and equitable capability rather than fluency, speed, engagement, or task completion alone.** --- ## [How Can AI Agents Support Students and Instructors?](https://edtechdev.github.io/aied/faqs/ai-agents-support-students-instructors/) # How Can AI Agents Support Students and Instructors? **AI agents can move beyond single-turn question answering** by planning, using tools, remembering relevant context, coordinating subtasks, and adapting support over a sequence of interactions. For students, plausible roles include adaptive tutoring, study planning, [[formative-assessment|formative]] feedback, guided [[problem-solving|problem solving]], practice generation, simulation, prerequisite recommendations, and reflective or metacognitive [[prompt-engineering|prompting]]. For instructors, agents can assist with material development, [[automated-question-generation|question generation]] and validation, feedback triage, course [[learning-analytics|analytics]], instructional-design workflows, resource retrieval, and the orchestration of specialized agents. ## Recurring agentic capabilities The article [[agentic-workflows-education|Agentic Workflows in Education]] describes four recurring agentic capabilities: **reflection, planning, tool use, and multi-agent collaboration**. Each adds possibilities but also introduces [[explainable-ai|interpretability]], coordination, trust, latency, and oversight challenges. ## What AI agents can do well (positive implications) - **Sustained, adaptive support.** Unlike single-turn [[conversational-ai|chatbots]], agents can maintain a learning conversation over many turns — remembering what a learner knows, adapting difficulty, and sequencing multi-step [[scaffolding]]. This supports [[adaptive-learning|adaptive]] and [[personalized-learning|personalized]] learning at scale. - **Unburdening instructors.** Agents can draft materials, generate and validate questions (e.g., a generator + validator pairing), triage feedback, and orchestrate specialized sub-agents, freeing teachers for higher-value interaction. - **Rich interaction and [[desirable-difficulties|productive friction]].** Multi-agent classrooms and simulated peers create varied dynamics — peer-like discourse, constructive disagreement, role-play — that support [[collaborative-learning|collaborative learning]] and [[socratic-method|Socratic-style probing]]. Agents designed to challenge rather than agree can push learners toward deeper reconsideration (constructive-conflict agents improved design outcomes in research). - **Low-risk practice and simulation.** Agent-based [[simulation|simulations]] ([[simulating-students|simulated students]], [[medical-education|clinical]] scenarios) let learners practice in safe, repeatable environments before real-world application. ## Key risks and caveats (negative implications) - **Over-automation can hollow out learning.** The more an agent automates, the less cognitive work the learner does. Proactive agents can leave students as passive consumers, weakening the effortful processes that build durable learning and raising [[cognitive-offloading|over-reliance]] risk. - **Reduced metacognitive engagement.** If agents handle planning and monitoring, learners may not develop the [[metacognition]] and [[self-regulated-learning|self-regulation]] that education aims to build. Agents should elicit, not replace, these processes. - **Misplaced trust and verification gaps.** Autonomous agents can produce plausible but unvalidated output; learners and teachers may [[trust-calibration|over-trust]] it. Robust verification and [[ai-literacy]] become more important as agents gain [[agency|autonomy]]. - **Opacity and accountability.** Multi-agent systems complicate [[human-in-the-loop-ai|human oversight]] — which agent is accountable for an error, and where does a human intervene? Coordination failures and persona drift can undermine reliability and [[pedagogical-safety]]. - **[[equity-in-ai-education|Equity]] and bias.** Agents can reproduce training-data bias at scale, and unequal access to capable agentic systems can widen inequity. ## What the evidence shows so far Concrete results — and cautions — are accumulating. [[wang-tutor-copilot-human-ai-live-tutoring-rct-2024|Tutor CoPilot]] — the first [[rct|randomized controlled trial]] of a human-AI system in live tutoring — gave novice tutors real-time expert-like guidance drawn from experienced tutors' reasoning: across **900 tutors and ~1,800 students**, students of tutors with access were **4 percentage points more likely to master topics**, rising to **9 p.p.** among the lowest-rated tutors, whose students caught up to those of higher-rated tutors in the control group. It cost about **\$20 per tutor annually**, and analysis of **550,000+ tutoring messages** showed tutors shifting toward asking guiding questions rather than giving away answers — a clear case of AI augmenting the [[teacher-role|teacher]] rather than replacing them. But "agentic" is not automatically better. [[ilieva-agentic-genai-higher-education-2026|A study with 130 students in an e-commerce course]] found both [[generative-ai]] chatbots and GAI agents rated above traditional [[online-teaching-and-learning|e-learning]] on learning enhancement and [[personalized-learning|personalization]] — yet **no statistically significant difference between the chatbot and the agent conditions**, so added autonomy did not translate into added learning value. The proposed Agentic GAI-Supported Learning Framework accordingly treats agents as **bounded, human-supervised learning partners**, with goals, checkpoints and final decisions reserved for humans. Two further cautions matter. First, **withdrawal**: randomized trials show that even brief AI assistance can depress subsequent unassisted performance, and the post-withdrawal interval — named [[cognitive-washout-ai-skill-decay-2026|cognitive washout]] — is almost entirely unmeasured, so the durability of agent-assisted [[learning-gains|learning gains]] is unknown. Second, **evaluation**: [[zhang-platform-scores-miss-ai-teaching-agents-2026|deploying eight AI teaching agents across a medical curriculum]], platform-generated scores ranked agents differently from an independent expert rubric (the platform's third-ranked agent came last on teaching quality) because platform scores index student performance rather than [[pedagogical-agent|agent]] teaching quality. The [[governance]] review [[beyond-agent-label-agentic-ai-governance-2026|Beyond the Agent Label]] adds a proportionality rule — **autonomy should not exceed the maturity of the evidence or the strength of accountable human control** — noting that evidence is strongest for artifact-level outcomes and weakest for durable learning and equity. For how to build one, see [[developing-ai-tutor]]; for how to evaluate whether it works, see [[evaluating-ai-interventions-methods]]. ## The state of the evidence The evidence base is still emerging. The knowledge base's [[agentic-ai|Agentic AI in Education]] synthesis draws on a [[meta-analysis-systematic-review|scoping review]] of 474 studies but notes substantial concentration in [[higher-ed|higher education]], [[stem-education|STEM]], short-term designs, and text-based tutoring; only a minority of the reviewed work explicitly grounded its systems in educational theory, and rigorous long-term classroom validation remains limited. The key design warning is therefore not to equate greater autonomy with better learning. [[agentic-ai-pedagogical-best-practice-2026|Agentic AI and Pedagogical Best Practice]] recommends intentional friction, dynamic scaffolding, and human oversight so that agent initiative does not remove the learner's own planning, monitoring, judgment, and effort. See also [[intelligent-tutoring|Intelligent Tutoring]] and [[human-in-the-loop-ai|Human-in-the-Loop AI]]. --- ## [How Does AI Affect Student Anxiety and Well-Being, and What Can We Do?](https://edtechdev.github.io/aied/faqs/ai-anxiety-wellbeing/) # How Does AI Affect Student Anxiety and Well-Being, and What Can We Do? **The situation you are likely in.** A student tells you, in office hours or on an intake form, that finishing the degree no longer seems worth it because AI can already do the job they were training for. Another goes quiet, submits less, stops asking questions. You want to help, you are unsure whether this needs a referral or a strategy, and you have ten minutes to decide. **The bottom line.** AI anxiety is a measurable emotional response, most of its drivers sit in career and competence rather than in the technology, and reassurance does not move it. What moves it is mastery experiences, evidence that AI-related work is worth the effort, and real career adaptability. A few students need a referral rather than a strategy — and the evidence tells you less about that boundary than you would like. **What this page is.** The practitioner page: what to do. The concept treatment of the emotion itself — proctoring stress, the productive side of anxiety, educator anxiety — is at [[anxiety-and-stress]]; the positive-state framing across emotional, psychological and social dimensions is at [[well-being]]. ## The short version - **Name the driver before you respond.** Career fear and competence doubt need different answers; they carry most of the association with anxiety. - **Do not lead with reassurance.** Encouragement that does not change what a student believes about their capability will not move the emotion. - **Teach scrutiny, not avoidance.** Anxious students who learn to evaluate AI output become deliberate users, not avoidant ones. - **Build career adaptability, not confidence alone.** Self-belief alone did not buffer the career harm in the study that tested it. - **Know your referral route before you need it,** and watch for hopelessness about the degree, not just tool frustration. - **Be specific about institutional expectations,** and never deploy a well-being tool whose error handling, data storage and escalation path you cannot explain in a sentence. ## What to watch for: ordinary adjustment, or something that needs more [[kim-ai-anxiety-comprehensive-analysis|Kim et al.]] define AI anxiety as apprehension or fear produced by the acceleration of AI, distinguish it from older automation anxiety because AI threatens cognitive and professional roles rather than manual tasks, and name the **fear of replacement by AI** as its primary contributor — with uncontrolled AI growth, [[privacy]] concerns, AI-generated misinformation and AI bias as secondary causes. Read that against what you see: the student worried about being replaced is worried about the labor market they are entering, not the tool you demonstrated last week. That is why career and job-market anxiety is the best-evidenced source. [[ustun-ai-anxiety-job-finding-anxiety-2026|Üstün and Danacıoğlu (2026)]] surveyed 1,057 students across 35 Turkish universities: higher AI anxiety and more negative attitudes toward AI went with higher post-graduation job-finding anxiety, with female, social-science and second-year students reporting higher anxiety. [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al. (2026)]] found a moderate positive correlation between AI anxiety and job-search anxiety among 821 health-sciences students (r = 0.233, p < 0.001), with AI anxiety still a significant predictor after controlling for socio-demographics (β = 0.234, p < 0.001). [[duan-ai-anxiety-career-decisions-college-2026|Duan et al. (2026)]] extended this to consequences: in a structural equation model of 315 Chinese college students, AI anxiety predicted poorer career decisions directly and indirectly by undermining career adaptability, a mediated pathway accounting for 63.35% of the total effect; their moderation test found [[self-efficacy]] did not buffer the harm. Ordinary adjustment sounds like *"I don't know how to use this well yet."* Act on it when it sounds like *"there is no point in me finishing this,"* when the student withdraws rather than complains, or when the worry is carried alone because raising it feels like admitting weakness — the stigma-driven disclosure pattern [[culturally-aware-student-stress-chatbot-2026|Bashir and Afzal (2026)]] describe for Pakistani university students. ## What to do in class **Competence worries work through appraisals, not reassurance.** A two-wave survey of 547 Chinese undergraduates modeled [[anxiety-and-stress|AI learning anxiety]] within control-value theory ([[school-support-ai-learning-anxiety-control-value-2026|Jiang, Chen and Chen, 2026]]): anxiety is generated by two appraisals — whether a student feels capable of handling AI-related learning tasks (control, measured as AI learning self-efficacy) and whether they judge AI useful for academic work (value, via the [[technology-acceptance-model|perceived-usefulness]] construct). Self-efficacy had the stronger negative association with anxiety and also predicted perceived usefulness; perceived school support was associated with lower anxiety even with both appraisals modeled, but the indirect routes through them carried most of that association, leaving roughly a third as a direct link. Modeled separately, only *informational* support retained a reliable path to self-efficacy — which is why [[ai-literacy|AI literacy]] training, responsible-use guidelines and accessible technical support matter. The design follows: give students mastery experiences and concrete evidence that AI is worth the effort — low-stakes practice, guided evaluation of AI outputs, feedback-driven revision — and offer support that is informational rather than purely comforting. Perceived support precedes favorable appraisals. **Teach scrutiny rather than avoidance.** Anxiety and engagement can move together, which changes what you do. In a survey of 107 students, [[ai-anxiety-strategic-regulation-writing-2026|Kim (2026)]] finds higher AI anxiety positively associated with verification and revision behavior (β = .24, p < .01) and evaluative capacity predicting active engagement (β = .46, p < .001), sorting students into four regulatory types from uncritical reliance (18.7%) through selective integration (34.6%) and evaluative transformation (31.8%) to strategic rejection (14.9%). The implication runs against anxiety reduction as the goal: teaching scrutiny and evaluative judgment turns anxious students into more deliberate users of [[generative-ai|generative AI]] rather than more avoidant ones, which connects this page to [[verify-ai-output|teaching students to verify AI output]] and to [[ai-literacy]]. **Build career adaptability rather than confidence alone.** [[duan-ai-anxiety-career-decisions-college-2026|Duan et al. (2026)]] conclude that career adaptability is the key protective mechanism and recommend universalizing AI literacy and career-planning courses alongside industry–education integration. [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al. (2026)]] reach the same conclusion from health sciences: enhanced programs plus AI literacy and career counseling, because AI anxiety shapes students' professional futures rather than their attitudes toward a tool. **Pair AI instruction with emotion regulation and metacognitive support.** [[zhang-ai-anxiety-academic-motivation-emotion-2026|Zhang, Shi and Lu (2026)]] argue that reducing AI anxiety rather than expanding tool access is what sustains motivation, and recommend integrating emotion-regulation and [[metacognition|metacognitive]] support into AI literacy education with gender-aware differentiation — while noting that their design cannot show causation. **Decide about well-being tools deliberately.** [[culturally-aware-student-stress-chatbot-2026|Bashir and Afzal (2026)]] trained a Random Forest stress classifier on 1,100 survey responses across 20 features, reaching 89.09% accuracy over three severity levels, paired with a tiered Stepped Care dialogue layer in English, Urdu and Roman Urdu. Their [[explainable-ai|feature-importance]] analysis placed blood pressure first (15.6%) and **teacher-student relationship second (10.0%)**, with anxiety level ninth (4.8%) — evidence, the authors argue, that student distress is multi-dimensional rather than driven by one indicator. Note what they concede: a [[machine-learning|trained classifier]] decides the support tier, so what happens on a harmful misclassification, how disclosures are stored and what [[privacy]] protections apply, and how escalation to human counseling works are all left open; the boundary between [[human-in-the-loop-ai|human oversight]] and automated encouragement remains a design question, and the system is not presented as clinical or therapeutic. ## What to say, and what not to say **Do not say:** - *"AI can't really do your job."* You may be wrong, and the student has already tested it. - *"Everyone feels like this."* True and useless; it makes a specific fear generic. - *"You just need to learn to use it properly."* This turns a career worry into a competence verdict — precisely the appraisal that drives the anxiety. - *"It'll be fine."* A prediction you cannot support and a plan they cannot act on. - *"You need to be more resilient."* Self-belief alone did not buffer the career harm; see the objections below. **Do say:** - Name the driver out loud: *"What worries me isn't that you can't use the tool. It's that you don't know what it means for the job. Those are different problems and I can help with both."* - Separate the tool from the person: capability is built, not possessed. - State your expectations for AI use plainly. Silence reads as prohibition to some students and abandonment to others. - Watch your differentiation: gender moderated the regulation–motivation link in Zhang, Shi and Lu's data, stronger among male students, so one script will not land the same way in every group. ## When to refer Refer when the worry has stopped being about a tool and started being about the student's worth, their place in the program, or their future: hopelessness about finishing, withdrawal from work and from people, statements that they see no professional future. Anything suggesting risk to safety is an immediate referral and a matter for your institution's protocol, not a conversation you manage alone. Career anxiety belongs at [[career-development-and-readiness|career services]], because it is a counseling matter rather than a morale problem, and [[duan-ai-anxiety-career-decisions-college-2026|Duan et al. (2026)]] found that self-efficacy alone did not buffer the damage to career decisions. Institutions are the level the evidence points at: [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al. (2026)]] recommend AI literacy and career counseling together with enhanced programs, not one instead of the other. And you have a lever the referral form does not: [[culturally-aware-student-stress-chatbot-2026|Bashir and Afzal (2026)]] rank the teacher-student relationship second in their feature-importance analysis (10.0%), ahead of sleep quality (9.3%), depression (8.3%) and social support (7.6%). The authors treat that as a hypothesis needing locally collected data, but it is a reason to make sure students know you are a route to help. Do not route a distressed student into a student-facing AI tool because it is always available. Such systems appeal to students who do not raise distress with parents, teachers or peers — the disclosure pattern [[culturally-aware-student-stress-chatbot-2026|Bashir and Afzal (2026)]] describe — but that conversational layer was checked with simulated inputs and informal usability testing rather than evaluated with students on cultural appropriateness, emotional safety or satisfaction, and it has no student outcome data at all. ## What the evidence actually shows, and where it stops Almost everything here about AI anxiety itself is correlational. The career-anxiety studies are cross-sectional surveys ([[ustun-ai-anxiety-job-finding-anxiety-2026|Üstün and Danacıoğlu, 2026]]; [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al., 2026]]; [[duan-ai-anxiety-career-decisions-college-2026|Duan et al., 2026]]), and the motivation study is cross-sectional and self-reported ([[zhang-ai-anxiety-academic-motivation-emotion-2026|Zhang, Shi and Lu, 2026]]). Each with its limits attached: - **Motivation.** In a survey of 1,484 Chinese undergraduates, [[zhang-ai-anxiety-academic-motivation-emotion-2026|Zhang, Shi and Lu (2026)]] found AI anxiety negatively correlated with emotion regulation (r = −0.172) and academic [[motivation]] (r = −0.175), while emotion regulation correlated positively with motivation (r = 0.457), all p < 0.001; bootstrap analysis confirmed a negative indirect path from anxiety to motivation through emotion regulation (indirect effect = −0.076, 95% CI −0.108 to −0.047), with the direct path still significant. Daily AI use duration correlated only weakly with motivation (r = 0.053) — a caution against reading time on tool as quality of [[student-engagement|engagement]]. - **Appraisals.** The school-support study separated predictors and outcome by about six weeks, but its authors state the design is still cross-sectional and cannot establish causal direction, temporal precedence or causal mediation; the sample was convenience-sampled, skewed toward STEAM majors, measured with [[self-report-measures|self-report]], and anxiety was generally low with limited variance, which may have attenuated the estimates ([[school-support-ai-learning-anxiety-control-value-2026|Jiang, Chen and Chen, 2026]]). - **Tools.** A dissertation on AI-driven campus well-being tools ([[ai-campus-wellbeing-tools|Tang, 2026]]) reports a survey chatbot (TigerGPT) reaching 75% usability and 81% satisfaction, and an adaptive follow-up question framework (AURA) producing a +0.12 mean quality gain (p = 0.044, d = 0.66) by using [[reinforcement-learning|reinforcement learning]] to choose whether to validate, specify, reflect or probe, plus a mental-health assessment component grounded in DSM-5 and PHQ-8 guidelines rather than black-box classification and a stacked multi-model architecture claimed to cut [[hallucination-risk|hallucination risk]] and outperform single models on the DAIC-WOZ [[benchmark]]. These are system-development results: they measure whether the tool works as designed, not whether students' [[well-being]] improved. The 89.09% accuracy comes from a single held-out split; the training data is not representative of Pakistani students, the system is English at its core with prompted Urdu expressions, and the conversational layer has no outcome data. Reading the corpus honestly: the sources and mechanisms of AI anxiety have measurable support, and the percentages and correlations above are associations, not demonstrated intervention effects. No study here randomized students to a support condition and measured anxiety afterward, so anxiety-reduction claims for any specific practice, including those recommended here, remain untested. What the evidence licenses is the direction of the work, not a promised size of effect. The policy environment around AI and [[social-emotional-learning|social-emotional]] learning is thinner than the deployment ambitions. [[policy-deficit-ai-sel-2026|Tran, Liu and Nguyen (2026)]] reviewed 65 peer-reviewed papers at the AI–SEL intersection: nearly three-quarters made no mention of policy implications at all, only about one in four offered any policy recommendation, and few gave actor-specific guidance on [[privacy]], teacher training or resource allocation. Their "WH" framework asks who should act, what action is recommended, why, when, where and how strongly it is framed — questions most left unanswered. Policy engagement correlated with publication venue, which the reviewers read as an incentive structure favoring technical novelty over [[governance]] and producing a "techno-solutionist trap": technical potential foregrounded, conditions for responsible use unspecified. For [[equity-in-ai-education|equity]], plausible AI-for-SEL tools can therefore be deployed without the safeguards — privacy, teacher preparation, resource equity — that adoption requires. ## The objections you will hear **"They need to learn resilience, not accommodation."** Partly right, and not in the confident version. Zhang, Shi and Lu found emotion regulation correlated positively with motivation (r = 0.457) while anxiety correlated negatively with both — an argument for teaching regulation skills. But in [[duan-ai-anxiety-career-decisions-college-2026|Duan et al. (2026)]]'s model, [[self-efficacy]] did not buffer the harm to career decisions. Resilience as "believe in yourself harder" is what failed; resilience as rehearsal, mastery experiences and career adaptability is what the studies point toward. **"This will pass as students get used to the tools."** Time on tool is not the mechanism: daily AI use duration correlated only weakly with motivation (r = 0.053) in the 1,484-student survey, and [[dag-ai-perceptions-career-anxiety-health-2026|Dağ et al. (2026)]] found AI anxiety still predicting job-search anxiety after controlling for socio-demographics (β = 0.234, p < 0.001). **"I am a subject teacher, not a counselor."** You are not being asked to treat anyone. Three of the four moves here are instructional — low-stakes practice, guided evaluation of AI output, feedback-driven revision — and the fourth is knowing the referral route. The one finding that speaks to your position is the teacher-student relationship sitting second in Bashir and Afzal's feature-importance analysis (10.0%), above sleep quality (9.3%) and social support (7.6%) — a hypothesis, they say, but a reasonable bet for a role you already occupy. ## Do this week - **Take one class period's temperature.** Ask students, in writing, what AI means for the job they are aiming at: you get the driver in their own words, and the students worth a private follow-up. - **Replace one reassurance with one appraisal move.** Swap five minutes of "you'll be fine" for a task where students evaluate and correct an AI output — the route with a measurable association to lower anxiety. - **Route career worry to the people who own it.** Identify your career services contact and the referral path now, so the conversation ends with a name rather than a shrug. - **Write your AI-use expectations down** — what is allowed, what must be disclosed, what support exists — and say them out loud. Perceived support precedes favorable appraisals. - **Before adopting any well-being or affect-aware tool,** get written answers on accuracy validation, how the support tier is decided, what happens on a harmful misclassification, how disclosures are stored, and how escalation to human counseling works. No answers, no deployment. For the surrounding picture, see [[how-ai-impacts-students]] for the general impact of AI on students and [[redesign-assessment-ai-era]] for the assessment side of institutional response. --- ## [How Can AI Support Disabled and Neurodivergent Learners in My Course?](https://edtechdev.github.io/aied/faqs/ai-disabled-neurodivergent-learners/) # How Can AI Support Disabled and Neurodivergent Learners in My Course? You are the instructor who just received an accommodation letter, the accessibility lead deciding whether to approve a license, or the disability services coordinator who has to answer a tool request by Friday. You do not need a literature review. You need a defensible decision: what to adopt, what to refuse, and how to tell whether the thing you approved actually helped the student in front of you. The bottom line, with its caveats left in place: AI can support disabled and neurodivergent learners, but what has been measured is narrow. Most of the evidence concerns tools that remove one functional barrier at a time; the strongest quantitative result comes from PK-12 special-education settings using mostly pre-generative AI; and the higher-education picture is a map of what has been tried rather than a measurement of what works. Nothing here justifies blanket adoption. Everything here supports a narrow, instructed, barrier-specific adoption — and names several practices that go wrong. ## Start with the barrier, not the diagnosis Most of what goes wrong for disabled and neurodivergent students happens at the level of a task rather than a diagnosis: a reading that assumes print, a video that assumes sustained attention, a group project that assumes unspoken social rules. Those barriers recur across diagnoses, so the productive question is not which tool suits which condition but which functional barrier a tool removes. Every recommendation below is organized that way, and so should your procurement be. ## What to adopt, ordered by how strong the evidence is **1. Barrier-removing tools with a measured effect — medium confidence, but PK-12 and mostly older AI.** The strongest quantitative result on this page is also the most general. [[zhang-ai-students-disabilities-meta-analysis-2024|Zhang, Carter, Liu and Peng (2024)]] synthesized 29 (quasi-)experimental studies of AI-based interventions for students with disabilities — 239 effect sizes from 41 independent samples — and found a statistically significant medium overall effect on learning outcomes, **Hedge's g = 0.588** (95% CI [0.349, 0.826], p < .01), with high between-study heterogeneity. Academic performance was the largest outcome subgroup (k = 80, g = 0.929) and social-emotional skills the smallest (k = 144, g = 0.382). Computer software — speech recognition, expert systems, [[intelligent-tutoring|intelligent tutoring systems]] — performed strongly (k = 74, g = 0.959) and [[educational-robotics|robots]] moderately (k = 144, g = 0.509), while intelligent VR (g = 0.528) was not significant. Publication bias was present (Egger's β = 2.837, p < .001); trim-and-fill reduced the estimate to g = 0.269, still significant. Read the confidence label honestly: every study was PK-12, 96.6% used step-based instruction in special-education or clinical settings, and most used older rule-based AI rather than [[generative-ai|generative AI]]. This is grounds for cautious adoption of structured, instructional software — not a warrant for the chatbot you were pitched. **2. Concrete single-barrier tools with small studies — low confidence, small samples.** Three results show what an individual tool can do, each tested on a handful of learners. For attention, [[adhd-video-segmentation-computing-education|Pimenova, Begel and colleagues (2026)]] segmented instructional videos post hoc into single-instruction chunks with fixed pauses; with 17 learners with ADHD and 10 without, it improved everyone and brought the ADHD participants' errors and hesitations to parity — an equalizing effect from a lightweight transformation rather than a diagnosis-specific product. For language access, [[llm-question-generation-deaf-hard-of-hearing-2026|Chen et al. (2026)]] added Visual questions (timestamps where visual information is likely to be misread) and Emotion questions (timestamps where prior Deaf and Hard of Hearing learners reported frustration) to an [[llm]] baseline set, producing 30 questions, ten per strategy; with 16 learners, self-efficacy averaged 5.70 (SD = 1.12) on a seven-point scale, and visual questions were chosen more by Deaf than hard-of-hearing participants. For reading and writing, [[dyslexlens-dyslexic-learners-ai|Rezazadegan et al. (2026)]] mined dyslexic learners' own forum discussions and found they value AI for literacy support while reporting uneven output quality and a lack of equitable accommodations — the inconsistency is itself a barrier. Adopt these as low-cost, reversible experiments, not as fixed infrastructure. **3. The higher-education inventory — a map, not a measurement.** For [[higher-ed|higher education]] the picture is a map rather than a measurement. [[assistive-tech-neurodivergent-higher-ed-review-2026|Rempel, Heimann and Prilop (2026)]] screened 766 records from five databases down to 40 included studies published between 2015 and 2025. Generative AI appeared in 15 of the 40 and [[virtual-and-augmented-reality|virtual reality]] in 11. Barrier coverage was uneven: reading and writing (n = 13) and study management (n = 12) dominated, attention (n = 4) and social communication (n = 5) were neglected, and only 6 studies addressed more than one barrier. The review covers higher education only — it excluded [[k-12]] and specialized-institution settings by design, so it is not evidence about younger learners. Use it to see which barriers have been addressed at your level and which have barely been touched. **4. Structure and universal design, which are evidence-backed and free.** Learners' stated preferences supply the design brief. [[neurodivergent-computing-students|Zastudil et al. (2026)]] surveyed 24 neurodivergent computing students (autistic and/or ADHD) and 20 neurotypical peers with four follow-up interviews: the neurodivergent group was uncomfortable with unstructured or ambiguous assignments and strongly preferred smaller teams that work together consistently with explicitly defined roles, coping by self-selecting roles and disclosing strategically. Their warning for [[agentic-ai|AI]]-mediated work is that a tool's interaction model can reproduce the same structural ambiguity. Structure is the cheapest intervention on this list, and the one most often omitted. ## What to avoid, including accommodations that backfire - **Do not supply AI as an accommodation without teaching it.** The most common failure is not the wrong tool but an untaught one. The case *W.A. v. Clarksville/Montgomery County School System* (2024) is the documented version: a student given AI as an accommodation graduated unable to read, and the court found that a free, appropriate public education had not been provided. - **Do not let a blanket AI prohibition remove an assistive tool.** [[wright-transcription-not-generation-2026|Wright (2026)]] argues that rules treating all "AI use" alike are over-inclusive because they do not separate speech-to-text transcription and OCR from generative drafting. Students with conditions affecting fine motor control, handwriting legibility or typing accuracy have relied on standalone voice-to-text products such as Dragon NaturallySpeaking and standalone OCR, several of which have been discontinued or degraded, with AI-powered transcription filling the functional gap — so a rule written as "no AI" can withdraw the student's primary means of producing legible work, and the scale of that displacement has not been measured. Treat transcription and OCR as accommodations to name explicitly in the policy rather than casualties of it; the exposure this creates belongs with [[legal-issues-and-risks]]. - **Do not organize support around diagnosis alone.** The scoping review's central finding is a critique: the literature is dominated by individual accommodations and single-neurotype tools organized around formal diagnosis, while the functional differences they address cut across neurotypes and are hidden by diagnostic categorization. Dyslexic students tend to benefit from reduced reliance on written-only material, and both autistic students and students with ADHD can experience frequent fluctuation in focus, which makes performance sensitive to consistent structure and enough time for task switching. A tool matched to a label can miss the barrier the student actually has. The same caution applies to group statistics: the largest meta-analysis in this knowledge base reports a considerably larger effect for students with specific learning disabilities, intellectual and developmental disabilities or who are deaf than for students with autism spectrum disorder, with that difference reported as not statistically significant across its effect sizes (see [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]]). - **Do not make self-identification the price of access.** Accommodation requires the student to identify themselves as needing special provision, which carries stigma; universal design removes that step. - **Do not ignore inversion effects.** Chatbot over-reliance and the cognitive overload of immersive environments are documented as inversion effects, where adding the technology lowers learning. A tool that only reproduces existing teaching in digital form should be expected to yield limited [[learning-gains|learning gains]]. - **Do not deploy tools whose reliability is unproven.** Dyslexic students using AI for literacy support reported real value alongside unreliable output quality; a tool that cannot be relied on adds work instead of removing it. - **Do not accept products as they arrive.** State universal design and participatory development as procurement requirements rather than accepting whatever is sold to you — the scoping review recommends building tools with neurodivergent students rather than for them, and its own sample shows how rare that is, with only 3 of 40 studies treating the environment or neurotypical peers as the object of intervention. - **Do not buy the most immersive option by default.** [[virtual-and-augmented-reality|Virtual reality]] is the most resource-intensive option in the set. Assistive technology needs devices and connectivity, and tools free in a pilot routinely move behind a paywall. - **Do not let automation bias stand in for judgment.** [[ludia-udl-ai-thought-partner-2026|LUDIA]]'s authors name a limit that belongs on every deployment plan: automation bias means the educator most likely to accept a poor suggestion is the one who trusts the tool knows [[universal-design-for-learning|Universal Design for Learning]]. Nobody has tested its outputs for cultural or linguistic bias either. ## How to tell whether a tool helps *this* student Aggregate effects tell you a direction; they do not tell you whether the tool works for the student in your section next week. Three cautions should shape any trial. First, the subgroup ordering is suggestive rather than established. The 2024 meta-analysis cannot say which design choice is better: no moderator of any kind reached significance, so rankings such as teachable agents (g = 1.100) ahead of social-emotional coaches (g = 0.336), or dyadic interaction ahead of triadic (g = 0.973 versus g = 0.385), should not be treated as settled. The defensible position is that the direction of effect is supported while the size, design optimum and durability of any gain are not. Second, the corpus describes deployments and perceptions more than measured learning. Attention (n = 4) and social communication (n = 5) are the least-studied barriers, and only 6 of 40 studies served more than one. The scarcity of (quasi-)experimental and longitudinal designs, and the heterogeneity of tools and outcomes, are why the higher-education work is a scoping review rather than a meta-analysis. Coverage skews undergraduate and Global North, the search was English-language peer-reviewed work, and postgraduate students are under-represented. Treat engagement and satisfaction as weak indicators; measure the barrier you set out to remove. Third, a tool that fixes one barrier may leave another. Only 6 of 40 studies served more than one barrier, so check the tool against the specific difficulty the student reports — reading and writing, study management, attention, or social communication — and evaluate against that barrier, not against a satisfaction survey. Ask the student; Zastudil et al.'s participants and Rezazadegan et al.'s forum posters both describe their needs more precisely than the products built for them. ## Objections you will hear **"We cannot require a tool for everyone."** You can, where the tool is a design feature rather than a personal accommodation. Video segmentation with fixed pauses improved all learners, not only those with ADHD, and bringing the ADHD participants' errors and hesitations to parity is exactly the universal-design outcome that removes the need for disclosure. Pimenova and Begel's result is an equalizing effect from a lightweight transformation rather than a diagnosis-specific product. Pair that with accessible defaults: read-aloud, OCR and text-leveling tools recur in the validated policy items of [[shin-ai-policies-sld-2026|Shin et al. (2026)]], whose LLM-based topic modeling and two rounds of Delphi surveys of 17 experts across 12 U.S. AI-in-education policy documents from 2015 to 2025 found that **only 2 of the 12 specifically address learning disabilities**, while 18 topics present in general AI or other-disability policy — data protection, legal risk management, ethical guidelines among them — are missing from specific-learning-disability policy. **"Captioning and translation are not academic support."** Captioning is a start, not the whole of language access. Text-based prompts mismatch sign-based first languages: pair transcripts with visual and emotional context cues and cut unnecessary complexity, as Chen et al.'s Visual and Emotion questions do. And the policy experts' own ranking disagrees with the objection: their 36 validated items fall into five themes — inclusive and personalized learning (11 items, 30.56%), ethics, equity and inclusion (9, 25.00%), student empowerment and AI literacy (6, 16.67%), assessment and research (6, 16.67%), and educator preparation (4, 11.11%) — and they ranked **student empowerment and AI literacy** as most essential: teaching these students to use AI responsibly and independently while guarding against over-reliance. **"Students must be independent."** Independence is what the instruction is for, which is why the tool without the teaching is the failure mode. Shin et al.'s educator-preparation theme exists for the same reason, and LUDIA's authors name the obstacle as the "knowing-doing divide": guidance exists in abundance but is written in general terms, while a barrier is always particular to this room, this week. The scale of need is easy to understate: roughly 240 million children worldwide live with disability, about half of them out of school in low- and middle-income countries, and about a third of teachers across OECD systems report lacking the competencies to support students with specific needs. The legal frame is the U.S. Assistive Technology Act (2004) and IDEA (2004), which place assistive-technology evaluation inside each Individualized Education Program and require a free, appropriate public education, yet no formal evidence-based guidelines help educators or families implement AI for these learners. Independence built on an untaught tool is what the case law punishes. **"We cannot afford it."** Then refuse the expensive default. The assistive-technology field is concentrated in the Global North, its most immersive tools are the least scalable, and virtual reality is the most resource-intensive option in the set. LUDIA is one worked version of the universal-design alternative: it connects educators with Universal Design for Learning while a design decision is still open, and removes access barriers by architecture — no cost, no accounts or stored chats, WCAG 2.2 Level AA conformance, and 13 machine-translated languages, though non-English versions may carry the assumptions of the English they came from. It is a thought partner rather than a solution engine, organizing its inquiry around "proof of trust" rather than proof of impact, and stating plainly that the authors hold no evidence the tool improves learning. Budget honestly for the other cost: [[shin-ai-policies-sld-2026|Shin et al.]] flag the [[digital-divide|digital divide]], since the price of advanced AI tools can widen disparities between wealthier and poorer schools, alongside algorithmic bias and teacher-training gaps. ## Do this this week - **Name the functional barrier first.** Reading and writing, study management, attention, social communication — then check whether the tool addresses that barrier or only a diagnosis. - **Use what is evidenced.** Video segmentation with fixed pauses narrowed the ADHD performance gap to parity, and read-aloud, OCR and text-leveling tools recur in the validated policy items. - **Treat captioning as a start, not the whole of language access.** Pair transcripts with visual and emotional context cues and cut unnecessary complexity. - **Give structure rather than asking students to infer it.** Small consistent teams, defined roles, explicit expectations and role self-selection were the preferences neurodivergent computing students reported. - **Build accessibility into procurement.** State universal design and participatory development as requirements, prefer free and privacy-preserving tools, and ask what happens when the free tier ends. - **Train people, not just deploy.** Budget time for students to learn the tool and for staff to work it into the assignment; an accommodation without instruction failed in court. - **Design for overlap, and watch for inversion effects.** Only 6 of 40 studies served more than one barrier; chatbot over-reliance and cognitive overload in immersive environments can reduce learning. - **Report outcomes honestly.** Most of the evidence describes deployments and perceptions, and no design moderator has been established; treat engagement and satisfaction as weak indicators. - **Fund the training and workflow integration** that turn an accessible tool into usable support, and write the cost of devices and connectivity into the plan. ## Where to go next This page is the practical, cross-learner entry point: [[universal-design-for-learning]] covers the design philosophy that prevents barriers before they appear, [[assistive-technology]] the tool layer a student uses, [[accessibility]] whether a format can be perceived and operated at all, and [[neurodiversity]] the lens that treats difference as diversity rather than deficit — all within the umbrella of [[inclusive-learning]]. For the research-side obligations, see [[equity-ethics-pedagogical-safety-research|incorporating equity, accessibility, privacy, ethics and pedagogical safety into AIED research]]. --- ## [How Can AI Help Me Give Better Feedback at Scale?](https://edtechdev.github.io/aied/faqs/ai-feedback-at-scale/) # How Can AI Help Me Give Better Feedback at Scale? It is week six. You have a stack of drafts, a rubric you believe in, and a growing suspicion that a third of your comments will be skimmed and forgotten. A colleague mentions that a tool now comments on every draft within the hour. You want to know whether that is a real upgrade or a faster way to produce text nobody reads. **The bottom line:** AI feedback is worth adopting for coverage, timeliness and consistency inside a narrow band — copy-editing-scale language work and rubric-guided first-pass comments — and it is worth adopting only if you design what happens around the comments. Reach is the easy part. In [[genai-feedback-design-multisite-experiment|a cluster-randomized experiment with 1,176 first-year undergraduates]], reflective and hybrid designs beat direct AI critique on delayed, AI-free [[transfer-of-learning|transfer]]; in [[farrokhnia-genai-feedback-student-revisions-2026|a randomized essay experiment with 70 students]], higher-quality AI feedback did not produce better revisions than an experienced teacher's. Across the evidence, what students do with a comment matters more than how fast the comment arrives, and [[feedback-literacy|feedback literacy]] — not AI access — separates the students who gain from the ones who copy. ## The distinction that decides everything: volume is not quality, and delivered is not read Two separations do most of the work on this page. **More feedback is not better feedback.** AI can produce comments on every draft every week; the quantity is nearly free. Quality is not. [[zhan-boud-dawson-genai-feedback-engagement|Zhan, Boud, Dawson and Yan (2025)]] make prompt quality the hinge of the process: vague prompts yield generic, useless output, and generated comments can be hallucinated, biased or overlapping, which pushes [[evaluative-judgment|evaluative judgment]] back onto the reader. Calibration is sharper still: in [[farrokhnia-genai-feedback-student-revisions-2026|Farrokhnia et al. (2026)]], GenAI feedback quality was significantly associated with the strength of a student's initial draft, while teacher feedback quality was not — the teacher calibrated more consistently. Volume scales. Calibration does not. **Delivered feedback is not enacted feedback.** This is the more expensive confusion. Farrokhnia et al. randomized 70 students writing argumentative essays in Persian into three groups. Chain-of-thought prompting produced significantly higher-rated feedback (M = 12.90) than both zero-shot prompting (M = 11.25, p = .01) and an experienced human teacher (M = 11.20, p = .008). The chain-of-thought group nevertheless did not revise its essays significantly more than the teacher-feedback group, which achieved comparable gains. Better-rated comments, the same revision. If your metric is the quality of what the model writes, that reads as a win; if your metric is what changed in the essay, it does not. Who enacts feedback depends on the learner. In [[hawkins-feedback-literacy-ai-essay-writing|Hawkins, Taylor-Griffiths and Lodge (2026)]], 32 psychology students did a screen-recorded 25-minute essay task with unrestricted access to ChatGPT, then watched the recording back in a video-stimulated interview. Feedback literacy was the only significant positive predictor of essay grade (β = 0.46, p = .017); [[feedback-futures-genai|the special-issue editorial by Zhan, Wood, Carless and Yan (2026)]] notes that frequency of GenAI use, [[trust|trustworthiness of the source]] and [[prior-knowledge|prior knowledge]] did not predict performance. Fewer than a third of the students compared AI output against another internet source; half expressed deliberate AI avoidance for [[academic-integrity|academic integrity]] or wanting the essay in their own words; and most requested task-level feedback, which transfers poorly to other tasks. Uptake is also unevenly distributed. In the university-wide Deakin pilot reported by [[tubino-adachi-ai-automated-feedback-literacy|Tubino and Adachi (2022)]], the AI tool was optional, and usage averaged roughly 13% of undergraduates and 12% of postgraduates despite almost 4,000 students having access. Proactive, high-achieving students used it more. That is what "making the tool available" buys you, and it is why the rest of this page is about design rather than access. ## What AI is genuinely good at, and where it stops **Good: coverage and consistency inside a narrow band.** Tubino and Adachi report that pilot at Deakin in 2021 using FeedbackFruits' AI automated feedback tool across 29 units and almost 4,000 students. It works on micro-level text features — sentence length, punctuation, grammar, text structure — supporting the copy-editing stage: the instructor sets parameters, the student uses the tool independently and gets timely, actionable feedback. The division of labor is what scales, not the model. **Bad: generality, correctness and calibration.** That is where it fails (Zhan et al., 2025), and the remedy is a prompt, a rubric or a human — not a bigger model. [[learner-centered-feedback-ai|Aldino et al. (2026)]] add instructor-side failure modes from 21 higher-education teachers: inconsistent tone (n = 7), potential misinformation (n = 5) and trust issues (n = 5), with revision effort spent deleting exaggerated praise and generic suggestions rather than adding content. ## A deployment pattern that holds up The evidence supports one shape more than any other: the AI drafts, the student evaluates first, and a human decides. **What to automate.** Micro-level language work and rubric-guided first-pass diagnostics. Give the model a rubric and make it reason step by step — Farrokhnia et al.'s quality gain came from a chain-of-thought prompt walking the model through a step-by-step evaluation against an argumentation rubric, not a bare instruction. [[scaffolding-srl-feedback-genai-human-peers|Gu, Chen and Yan (2026)]] likewise gave their GenAI group ChatGPT-4o with pre-trained rubrics and prompt guidelines, which is what made the output checkable against stated criteria. Rubric-referenced generation is automatable; judgment is not. **What to keep human.** The final message, the tone, and anything that becomes a record. [[becerra-aicofe-feedback-2026|Becerra, Palma and Cobos (2026)]] describe AICoFE, which runs three independently fine-tuned models (GPT-4.1-mini, Gemini 2.5 Flash, Llama 3.1) over rubric scores and qualitative observations to produce independent draft comments, then has the instructor compose the final message by selecting sentences or paragraphs, with a legend showing which model contributed what. AI is a draft generator, not a final deliverer. Aldino et al. find the same "assist but verify" pattern: the ML component most often flagged **Meeting Learning Objective** as missing (20 of 21 teachers had omitted it; 16 accepted) and **Student–Teacher Relationship** (14 omitted; 12 accepted), with teachers deciding each suggestion on professional judgment. **How to sequence it.** Do not stack sources side by side; stage them. In [[genai-feedback-design-multisite-experiment|Ateş's (2026)]] experiment, the reflective condition required self-evaluation before AI critique, and the hybrid ran self-evaluation → peer feedback → GenAI critique. Tubino and Adachi propose templates for three drafting stages for the same reason. Combine AI with peer feedback rather than substituting one for the other: Gu et al. ran two parallel English classes (N = 118 first-year undergraduates in China; GenAI n = 56, peer n = 62) through three self-assessment cycles over one semester, and GenAI scaffolding improved feedback literacy slightly but significantly over human peer review (ANCOVA p = 0.049, ηp² = 0.03). GenAI students refined prompts iteratively and verified or challenged inaccurate output, while peer-group students chose sources by social convenience and had their evaluative judgment distorted by friendship bias. Peer review kept distinct value for audience awareness, so the authors propose multi-stage designs with anonymous peer feedback. **How to frame it for students.** As a draft opinion to be judged, not an answer to be copied. Say plainly that model comments can be hallucinated, biased or overlapping and that the student owns the revision. The pair of IELTS writers in Zhan et al. shows the extremes: a low-literacy student with a vague prompt got generic feedback, trusted or over-copied it and engaged superficially; a high-literacy student wrote a criteria-referenced prompt, followed up, cross-checked sources and monitored revisions. The difference was instruction and habit, not tool access. **How to check the output.** Sample generated comments against your own reading of the same drafts, looking for the failure modes named above — tone, misinformation, exaggerated praise, overlap. In the Deakin pilot, positive average ratings coexisted with student-flagged errors, so satisfaction is a weak quality signal; a smiley-face survey is not accuracy evidence. ## Failure modes, with the numbers **Design, not availability, drives the effect.** The strongest design evidence is a multisite, cluster-randomized, longitudinal field experiment: 1,176 first-year undergraduates in 48 sections across 4 universities and 3 science domains, randomized at section level into peer-only, direct GenAI, reflective GenAI and hybrid conditions. Hybrid produced the highest argument-quality gains and the clearest advantage on conceptual learning. Reflective and hybrid both produced stronger feedback uptake and [[self-regulated-learning|self-regulated learning]] than direct GenAI, and both outperformed it on delayed, AI-free transfer. Direct GenAI did improve immediate argument quality over peer feedback — the advantage it has — but with weaker transfer. Educational value therefore depends less on AI access than on whether the feedback environment preserves student [[agency|learner agency]], evaluative judgment and ownership during revision: direct critique invites passive uptake, while reflective and hybrid designs force the student to interpret it, compare it against criteria, judge its relevance, and then revise. Other effect sizes temper this. Gu et al.'s advantage over peer review was small, and they read it as real but not transformative alone — teacher scaffolding, prompt guidelines and worksheets did the supporting work. The editorial's framing is that GenAI and human feedback are complementary only if the complementarity is specified, through sequencing, comparison, editing and governance rather than by leaving two sources side by side. **Over-trust at both ends of the pipeline.** Students can over-rely uncritically when enacting feedback, and instructors can over-defer to suggestions. Teacher feedback is also often perceived by students as more negative or riskier than GenAI feedback, which inverts the usual quality assumption. Treat feedback literacy as a prerequisite rather than an assumption: the strongest predictor in Hawkins et al. was the learner's feedback literacy, not how often they used AI. **Workload that moves rather than shrinks.** [[ai-save-instructor-time|The workload question]] is the one most often promised away. Asked what the AI feedback tool gave them, 21 teachers named saving time (n = 2) only rarely, and named reflection (n = 14), improved language and structure (n = 11) and identification of missing components (n = 10) far more often. The visible gain is a draft and a diagnostic prompt, not a finished comment. What appears instead is review, calibration and verification work. Twelve of the 21 teachers made further sentence-level revisions; editing (f = 32) and removing (f = 27) dominated, with adding rare (f = 8). The dominant pattern was calibrating tone — editing praise down (f = 11), removing suggestions (f = 9), removing encouragement (f = 8) — to protect authenticity and professional voice, and revisions clustered in the relational dimension, especially Student–Teacher Relationship (f = 24). Teachers named the need for human editing (n = 9), inconsistent tone (n = 7), misinformation (n = 5) and trust issues (n = 5); instructors with more than five years of experience reported more of them, while less-experienced teachers valued the scaffolding — a reason to watch whether novices who defer to AI suggestions build less independent judgment. The editorial's rule: GenAI redistributes rather than removes teacher labor, and poorly designed tools can increase it. ## Three objections, taken seriously **"AI feedback is not personal."** Often true, and it is precisely the reason to keep the last word human. AICoFE exists because generic model prose is not what a student should receive — the instructor selects, edits and signs the final message. The 21 teachers in Aldino et al. rewrote tone for the same reason, deleting exaggerated praise and cutting suggestions that did not match their relationship with the student. Personalization is not the model's job in this design; it is yours, and the model is doing the setup work. **"My students will not read feedback anyway."** They read less than we hope, and the evidence says the fix is design, not volume. Uptake in the Deakin pilot ran at roughly 13% of undergraduates and 12% of postgraduates. Farrokhnia et al.'s higher-rated AI comments produced no more revision than teacher comments. What moved the needle was sequencing that forces engagement — self-evaluation first, then peer feedback, then AI critique — plus three revision cycles across a semester that produced measurable feedback-literacy differences. Multiple same-day resubmissions are a usable trace of whether students engaged at all. **"This is not what students pay for."** Then do not let the AI be the only voice they hear from you. The defensible version of the tool is the one where students get more feedback moments while you spend your hours on the judgment-intensive parts: the final word, the calibration, the comments that carry your authority. What students lose under careless deployment is not your presence but the coherence between criteria, comments and outcome. Specify the complementarity — sequencing, comparison, editing, governance — and the objection dissolves; leave two sources side by side and it does not. ## This week **1.** Sort your own last set of comments into language-level fixes and judgment-level calls. Only the first category is a candidate for automation. **2.** Write the rubric-referenced, step-by-step prompt you would hand a model, and test it against three drafts you have already graded. **3.** Sample the output against your own comments and log where the model was generic, wrong or flattering. **4.** Stage one assignment as self-evaluation → peer feedback → AI critique, and tell students the AI is a draft opinion they are expected to challenge. **5.** Ask for one comparison against a second source or against the criteria — fewer than a third of students do this unprompted. **6.** Budget the review time you will actually spend on tone and deletion, and compare it with what you freed. **7.** Watch the expert–novice split. If experienced colleagues report more problems with the tool's output than you do, that is information about your own calibration. **8.** Decide what outcome counts and measure that: usage rates, revision traces and delayed unaided performance — not satisfaction ratings. The Deakin pilot's roughly 13% and 12% are a useful baseline, and AICoFE's curation tracking models the instructor side: log how much of the final feedback came from the model and how much you changed. **Where this page sits.** [[feedback]] is the umbrella concept for the whole system — provision, loop, uptake and assessment context — and [[ai-feedback-quality|AI feedback quality]] covers what makes generated comments accurate, useful, timely and pedagogically sound. This page is the practical middle: which combinations of those parts produce better feedback at scale. Questions about what a grade should mean when AI is involved — construct substitution, authenticated process evidence, integrity rules — belong to [[redesign-assessment-ai-era|redesigning assessment in the AI era]] and [[assessment-validity|assessment validity]] rather than here. For the writing-instruction special case, see [[writing-instruction-ai-best-practices|best practices for writing instruction in the context of AI]]. --- ## [How Should Parents and Teachers Approach AI with Children Under 13?](https://edtechdev.github.io/aied/faqs/ai-guidance-children-under-13/) # How Should Parents and Teachers Approach AI with Children Under 13? The AI is already in your house or your classroom: the toy that talks back, the tablet companion, the tutor your child opens while you make dinner, the writing tool a colleague is enthusiastic about. You are deciding what a child under thirteen may do with it, with or without a policy. The bottom line: **generative AI is the advanced tier for this age band, not the entry point.** Let a young child meet AI first as something people design and children can shape, keep an adult in the loop wherever the child cannot judge the output, and protect the child's own thinking. A blanket ban is not the safest option, because most of this age group's exposure happens at home where a school ban does not reach — and no study of age-based bans exists. ## The short version: allow, block, supervise - **Allow.** Under 9: unplugged play, tangible coding, storytelling, role-play and hands-on making, plus explanations of what a machine is. Ages 9–12: scaffolded, often non-generative activities on vetted platforms, plus explicit teaching about data and what AI cannot know. - **Block.** Unsupervised generative chat for the youngest; any tool that keeps recordings, transcripts or a child's data without a plain answer to "what happens to it"; any tool you cannot vet in your families' languages. - **Supervise.** Co-use, not quiet monitoring: ask what the tool did with the child's words, and keep unaided work visible — rising output quality with falling unaided work is the pattern to catch. - **Do not build a rule you cannot enforce.** Detection-based enforcement is not defensible ([why](#do-not-build-your-rule-on-a-detector)). ## What a child under thirteen should be allowed to do **Tiered design already exists, and it puts generative use last.** [[age-tiered-ai-literacy-guidebooks-2026|Wang, Chuang and Wu (2026)]] evaluated an age-tiered AI literacy resource for Taiwan's K-12 system in two editions: an **Elementary edition for ages 9 to 12** of scaffolded, platform-based, **non-generative** activities, and an **Advanced edition for ages 13 to 18** with ethical reasoning and supervised [[generative-ai|generative AI]] use. In a single post-exposure survey of 831 participants, the Elementary cohort rated the materials significantly higher on performance expectancy, effort expectancy, playfulness and behavioral intention, with perceived playfulness the strongest correlate of intention in both groups — though the authors are explicit that this is short-term acceptance after roughly thirty minutes of exposure, not [[learning-gains|learning]], ethical competence or actual use. **Play is the mechanism.** [[play-ai-pre-k-kindergarten-ai-literacy-2026|Play With AI (PL-AI)]] is a play-centered pre-K and kindergarten curriculum built on [[embodied-learning|embodied play]], tangible coding as a procedural bridge, guided dialogue as reflective [[scaffolding]], and teacher co-design as the [[sustainability]] anchor; teachers' self-rated comfort rose across design cycles, and children reasoned about AI as human-designed, built sequencing and debugging through tangible coding, and reflected on fairness through guided dialogue. [[ai-play-framework-early-childhood-2026|AI-Play]] takes the unplugged route, with screen-free activities parents can replicate at home. **For the youngest children the risk is belief.** [[creative-project-approach-ai-early-childhood-2025|Yang, Li and Lee (2025)]] caution that generative social robots can produce content preoperational children accept as true, that these systems are not developmentally calibrated in the feedback they give, and that their cost and availability widen existing [[digital-divide|divides]]. [[ai-toys-child-development-2026|Xu, Girouard and Shi (2026)]] call the mechanism a **redistribution of agency in play**: control shifts toward the child–toy interaction, constraining play when the toy organizes it and enabling it when the toy follows the child's imagination. **Teach how machines work.** The elementary tier is deliberately non-generative, and the kindergarten [[educational-robotics|robotics]] evidence agrees: [[tsingidou-ct-robotics-kindergarten-2026|Tsingidou et al. (2026)]] find problem-based learning, storytelling and scaffolding the most-used strategies in robot-mediated computational thinking, with sequencing, debugging and algorithmic design the most-assessed skills (most often via TechCheck-K, though many tools are developed ad hoc without validation). Skip this and the risks appear later as [[cognitive-offloading|over-reliance]]. **Cross-cutting skills can be taught cheaply.** [[demir-akar-ai-media-literacy-children-2026|Demir and Akar (2026)]] ran an 18-hour, 5E-based critical media literacy program for 36 Turkish fourth-graders covering data privacy, safe communication and media ethics, with between-group Cohen's *d* of 1.12 (reading), 1.18 (writing) and 1.31 (total literacy). ## Supervision and disclosure: who is in the room **Expect the responsibility to be unclear; settle it explicitly.** In interviews with 33 US K-12 teachers, [[k12-teachers-ai-companion-literacy-2026|Xiao et al. (2026)]] found 10 named the parent-facing companion scenario as the most concerning, 17 said a teacher should get involved and 11 said it depends; teachers' visibility test left home AI use outside their jurisdiction absent observable effects, and parents were named as primarily responsible while described as unaware or overstretched. [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] formalize family–school arrangements as additive, synergistic and compensatory models, and state plainly that the synergistic version has no supporting evidence. **Give the adult a defined role, or the tool takes it.** [[paratutor-parent-child-tutoring|ParaTutor (Luo et al. 2026)]] separated parent and child roles in LLM-mediated home mathematics tutoring with 23 parent–child dyads (children aged 10–12) across four conditions: generic assistance tended to displace the parent's tutoring role, while the role-separated interface preserved it. Parents also struggled with content knowledge and communication. **Companions that read to children need the same care.** [[liao-role-adaptive-ai-companion-book-talk-2026|Liao (2026)]] found a fixed peer-role [[conversational-ai|AI companion]] produced significantly longer interactions but a lower share of words and sentences from the student, and was stronger on factual recall than on emotional or future-oriented reflection — the reason the proposed framework adds a **Parent Advisor** role and names a support vacuum for teachers and parents. **Keep the human in the parts that matter most.** [[human-ai-complementarity-social-emotional-learning-2026|Raave and colleagues (2026)]] compared a generative AI conversational agent with human educators facilitating story-based social-emotional learning activities with a static AI child — 18 simulations per facilitator, 108 observations, rated blind to facilitator type. The agent was strong on respectful tone and routine procedural scaffolding; human educators were stronger at deeper SEL instruction, guiding reflection and promoting social-emotional knowledge, so early [[social-emotional-learning|social-emotional learning]] should default to complementarity. **Disclosure is the school's obligation, not the family's detective work.** Write down what the school uses and how a family can ask; see [[privacy]], [[pedagogical-safety]], [[ai-use-disclosure]] and [[equity-ethics-pedagogical-safety-research|equity, ethics, privacy and safety in AIED research]]. ## Privacy and data: what the tool keeps **Assume the safety layer was not built for your child.** [[child-safety-genai|A child-safety evaluation framework grounded in expert guidance and real incident data]] reports that most AI safety frameworks and benchmarks target adult users despite evidence of heavy youth engagement — a national survey cited in that work found 72% of US adolescents had used AI companions — and that when three Llama Guard models were tested on education-related unsafe prompts, they struggled to identify them. **Ask four questions before adopting a tool.** Which child-specific hazards were tested? What does it do with recordings, transcripts and voice data by default? What will it cost families if the free tier disappears? Does it work in the languages your families speak? [[all-girls-genai-makerspace-gender-equity-2026|An all-girls makerspace case study (Liu et al. 2026)]] documents parents reporting that girls became scared and spoke less in mixed-gender settings, a free image generator that dropped its free functions for a paid tier the nonprofit could not absorb, and English-only prompts as a further barrier. **For school leaders, staff confidence is a privacy control.** [[preschool-teachers-ai-behavioral-intention-2026|A study of early-childhood settings]] finds perceived usefulness and ease of use central to preschool teachers' intentions to use AI, **AI self-efficacy** a meaningful predictor, **AI anxiety** a deterrent, and subjective norm — colleagues and leadership — shaping intention. ## Integrity: what is the child's own work? **Do not build your rule on a detector.** [[bassett-ai-detectors-education-2026|Bassett et al. (2026)]] argue AI detection should not be used in education at all: its estimates are probabilistic and cannot be independently verified because real-world text origin is unknown, its use violates procedural fairness since detector scores do not meet the balance-of-probabilities standard integrity investigations require, and its human-versus-AI dichotomy is meaningless for work created *with* rather than *by* AI. Detection "does not safeguard academic integrity; it undermines it," eroding [[trust]]. **Surveillance pushes concealment; it does not remove it.** [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi (2026)]] model assessment as imperfect information: the student knows how the work was produced, the institution sees only the artifact. Deterrence runs through a detector's *discrimination* between hidden use and legitimate work rather than its raw catch rate — when false positives rise faster than true positives, stronger monitoring makes concealment relatively *more* attractive. [[qu-wang-disclose-or-not-genai-2026|Qu and Wang (2026)]] found non-disclosure among 409 undergraduates was strategic adaptation to perceived peer norms and low interpretive trust, not moral negligence. **Watch unassisted work, not the product.** [[stromberg-generative-ai-learning-penalty-secondary-2026|A 26,811-student secondary analysis]] found homework scores rose while closed-book exam scores fell — hence the recommendation to monitor inputs rather than outputs. [[elementary-writing-genai-systematic-review-2026|A systematic review of 8 studies of AI in elementary writing instruction (2019–2025)]] finds conversational tools supporting writing practice and multimodal tools extending composition beyond text, while concluding that automated assessment still requires [[human-in-the-loop-ai|human oversight]] for fairness and accuracy. [[genai-writing-program-primary-l2-motivation-engagement|A nine-week GenAI-supported opinion-writing program]] with 301 Grade 5 and 6 students raised ideal writing self and academic buoyancy and lifted behavioral and emotional engagement, but did not move growth mindset, cognitive or metacognitive engagement, or organization — and names the risks: over-reliance, shortcut-seeking and diminished self-monitoring. **Know where the line is before you announce it.** [[genai-policies-higher-ed-computing|A comparative analysis of institutional policies and course syllabi]] found institutions broadly encouraging GenAI use (63% of 116 policies) while half of the 98 course syllabi outright prohibited it. [[nash-preservice-teachers-classroom-ai-policies-2026|Nash and Burriss (2026)]] found the same reflex among 27 preservice teachers writing classroom policies: 26 of 27 permitted some AI use on teacher-specified terms, only one prohibited it entirely, and the limits were rarely operationalized — one policy allowed AI "to get your thinking started" and then declared "this is where the line should be drawn" without saying where. [[adarkwah-genai-unesco-policy-2026|A UNESCO-framework analysis of 30 universities' GenAI policies]] adds that core ethics and governance principles are widely adopted while inclusion, equity, internet access and environmental impact are often overlooked, with many provisions remaining declarative. ## What to do when something goes wrong Treat the wrong output as a teaching moment: ask how the child would check it instead of deleting it. When a child says the companion "understood" them, that is a conversation about what a system predicting the next word can know. [[lu-ai-multimodal-writing-critical-thinking-2026|A Grade 5 multimodal writing study]] found that visualizing children's stories produced sustained gains in interpretation, analysis, evaluation and explanation, but **no gain in inference**, and children reported less need to infer implicit meaning once images made it explicit. Fluency buys belief. The same caution applies to a fluent AI answer. **Robots and role-play are legitimate vehicles for hard conversations.** [[remind-robot-mediated-roleplay-antibullying-2026|REMind]] uses robot-mediated applied drama for anti-bullying bystander intervention — children observe a bullying scenario enacted by social robots and rehearse defense strategies by puppeteering a robotic avatar; in a mixed-methods play-test with 18 children aged 9–10 it supported self-efficacy, perspective-taking and understanding the outcomes of defending. ## What the evidence does and does not support **The largest synthesis here is careful.** [[young-people-learning-generative-ai-rapid-review-2026|Arthars and colleagues (2026)]] reviewed 271 empirical papers on generative AI in PreK-12 and found **no single "GenAI effect"**: affective gains are common but weak indicators of learning, the most consistent evidence concerns improved *immediate* performance and product quality, and evidence on durable learning, [[transfer-of-learning|transfer]] and sustained self-[[self-regulated-learning|regulation]] is uneven. Outcomes depend on five entangled conditions — learner, tool, task, social arrangements, cultural and institutional context — and the framework separates four pedagogical functions, learning *from* AI, *with* it, *about* it, or *by shaping* it, and asks whether a student **surrenders, offloads or exercises agency**. For children under thirteen: *about* and *by shaping* before *from*, adult co-use wherever the child cannot yet evaluate the output, and no claims about durable learning from a good afternoon's product. **There is no study of age-based bans in this knowledge base.** That is the honest state of the evidence, said plainly rather than smoothed over: nothing here can settle whether the bans now appearing in some school systems work. The evidence that exists — on the risks, on enforcement, and on restriction versus teaching — points the same way. **Where access was taught rather than merely permitted, outcomes recovered.** [[moral-panic-genai-classroom|The study of GenAI availability versus integration]] found that making GenAI available without teaching students how to use it was associated with significantly lower performance on applied questions (ω² = 0.35) — presumably because students could not evaluate the output — whereas explicitly teaching its use for tasks like data summarizing returned applied performance to pre-GenAI levels and exceeded baseline on one harder quiz. The alternative to a ban is not laissez-faire: unrestricted access without instruction also failed. This age band can be taught — [[demir-akar-ai-media-literacy-children-2026|the fourth-grade media literacy program]] produced between-group effects of *d* = 1.12 to 1.31. ## The objections you will hear ### Should AI be banned for children under thirteen? Not as a blanket rule — and the honest reason is that no study of these bans exists, so nobody can tell you they work. The enforcement path is bad: [[bassett-ai-detectors-education-2026|detection's estimates are probabilistic and cannot be independently verified]], and [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|when false positives rise faster than true positives, stronger monitoring makes concealment relatively more attractive]]. A ban also governs the least important channel: companion apps, AI toys and family tutoring happen at home, beyond school reach, in the [[k12-teachers-ai-companion-literacy-2026|jurisdictional vacancy]] those teachers described. A school ban leaves the household ungoverned where the risk concentrates, while families with resources keep access at home. The defensible line is age tiering plus vetting plus teaching: scaffolded, often non-generative activity for roughly ages 9 to 12, and supervised generative use with a stated purpose after that. ### "Everyone else's child is already using it." Partly true: a national survey cited in [[child-safety-genai|the child-safety evaluation work]] found 72% of US adolescents had used AI companions. That figure is about adolescents, not your ten-year-old. [[young-people-learning-generative-ai-rapid-review-2026|The 271-paper PreK-12 review]] found no single "GenAI effect," so what peers do tells you what is normal, not what is safe or useful. ### "The school should handle it." Schools can act on what they provision, procure and teach: procurement conditions, child-specific hazard evaluation, data minimization and disclosure to families belong in the policy. But the home channel is where this age group's exposure mostly happens. [[family-school-autonomy-support-genai-2026|Fan, Li and Zhang (2026)]] name the synergistic family–school model and state plainly that it has no supporting evidence. It has to be built — which means telling parents what to do at home, not only what not to do. ## Checklists **For home.** - Under 9: screen-free and embodied first — tangible coding, story, role-play, hands-on making — with AI taught as something people design rather than something that knows. - Ages 9–12: scaffolded, often non-generative activities on vetted platforms, plus explicit teaching about data privacy, persuasion and what AI cannot know; generative tools come later, with a named adult role. - Decide who monitors which tool and what the school will tell you, use companions and tutors as joint activity rather than quiet supervision, and check what an app does with voice and transcript data. - Keep the child's own meaning-making visible — drafting, explaining, reflecting, making — and watch for output quality rising while unaided work falls. **For school leaders and teachers.** - Before adopting a tool, ask which child-specific hazards were tested, what it does with data by default, what it costs families if the free tier disappears, and whether it works in your families' languages. - Prefer age-tiered access rules and tool-vetting conditions to a blanket ban, and write the line down operationally — "AI to get your thinking started" is not a policy. - Monitor inputs as well as outputs, per [[stromberg-generative-ai-learning-penalty-secondary-2026|the 26,811-student secondary analysis]] in which homework scores rose while closed-book exam scores fell, and invest in teacher confidence, since [[preschool-teachers-ai-behavioral-intention-2026|AI self-efficacy predicts early-childhood adoption while AI anxiety deters it]]. - Publish an evaluation plan. The corpus has no evidence yet that bans change use or learning, and the closest analogues — detector enforcement, prohibition-heavy syllabi, untaught availability — all performed worse than expected. - Prefer unaided performance and delayed [[transfer-of-learning|transfer]] over engagement and satisfaction when reporting what a child learned. See also [[how-ai-impacts-students]], [[incorporating-ai-literacy]], [[ai-literacy]], [[parents-and-families]] and [[early-childhood-elementary-ai-education]] for the research base behind this page. --- ## [What Is the Evidence on AI Literacy Interventions in Higher Education?](https://edtechdev.github.io/aied/faqs/ai-literacy-evidence/) # What Is the Evidence on AI Literacy Interventions in Higher Education? The evidence on **[[ai-literacy|AI literacy]] interventions in [[higher-ed|higher education]] is promising, but still methodologically immature**. Across the knowledge base, the strongest recurring finding is that AI literacy develops more effectively through **active, contextualized practice with AI—especially critique, comparison, reflection, and collaboration—than through tool demonstrations or lectures alone**. At the same time, relatively few studies measure durable, independently demonstrated competence; many rely on [[self-regulated-learning|self-report]], short interventions, observational comparisons, or [[design-based-research|design-based research]]. The intervention literature converges on several design principles. ## The overall quantitative picture A [[liu-ai-literacy-interventions-meta-analysis-2026|three-level meta-analysis of 59 empirical studies]] (172 effect sizes, 7,211 participants) provides the field's clearest [[quantitative-research|quantitative]] estimate: AI literacy interventions show a **large overall effect (g = 0.837, p < .001)** — but with a wide 95% prediction interval [−0.292, 1.966], so effectiveness varies considerably across settings and may not generalize uniformly. Two moderators were significant: interventions in **East Asia and Europe outperformed those in North America**, and **knowledge-focused interventions outperformed those targeting skills, attitudes, or ethics**. Larger (though non-significant) effects appeared for mixed or reflective pedagogies, GenAI-supported tools, and performance-task outcomes. The authors argue AI literacy education should therefore move beyond knowledge toward skills, practices, ethics, and attitudes, via integrated and reflective pedagogies and GenAI-supported tools — and that [[culturally-relevant-pedagogy|culturally relevant]], context-sensitive intervention is needed.([[liu-ai-literacy-interventions-meta-analysis-2026]]) Two implications matter for practitioners: (1) the outcome you measure shapes the apparent effect — knowledge-focused interventions look stronger than skill-, attitude-, or ethics-focused ones, so evaluate what you actually care about; and (2) the wide prediction interval means a strong average does not guarantee a strong effect in any particular local setting, reinforcing the case for context-sensitive design. ## Critiquing AI rather than merely operating it In undergraduate psychology, Richmond and Nicholls had students grade a [[generative-ai|ChatGPT]]-generated media release against their course rubric, identify errors, and revise it. Students generally recognized that the output was stylistically polished but weak in accurately representing research aims, methods, and findings. Compared with the previous [[peer-assessment]] cohort, the AI-critique cohort performed modestly better on the subsequent script revision (*d* = 0.36), although there was **no significant advantage on the final video**. This provides encouraging evidence for critique-based AI literacy, but the historical comparison prevents a strong causal conclusion. Source: [[richmond-nicholls-genai-psych-feedback-ai-literacies|Using Generative AI to Promote Psychological, Feedback, and AI Literacies in Undergraduate Psychology]]. ## Embedding AI literacy in disciplinary work Beck and Brodersen's economics approach has students first solve or interpret an authentic economics problem themselves, then obtain ChatGPT's response, compare the two, critique the AI, reflect, and discuss with peers. The intervention combines AI literacy with existing [[active-learning]] techniques such as Think-Pair-Share. Its principal contribution is a transferable [[learning-design|instructional design]] rather than strong experimental evidence of learning effects. Source: [[beck-genai-literacy-economics-hands-on|Fostering Generative AI Literacy in Economics]]. ## Sustained experiences rather than one-off workshops The NC State **AI Literacy Continuum** describes progression from *Not Yet Engaged* and *Uncritical Use* through *Informed Use*, *[[critical-thinking|Critical Evaluation]]*, and *Improvement*. Its implementation involved more than 330 participants across courses and workshops. Observations suggested that brief experiences could move students toward informed use, whereas evidence of critical evaluation and improvement was more apparent in sustained, discipline-embedded experiences. However, there was **no validated pre/post measure or comparison group**, so this is practice-based rather than causal evidence. Source: [[ai-literacy-continuum-higher-education|Beyond Tool Adoption: A Practical Five-Stage Developmental Continuum for AI Literacy in Higher Education]]. ## Collaborative and cognitively active designs Hingle and Johri's [[meta-analysis-systematic-review|systematic review]] organizes AI-literacy activities using the [[icap-framework|ICAP framework]]: passive exposure, active manipulation, constructive generation, and interactive co-construction. The review finds interventions across all four modes and argues against reducing AI literacy to a one-way information session. The evidence supports designing opportunities for students to create, critique, explain, and debate AI outputs with peers. Source: [[hingle-collaborative-ai-literacy-2025|Systematic Review of Collaborative Learning Activities for Promoting AI Literacy]]. ## Teacher education: confidence versus demonstrated competence Le et al.'s design-based GenAI-literacy intervention, piloted with 14 master's students and evaluated with 29 [[teacher-education]] students, increased reported AI-competency [[self-efficacy]] and produced shifts toward more critical [[pedagogy|pedagogical]] consideration of GenAI. However, ethics gains became only marginally significant, and the small design-based research study means that conclusions about objective competence remain limited. Source: [[genai-literacy-training-teacher-education-dbr-2026|Development and Evaluation of AI Literacy Training for Teacher Education Students]]. ## AI literacy training does not eliminate AI-related errors A particularly important caution is that **training students to prompt or use AI better is not equivalent to making them reliably critical users**. In the [[ai-sycophancy|contextual-sycophancy]] experiment, AI-literacy and [[prompt-engineering|prompting]] training reduced the tendency of the model to mirror participants' reasoning errors, but did **not eliminate downstream error propagation**. This suggests that literacy education cannot carry the entire safety burden; tool design and system-level safeguards are also needed. Source: [[contextual-sycophancy-ai-literacy|The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration]]. ## Measuring AI literacy matters The measurement literature reinforces this caution. The validated **GLAT** performance test predicted performance on an AI-assisted higher-education task, whereas a [[self-report-measures|self-reported]] ChatGPT-literacy measure did not. More broadly, the knowledge base reports substantial discrepancies between people's reported and demonstrated AI competence. Consequently, an intervention that raises confidence, attitudes, or perceived literacy should **not automatically be interpreted as improving AI literacy itself**. This points to the importance of [[educational-measurement|educational measurement]] and [[assessment-validity|assessment validity]] when [[ai-ed-evaluation|evaluating AI]]-literacy interventions. Newer work sharpens both *what* should be measured and *how*. Burriss et al.'s classroom study of 22 eleventh-grade students composing video public service announcements about AI ethics argues that existing AI-literacy scales and competency frameworks assume individually measurable performance and therefore exclude collaborative and creative expression; they propose that student reflections, [[multimodal]] artifacts, and civic discourse complement conventional measures. Their finding that students developed critical AI-ethics understanding through collaborative [[critical-pedagogy|critical pedagogy]] challenges skill-list conceptions of AI literacy — and, with them, the instruments built on those lists. It also shows the direction of travel: 15 of 18 end-of-unit responses affirmed new learning, while the authors note the project is resistant to traditional [[summative-assessment|summative assessment]]. See [[burriss-multimodal-composition-critical-ai-literacy-2026|Multimodal composition as a form of critical AI literacy pedagogy]]. A separate caution comes from acceptance research. Wang, Chuang, and Wu surveyed 831 Taiwanese [[k-12]] students and teachers after roughly 30 minutes with an age-tiered AI-literacy guidebook and found a stable four-factor [[technology-acceptance-model|acceptance]] structure — but perceived playfulness was the strongest correlate of behavioral intention, and the study measured short-term acceptance, not [[learning-gains|learning gains]], [[ethics|ethical]] competence, or actual use. The authors are explicit that acceptance constructs "must not be used to infer learning effectiveness or implementation success". See [[age-tiered-ai-literacy-guidebooks-2026|Measuring acceptance of age-tiered AI literacy guidebooks]]. Sources: - [[jin-glat-genai-literacy-assessment|GLAT: The Generative AI Literacy Assessment Test]] - [[ai-literacy-assessment-misalignment|AI Literacy Assessment: Self-Reported vs Performance Misalignment]] ## Diagnosis before instruction: students' existing AI cognition Where an intervention *begins* matters as much as its content. Şan and Orhan Karsak used a psycholinguistic Word Association Test with 436 Turkish undergraduates (1,376 coded responses, inter-rater κ = 0.87) to map their cognitive representations of AI. Associations clustered around instrumental utility — convenience (f = 145) and speed (f = 110) — while algorithmic transparency, privacy, and [[governance]] concepts were not merely rare but structurally isolated from the dominant utility cluster (Kendall's τ = −0.819). The authors argue that delivering an ethics module into a cognitive architecture with no existing schema to receive it is likely to fail: curricula must build explicit **bridges** from everyday tool knowledge to ethical and governance frameworks, and micro-credentials should sequence content to close that structural gap rather than reproduce what students already know. This is a **needs-assessment** argument for intervention design, and it is one of the few studies to supply a data-driven diagnostic before instruction. See [[san-orhan-karsak-ai-cognition-micro-credentials-2026|Knowing its name, not its nature]]. A contrasting, [[situated-learning|situated]] view comes from Dai and Chan's seven focus groups with 28 postgraduate researchers, who enacted [[ai-literacy|AI literacy]] differentially across the research workflow — heavier use in low-stakes procedural tasks (formatting, translation, explanation), markedly more caution where scholarly contribution was at stake — and drew their own boundaries between assisting versus substituting their reasoning. Their conclusion is that AI literacy is better treated as a situated capacity than a static competency list, and that responsible-use guidance should [[scaffolding|scaffold]] each literacy dimension across real tasks rather than enforce binary rules. See [[dai-chan-responsible-genai-research-ai-literacy-2026|Shaping responsible GenAI use in research through AI literacy-oriented guidelines]]. ## Credentials and program-level evaluation AI literacy is increasingly delivered through certificates and micro-credentials, and this is where evaluation is thinnest. Wu and Li built an expert-weighted evaluation indicator system (AHP / Fuzzy-AHP with Monte Carlo robustness verification, 18 experts) for AI certificate programs and found that **faculty professional competence showed the largest importance–satisfaction gap** — high perceived importance (0.1976) against low current satisfaction — identifying faculty capability, not technological infrastructure, as the primary constraint on pedagogical transformation; cross-cultural adaptability received the lowest strategic weight (0.0735), signaling a technology-first bias in early-stage credentialing. See [[wu-li-evaluation-indicator-ai-certificate-programs-2026|Evaluation indicator system for AI certificate programs]]. This foregrounds a question the intervention literature has barely tested: whether a credential certifies durable [[ai-literacy|AI literacy]] or merely records completion, and whether the [[teacher-education|educators]] delivering it are themselves equipped to judge. ## Overall assessment of the evidence Overall, the evidence can be characterized as **moderate for instructional design principles, but still weak-to-moderate for causal effectiveness**. The most defensible current model for [[higher-ed|higher education]] is to: 1. Embed AI literacy inside authentic disciplinary tasks. 2. Have students form an initial judgment or solution before consulting AI. 3. Require comparison, verification, and critique of AI outputs. 4. Incorporate peer dialogue and reflection. 5. Repeat these experiences over time rather than relying only on one-off workshops. 6. Assess students with **performance-based tasks requiring independent evaluation and responsible judgment**, rather than relying only on confidence or self-report surveys. Important gaps remain. The literature still needs more **multi-institution randomized or strong quasi-experimental studies, validated common outcome measures, delayed tests of retention and transfer, and evidence about whether literacy gains persist when students encounter new AI models, disciplines, or unfamiliar AI failure modes**. In short, current evidence supports treating AI literacy as a **discipline-embedded critical practice**, not simply as knowledge about AI or proficiency in prompting. The instructional case for critique, comparison, reflection, collaboration, and repeated authentic practice is increasingly coherent, but the evidence that particular interventions produce durable and transferable AI literacy remains less established. For how that conclusion translates into concrete course design, see [[incorporating-ai-literacy|How should I incorporate AI literacy into my course?]]; for where the evidence base itself remains thin, see [[research-gaps-aied|What are notable gaps in the research literature on AI in Education?]]. --- ## [How Can AI Save Me Time as an Instructor?](https://edtechdev.github.io/aied/faqs/ai-save-instructor-time/) # How Can AI Save Me Time as an Instructor? **AI is best suited to repeatable, lower-stakes drafting and transformation work, while consequential educational judgment remains with the instructor.** Productive uses include first drafts of lesson materials, examples, discussion questions, [[formative-assessment|formative]] quizzes, alternative explanations, rubrics, feedback suggestions, summaries, differentiated versions of materials, and routine administrative language. The key mechanism is **reallocation, not reduction**: AI frees time that instructors redirect toward higher-value instructional work — the same time then buys more one-on-one student interaction, deeper feedback, and higher-order teaching. ## What the evidence shows ### Lesson preparation: roughly a 30% time saving, with quality held The article [[ai-changing-teaching-workflows|How AI Is Changing Teaching Workflows]] summarizes an English [[rct|randomized trial]] involving 259 science teachers in which teachers using ChatGPT spent about **69% as much time** on lesson preparation as the control group — roughly a **31% reduction** — with no detectable loss in material quality according to blind expert reviewers. Teachers generally **reallocated** the saved time to other instructional work (planning, grading, student-facing activities) rather than simply eliminating work. A companion dataset of 104,000+ messages from 15,000+ educators showed the average [[teacher-role|teacher]] prompt touches **1.7 categories at once** (e.g., a single request combining lesson plan + differentiation + formative quiz), so AI often surfaces instructional elements the teacher didn't have to ask for. ### Where AI quality holds — and where it doesn't AI-generated materials aren't uniformly as good as human ones; the value depends on the task and level: - **Strong:** lesson conclusions/exit tickets (AI versions preferred **59.7%** of the time over professional designs), high school content (**59.2%**), and teaching *outside* your expertise (bigger time savings when you're less confident in the subject). - **Weak:** elementary-level materials (humans preferred ~65% of the time for developmental appropriateness), and targeted [[multilingual-learning|multilingual]]/[[special-education]] [[scaffolding|scaffolds]] (AI is "neutral" but not nuanced). So the safest, highest-value uses are **structured, well-specified materials you can review** — not fine-grained developmental or culturally-scaffolded content you'd need to rebuild anyway. ### Feedback and grading: the reallocation payoff is real A large-scale Brazil experiment across **178 schools and ~19,000 high school seniors** tested AI-automated essay feedback: - AI feedback produced **identical [[learning-gains|learning gains]]** to human graders — who cost ~\$0.85/essay and added **zero incremental benefit**. - Students in AI classrooms had **~35% more one-on-one conversations** with teachers about writing, and wrote **30% more essays**. - Teacher at-home work hours dropped **20%**; teachers reporting time as "very insufficient" fell from 23% to 9%. - **The largest learning gains were on the most complex writing task** — precisely what AI is *least* equipped to evaluate — because AI freed teachers to focus on higher-order instruction. **Caveat:** the bottom quartile showed no improvement — freed time alone wasn't enough for the students who needed the most support, so savings must be paired with intentional, [[equity-in-ai-education|equitable]] reallocation. ### A counterweight: the efficiency-gain illusion Time savings are easy to overestimate — including by the instructors experiencing them. [[efficiency-gain-illusion-ai-overreliance|Across three pre-registered studies (N = 2,691)]], people systematically **underestimated how often they actually used AI** and **overestimated the time and effort it saved**, believing tasks were faster and easier even when objective measures showed no difference; prior AI use in a session predicted further use, entrenching the miscalibration in a self-reinforcing loop. The practical implication is to treat [[generative-ai]] time savings as a claim to check against actual workflow data rather than accept from feel, and to build the [[metacognition|metacognitive]] calibration that lets you notice when a task has genuinely sped up. Not every perceived saving is illusory, though: [[pishtari-teacher-ai-training-learning-design-2026|a within-subjects study of 13 higher-education teachers]] found AI access sharply lowered the perceived cognitive effort of [[learning-design|learning design]] (median 6.33 to 3.33, p = .004) — but the [[prompt-engineering|prompting]] training layered on top did **not** further improve design quality, a reminder that effort saved is not the same as capability gained. ### Automated marking: promising, but not portable Automated marking is the strongest time-saving claim — and the one that needs the most caution. [[opraise-automated-marking-ai-assessment-2026|The OpRaise benchmark]] tested three frontier models against **761 authentic undergraduate Psychology essays from 125 students across three UK universities**, with 27 prompt configurations per model. Agreement with human degree-classification bands ranged from **35% to 65% by institution**; AI marks compressed toward the middle so the strongest and weakest essays were marked least accurately; the models agreed with each other far more than with human markers; and the authors' central finding is that **evidence collected in one context does not transfer to another**. Automated [[automated-essay-scoring|essay scoring]] may save time, but only institutions prepared to validate it locally and keep final authority with people should rely on it. See [[redesign-assessment-ai-era]] for the assessment-design counterpart. ### What teachers will and will not delegate [[reichert-human-centered-llm-chatbot-design-teachers-2026|Participatory design work with six secondary teachers]] found they designed [[generative-ai|generative AI]] as a **"bounded expert"** — capable within a strictly defined domain and under human supervision. They welcomed AI help for presenting content, supplying practice problems and giving [[formative-assessment|formative]] [[feedback]], but **refused to delegate objective-setting or [[summative-assessment|summative]] [[assessment]]**, and insisted on teacher override for ambiguous cases. That boundary marks where a genuine time saving ends and an unacceptable transfer of professional judgment begins — a useful test for any task you are about to hand over, and a reason to keep [[human-in-the-loop-ai|human oversight]] explicit. ## Concrete ways to use AI to save time - **Draft lesson materials** — first-draft slides, handouts, worksheets, or a sequence of examples, which you then review and edit. - **Generate discussion questions, formative quizzes, and exit tickets** from your own notes or readings. - **Create alternative explanations** — re-explain a concept at a different level, in a different metaphor, or for a different audience. - **Draft rubrics and feedback suggestions** — AI can propose rubric criteria or a first-pass feedback comment you then personalize; the [[ai-feedback-quality|AI Feedback Quality]] synthesis cautions that speed and volume don't guarantee usefulness, so keep the [[pedagogy|pedagogical]] judgment yours. - **Summarize and differentiate** — condense long sources into study guides, or produce differentiated versions of a task for varied readiness levels (review carefully for multilingual/special-education nuance). - **Write routine administrative language** — announcements, syllabus boilerplate, form letters, and correspondence. ## What to be careful about - **Don't assume the exact magnitude transfers to college.** The 31% figure is from K–12 [[science-education|science teaching]]; the safer general lesson is to use AI for a first pass and spend human time where disciplinary judgment, relationships, interpretation of student thinking, feedback prioritization, or high-stakes decisions matter most. - **The [[prompt-engineering|prompting]] gap:** most teachers in the research didn't iterate with follow-up prompts — they took the first result and edited manually. Investing a little time in [[ai-literacy]] and prompt refinement pays off in output quality. - **The assessment trap:** nearly half of educator–AI conversations involved assessment, but some requested grading without specifying rubrics or criteria — unguided AI assessment risks inconsistency and bias, so keep [[human-in-the-loop-ai|human oversight]]. - **Equity divides:** freed time is only net-positive if it isn't spent at the expense of multilingual learners or students with disabilities, and if under-resourced instructors use it to upgrade practice rather than merely keep pace. --- ## [How Should We Design and Facilitate Asynchronous Online Courses When AI Can Do the Work?](https://edtechdev.github.io/aied/faqs/asynchronous-online-courses-ai/) # How Should We Design and Facilitate Asynchronous Online Courses When AI Can Do the Work? You have already made the discovery that forces this question. Somewhere in your asynchronous course is an assignment whose first draft is now free — a discussion post, a reflection, a case analysis, a problem set — and the work that arrives is fluent, on-topic and tells you nothing about what the student can do. You are not in the room to watch the thinking happen, and the artifact that used to stand in for it no longer does. That was the version of the problem in which a student decides, at two in the morning, to let a model write the post. The harder version is already here. In three demonstrations on a live undergraduate psychology course, autonomous agents logged into the learning management system, read the course materials, and completed unproctored assessed work with no student involved: two quizzes finished in about **12 minutes** and in **under 5 minutes**, both scoring **10/10**, and a discussion post in which the agent mined its peers' posts and then fabricated a credible first-person life history to answer them. The wider record the authors assemble covers at least **15 documented agent runs** across Canvas, Moodle and Brightspace using **seven agent tools**. Their argument is that this is an [[assessment-validity]] problem rather than only an [[academic-integrity|integrity]] one: agent completion removes the assumption that the submitted work was produced by the person whose learning is being assessed ([[ai-agents-complete-lms-assessment-validity-2026|Hadjisolomou and El-Haddad, 2026]]). The bottom line: an asynchronous course has no presence or process visibility to fall back on, so the design has to manufacture them. Move the evidence of learning away from the finished product, keep one measurement the student must produce unassisted, guardrail whatever AI you supply, and pace the term deliberately — self-regulation is what self-paced formats quietly assume and rarely teach. Put plainly: stop designing around the assumption that producing the artifact demonstrates learning. That is not an argument for banning AI or for making tasks AI-proof, but for designing the course so that having AI produce the artifact is insufficient to accomplish the learning. ## The short version - **Make the thinking the deliverable.** Staged drafts, decision logs, and self-explanation are harder to outsource than a final artifact, and they show you the reasoning you otherwise cannot see. - **Keep one unassisted measurement** on the assessments that certify competence. Population data show the AI-access effect on retention disappears under proctoring, which tells you the shortcut is off-platform and conditional on the conditions you set. - **Use [[asynchronous-oral-assessment-2026|asynchronous oral assessments]]** where the stakes justify them — just-in-time prompts, time-limited unrevised recordings, rubrics embedded, transcripts automatic. - **Guardrail the AI you provide.** Hint-not-answer tutoring eliminated the exam penalty that unguarded access produced in a randomized trial; supplying a model without constraints is the version that harms learning. - **Embed the support inside the course** rather than linking out to a generic chatbot. The Open University's embedded assistant doubled time on task in its trial; the same institution's experience with an external, unattached chatbot went the other way. - **Pace the term on purpose.** Early-warning signals exist up to 7–8 days before module deadlines, and self-regulation and well-being decline across a term in step with assessment clustering — so spacing deadlines is a design decision with measurable consequences, not a scheduling detail. - **Facilitate discussions sparingly, and sequence them deliberately.** LLM facilitators are markedly more eager to intervene than human facilitators, and the summary tools that genuinely help students navigate a large forum do not keep participation from declining on their own. Ask for a committed position, a challenge, and a documented reconsideration rather than a post count. - **Commit before consulting.** Have students produce their own position, prediction, outline or decision first, then bring in the model for critique; the thinking that has to be learned should happen before the tool is opened. - **Pair a vulnerable task with a twin.** Keep the take-home analysis and add a closely scheduled second task that assesses the same outcome in a way the first cannot be delegated for. - **Give the AI a stated role,** and be able to finish the sentence: *the AI's job here is X, the student's job is Y.* If Y is thin, the activity is not ready. ## Decide what the assessment is actually measuring The reason an asynchronous course is more exposed than a face-to-face one is not that students there are less honest. It is that the course's evidence of learning is the submission, and the submission is now cheap to produce. Two pieces of causal evidence show what that costs. In a randomized trial with roughly 1,000 high-school mathematics students, unguarded AI assistance raised practice performance by **48%** while reducing unassisted, closed-book exam scores by **17%** — students who never had access outperformed those who did. The critical detail is the fix: a [[guardrails|guardrailed]] tutor that gave hints rather than answers eliminated the harm. The problem was never the model's presence; it was the absence of a constraint on what it would do. Population-scale behavioral evidence points the same way. Across **3.2 million ALEKS interactions**, study time on AI-susceptible problems fell **26.9%** after ChatGPT's public release, with a **25% decline in the odds of answering proctored retention items correctly** — an effect that vanished once proctoring was in place. That disappearance is the important part for course design: it locates the shortcut off-platform and shows the harm is a property of the conditions you set, not of the learner's character. There is also a subtler failure to design against. [[metacognitively-discordant-completion-genai-2026|Metacognitively discordant completion]] names the state of submitting correct, complete work while privately knowing that understanding never arrived. In a classroom you might catch that in a student's hesitation; in an asynchronous course the submission is your only channel, and a correct submission reads as success. [[assessment]] design is therefore the first lever, not the last resort. ## Decide what counts as evidence when there is no room to walk into If the finished product no longer proves the thinking, the evidence has to come from the process or from a condition the student cannot delegate. **Process-revealing artifacts** are the cheapest change and work at scale: staged submissions with a required intermediate decision log, a one-paragraph self-explanation attached to each answer, an annotation of a source, a plan revised after feedback. These capture [[cognitive-offloading|the reasoning]] rather than its residue, and they cost the instructor comparatively little to scan. **Asynchronous oral assessment** is the strongest option when competence must be certified. [[asynchronous-oral-assessment-2026|Pentland, Lowenthal and Krier (2026)]] evaluated web-based assessments in which prompts are delivered just-in-time, students record brief, time-limited webcam responses they cannot revisit, and instructors grade against embedded rubrics while transcripts generate automatically. Across two studies — an intermediate accounting pilot and a data analytics course — students scored higher on these assessments than on in-person multiple-choice exams (significant in the second study, a positive but non-significant trend in the first), with moderate cross-format correlations supporting convergent [[assessment-validity|validity]]. Students reported preparing differently, using more [[active-learning|active]] study strategies, and treating the format as professionally relevant. It answers the asynchronous problem directly: the thinking is performed live, at a time and place of the student's choosing, at administrative cost that does not scale with cohort size. **One unassisted measurement** belongs in any course whose grade certifies knowledge. The ALEKS result above is the argument: when proctoring was in place, the retention gap disappeared. Close the loop by telling students why the condition exists — the reasoning belongs in your [[course-ai-policy|writing a course AI policy]]. ## Decide the sequence, not just the artifact Once the finished product is unreliable evidence, the order in which the work happens becomes part of the assessment design. [[brcic-effortless-trap-productive-struggle-2026|Brcic and Frljic (2026)]] make placement the center of the argument: allow-or-ban is a false dichotomy, and the design question that matters is **where the tool sits**. The causal evidence they collect shows the outcome flipping on placement alone — the same unguarded helper that left high-school students about **17% worse** on an unaided exam did no harm once rebuilt to withhold answers, while a well-engineered [[intelligent-tutoring|tutor]] roughly **doubled** learning. Their six-move frame for placing the tool (prime, probe, point, attach, strengthen, test) is one usable menu, and their one-line diagnostic is the one to keep: *if letting AI in makes the task feel effortless, it is in the wrong place.* Poorly placed assistance does not merely fail to help; it leaves an illusion of learning that collapses on the unaided task — the state [[metacognitively-discordant-completion-genai-2026|metacognitively discordant completion]] names for a submission that is correct and not understood. For an asynchronous course that translates into a sequence you can write into the assignment itself: **Think → Commit → Use AI → Critique → Revise → Explain.** Students attempt the intellectual work and commit to an initial interpretation, prediction, outline or decision. Only then does the model enter, for feedback, alternatives, counterarguments or critique. Then students evaluate what it produced, revise, and explain what changed and why. The commitment step is what makes the rest assessable: an uncommitted student has nothing to revise, and an artifact written from scratch by a model has no revision history at all. One caution about the overcorrection: collecting every prompt and forcing documentation of every keystroke produces workload without evidence. Ask for meaningful decision points instead — which evidence was selected, which alternative rejected, what changed after critique — and for the reasoning you would actually read. ## Pair a vulnerable task with a twin The other structural option is to stop choosing between a pedagogically valuable task and a defensible one. [[roe-assessment-twins-2026|Roe, Perkins and Giray (2026)]] propose **assessment twins**: pair the AI-vulnerable task — a take-home essay or analysis — with a second, less vulnerable task that assesses the *same* outcomes, scheduled closely enough for cross-verification, and marked interdependently. Their mapping runs across Messick's six strands of validity evidence, and the design process has three steps: identify the vulnerabilities, align outcomes and choose the twin, then develop marking that connects the two. The take-home analysis keeps its value for extended inquiry; its twin might be a short case variation, an explanation of one key decision, or a brief oral or multimedia defense. The companion question is ownership rather than authorship. [[coauthorship-integrity-reconceptualizing-assessment-validity-for-the-age-of-gene|Ebrahimzadeh, Shibani and Buckingham Shum]] argue that coauthorship with generative AI undermines several forms of validity evidence, and propose **coauthorship integrity** as a source of validity evidence in its own right — violated when a student submits AI-generated content they do not understand. To check ownership at scale rather than by inspection, they report work on an **AI Viva**: a conversational agent that runs a hybrid viva voce, asking comprehension questions of controllable type and complexity, validated in depth by expert educators and assessment specialists. That is worth weighing against proctoring everything: a spoken defense of one's own reasoning scales in a way that room monitoring does not, and it produces evidence about understanding rather than about who was in the room. ## Decide what the AI inside your course is allowed to do The comparison that matters is not AI versus no AI. It is embedded, constrained AI versus an unattached chatbot, and the two produce different results. The [[new-systems-of-learning-for-distance-learning-institutions-a-six-study-review-of|Open University's AIDA assistant]] is the best-documented embedded case: six iterative design-based studies over 18 months with 498 students and 20 staff, at an institution serving 200,000-plus learners across more than 50 countries. In an exploratory randomized trial, students using AIDA spent **twice as long** and visited more pages in their course than the control group, and **96%** wanted the assistant available in their formal studies. Purpose-built and in-environment beat generic and external. [[lock-integrating-ai-online-learning-higher-ed-2025|Lock, Arteaga and Johnson's (2025) review]] — 63 citations across 32 countries — adds a social condition to the design condition: students who used ChatGPT *alongside* teacher tutoring perceived greater [[learning-gains]] than those who used it alone. The assistant supplements the instructor or it replaces the relationship, and only one of those is a course design. Two cautions are worth taking seriously. KhanMigo's failure was not technical: learners simply did not engage with the chatbot and evidence of gains was limited, which points at [[governance|organizational readiness]] and instructional fit rather than model quality. And the instructor side is not automatically ready: in a comparison of South African teacher preparation, self-reported TPACK for AI-integrated science teaching was **64.0%** at a campus-based university against **47.4%** at a distance university, with pedagogical knowledge the weakest domain in both ([[online-teaching-and-learning]]). Deploying an assistant into an asynchronous course does not train the people who must judge its output. ## Decide how the asynchronous discussion is actually facilitated Discussion forums are where asynchronous courses either build [[community-of-inquiry|community]] or quietly become submission boxes. Two findings should shape what you automate there. First, **timing is the hard part, not topic detection.** [[llm-facilitation-timing-online-discussions|Tsirmpas and colleagues]] built the PEFK corpus to compare facilitation datasets, then ran the first survey of facilitation *timing* with expert human facilitators and LLM judges. Humans were more cautious about intervening; LLMs were excessively eager. Both were more certain when judging that facilitation was **not** needed than when judging that it was. Trained classifier models outperformed the LLM setups, and even then the existing datasets capped performance. The practical reading: do not hand autonomous moderation to an agent, and if you use AI in a discussion at all, configure it to stay quiet by default. Second, **AI helps students navigate a forum without making them participate.** With 128 university students across three iterations, [[hao-peer-exposure-bridging-social-capital-ai-summaries-2026|Hao and Cukurova (2026)]] found AI-generated discussion summaries broadened exposure to peers' contributions and strengthened network connectedness — the weak ties social-capital theory calls bridging — by lowering the effort of finding meaningful posts. They did not stop viewing activity declining across the course. Summarization is a navigational [[scaffolding|scaffold]], not a substitute for generating participation. The structure that mitigates this is a sequence rather than a thread: **Position → Challenge → Reconsideration.** Students commit to an interpretation grounded in the course material, then meet another perspective, counterexample or critique, then explain whether and how their reasoning moved. What you grade is the movement between ideas, not the number of posts — and a submitted post plus two replies no longer demonstrates it, since the third demonstration above has an agent mining a peer's posts and impersonating a classmate's life history to order. Because what an agent imitates is surface behavior, the instructor's own contributions become the scarce resource. Replying mechanically to dozens of individual posts is the least valuable form of teaching presence; synthesizing across the discussion is the most — *three assumptions keep appearing in your analyses; several of you read this evidence differently; this argument is convincing until we introduce this counterexample.* That is [[teacher-role|orchestration]] rather than message production, and it is the part no agent in this literature performs. The framework question underneath both is accountability. [[reconceptualizing-community-inquiry-generative-ai|Ba, Gašević, Lim and Anderson (2026)]] argue that generative AI unsettles the [[community-of-inquiry]] assumption that presence indicators can be attributed to humans at all: they treat GenAI as an epistemic condition, and presences as sociotechnical accomplishments whose relationship to inquiry quality depends on where human accountability sits. The design implication for an asynchronous course is concrete — decide and state who is accountable for what in each exchange, rather than assuming presence arises because a forum exists. ## Decide how students keep pace without a room to walk into Self-paced formats assume [[self-regulated-learning|self-regulation]] and rarely teach it. The evidence on which behaviors actually correlate with staying on task is unusually practical. Surveying 530 college students with association-rule mining and clustering, [[decreasing-digital-distraction-college-online-learning-2026|Shi et al. (2026)]] found that self-regulated learning behaviors — **goal setting, environment structuring and time management** — co-occurred most consistently with low digital distraction, along with learner–instructor and learner–content [[student-engagement|engagement]] and technical competence. Notably, reliance on peer help-seeking and learner–learner engagement appeared *less* often in the low-distraction profiles. Structure the environment and the schedule before you design another group activity. Self-regulation is also not a fixed trait you can assume or dismiss. Across a full semester with 75 first-year students, [[song-genai-learning-partner-srl-over-time-2026|Song et al. (2026)]] found SRL functioning as both a stable aptitude and a fluctuating state: individual baselines held steady while metacognitive knowledge and [[well-being]] declined systemically across the term, driven by curriculum demands such as major assessment deadlines. Clustering deadlines may be an efficient administrative choice, and it is also a measurable tax on the regulation the format depends on. The intervention window is knowable. [[zhang-ml-student-progress-programming-2026|Zhang, Jeffries and Koprinska (2025)]] showed that interpretable machine learning on content-interaction logs predicts module-level progress and flags dropout outcomes up to **7–8 days before module deadlines** in large-scale online programming courses. That is enough notice to send a specific nudge to a specific student, and it beats discovering the failure at grading time. For [[adult-learning|adult]] and distance learners, the AI-ALOE design guidelines add the constraint that matters most: mobile access, offline capability, and genuine asynchronous availability, since these learners study in fragments of time between other obligations. ## What AI can take off your plate as the designer Two uses have reasonable evidence behind them, and both target instructor workload rather than learner thinking. **Production cost.** [[mooc-to-maic|MAIC]] reframes the MOOC's "one video for N students" broadcast as "N agents for one student," using specialized teacher, assistant, classmate and analyzer agents on a shared model foundation. Its authors report collapsing course production from roughly **\$25,000 and 60 hours** per MOOC to **under \$2 and 30 minutes**, piloted at Tsinghua across two courses with **100,000+ learning records from 500+ students**, and released as open-source OpenMAIC. Personalized media is the same idea at the asset level: in a large online course, [[personalized-ai-generated-videos-preference-2026|Tomlinson et al. (2026)]] found students preferred AI-generated personalized videos over non-personalized human-recorded ones, with the personalization effect outweighing the value placed on a human presenter. **Forum navigation.** The summary scaffold above reduced the cost of locating good contributions without burdening instructors with the summarizing. **What not to automate.** Deciding whether a piece of writing is the student's own thinking, judging whether a discussion needs intervention, and calibrating difficulty are the tasks the evidence says are either unreliable or accountability-bearing. Generated material also needs pedagogical review before it carries credit — cheap production is not the same as sound design. ## Give the AI a stated role, then check the student still has one "Students may use AI" is too broad a specification to design against. Naming the role is what makes an activity designable, and the roles carry consequences for presence: a study of generative AI in marketing education distinguishes **tutor, teammate and tool**, shows each influencing teaching, social and cognitive presence differently, and lists the familiar ethical exposures — data privacy, plagiarism, dependency and assessment fairness ([[genai-marketing-education-roles-2026|GenAI in Marketing Education]]). The same model can be a Socratic questioner before an assignment, a critic after a first draft, a simulated stakeholder during a case analysis, an opponent whose argument must be rebutted, a hint-giving tutor during practice, or an editor brought in only after the substantive reasoning exists. Those are different activities, and conflating them is how "AI is allowed" quietly comes to mean "nothing was designed". [[scaffolding]] supplies the test: support should help learners do what they cannot yet do alone, and should fade as competence develops. A model that keeps supplying complete solutions is not a scaffold — it is doing the task. So for every AI-enabled activity, the designer should be able to complete this sentence without hesitation: *The AI's job here is \_\_\_\_\_\_, while the student's job is \_\_\_\_\_\_\_\_ .* If the second blank contains little meaningful thinking, the activity needs redesigning rather than a stricter policy. ## Teach evaluation, not just operation [[ai-literacy]] is not the ability to operate a chatbot; it includes critical evaluation, [[metacognition|metacognitive]] judgment and knowing when not to use the tool at all — and the knowledge base is consistent that self-reported confidence is a poor proxy for demonstrated competence. That points at a change any asynchronous course can adopt immediately: **sometimes give students the AI output yourself.** Hand over two competing answers and ask which is stronger and why. Ask for the weakest claim, the missing assumption, the unverifiable source, the plausible explanation that is nonetheless wrong. Ask what evidence would change their mind. Framed that way, the capability being graded is no longer only *can this student produce an answer?* but *can this student recognize whether an answer deserves to be trusted?* — closer to what the discipline actually requires, and far harder to delegate. The [[verify-ai-output|verification practices]] and the AI literacy material are where the technique lives; what matters here is putting evaluation inside the graded task instead of leaving it as advice. ## The objections you will hear **"These are adults who chose an asynchronous course. If they let AI write it, that is their decision."** The choice argument would hold if the course certified nothing. It does. The ALEKS evidence shows the shortcut produces an appearance of competence that does not survive a proctored check, and the cost lands on the student later — in the next course, the licensure exam, or the job. There is an equity edge too: the learners most likely to offload are often those with the least time, which is precisely the group an asynchronous course exists to serve. **"Detection is the answer."** Detection is contested, and the proctoring result above shows why design beats policing: the harm disappeared when the conditions changed, not when policing intensified. Detection also carries false-positive costs and turns instruction into an arms race. Redesign the task and keep one unassisted measurement instead — see [[reduce-ai-cheating]] and [[redesign-assessment-ai-era]]. **"Teaching presence is impossible at a distance, so async is inherently inferior."** Presence in an asynchronous course is designed rather than implied, and the CoI reconceptualization above says its indicators are sociotechnical accomplishments in the GenAI era. The AIDA trial is a useful counter-example: embedded generative support doubled time on task in a course with no synchronous meetings at all. **"Proctoring is surveillance and I will not impose it."** A defensible position, and it does not leave you without options. Process-revealing artifacts, asynchronous oral assessment, self-explanation and the AI Viva described above all capture reasoning without monitoring anyone's room. If you do use a proctored condition, say so in the syllabus, explain the reasoning, and keep it to the assessments that certify competence. **"Then make it all synchronous and proctored."** Many students choose asynchronous study precisely because they have employment, caregiving responsibilities, disabilities or geographic constraints that make synchronous attendance difficult, and detection carries its own equity costs: tools that flag non-native writers disproportionately produce false positives that penalize honest work ([[ai-detection]]). The proportionate alternative is a small number of verification moments inside an otherwise flexible course: a short recorded explanation, a personalized application, a response to an instructor-selected question, an annotated decision trail, a low-stakes individual check. Verification should raise the validity of your evidence without removing the flexibility that made the format worth offering. ## What the evidence does not settle The [[ai-distance-education-systematic-review-2026|systematic review of AI in distance education]] (56 articles, 2020–2025) covers personalization (24 studies), assessment and feedback (19), human–AI interaction (17) and governance and equity, and concludes that the base is short-term and cross-sectional, with little longitudinal work. The [[ai-student-engagement-online-learning-review-2025|review of AI and student engagement in online learning]] (24 studies) is limited to one database, treats engagement only, and explicitly conflates synchronous with asynchronous contexts — so its conclusions should not be read as asynchronous-specific. Asynchronous oral assessment rests on two studies in two courses. MAIC is a pilot. Facilitation-timing datasets cap performance even with trained classifiers. And no study here follows an asynchronous cohort long enough to show whether redesigned assessment produces durable learning rather than better evidence of it. Treat all of it as strong enough to change your next course and too thin to justify a policy claim. ## Do this week **In ten minutes:** pick your highest-stakes assessment and add one condition the student completes unassisted. Tell them why it exists. **Before the next assignment goes out:** take the assignment AI now completes end-to-end and rewrite it so the process is the artifact — a decision log, a staged draft, a self-explanation, a comparison of two attempts. If the task can be finished without any of that reasoning, the reasoning was never required. **This term:** convert one major assessment to an asynchronous oral defense with embedded rubrics; set the AI permission per task in the syllabus and place it where students will actually read it; space the major deadlines instead of clustering them; and open a 7–8-day pre-deadline window in which you contact students whose interaction data has gone quiet. **Run three questions over your weakest assignment.** (1) Could an AI system complete this activity without the student understanding the material? (2) What cognitive activity is supposed to produce the learning here? (3) What evidence will show that the student actually performed that activity? If the first answer is yes and the other two are hard to answer, the problem is the learning design rather than the AI policy — and no wording of the policy will fix it. The goal is not a course that AI cannot participate in. It is a course where AI can participate **without displacing the learning the course exists to produce**. **Where to go next:** the course-level rules belong in [[course-ai-policy]]; the assessment rebuild is covered in [[redesign-assessment-ai-era]]; the reliance problem underneath it in [[reducing-over-reliance]]; instructor workloads and what AI realistically removes in [[ai-save-instructor-time]]; and the design-level view of embedding AI into a learning experience in [[designing-ai-into-learning]]. --- ## [How Do I Write a Course AI Policy and Communicate It to Students?](https://edtechdev.github.io/aied/faqs/course-ai-policy/) # How Do I Write a Course AI Policy and Communicate It to Students? You have one syllabus, one assignment sheet, and a room full of students who have already written their own private rules about [[generative-ai|generative AI]]. Whatever you publish has to be fair to the student who plays by the rules and to the one who does not, and has to be something you can actually apply to forty submissions without becoming either a pushover or a police officer. The bottom line: write fewer rules than you are tempted to, tie each one to the specific thing you are assessing, name a single standard place where students declare AI use, and make your enforcement route evidence you can see with your own eyes rather than a detector score. A policy that explains itself is doing assessment design and communication work at once — which is why it takes an afternoon and saves a semester of arguments. This page covers the course-level document for one module, course or program strand. The assumption throughout is that you have decided what you are assessing; the question is how to say so and make it stick. ## What course policies actually look like right now A content analysis of **116 institutional GenAI policies** from 131 US R1 universities and **98 computer-science syllabi** from 54 of them found institutions broadly pro-use — 63% encourage GenAI use, 41% offer detailed classroom guidance, 27% discourage it — while **half the syllabi (50%) outright prohibit it**, 41% permit partial use for specified activities and 7% communicate encouragement. [[genai-policies-higher-ed-computing|Ganguly et al. (2026)]] call this the top-down versus bottom-up gap: 92% of syllabi give explicit guidelines, but instructors write local rules that may not align with institutional guidance, and only 47 institutions had both levels detectable. Citation is the one thing both levels agree on, required by 83% of syllabi. Your local rule is the one students meet first, and being stricter than your institution is normal — but it is a commitment you take on alone. [[chirikov-regulate-ai-syllabi-2026|Chirikov (2026)]] tracked over 31,000 syllabi at a large public research university in Texas from 2021 to 2025: regulation rose from near zero before ChatGPT to 55% of courses by Fall 2025, while the share of fully restrictive policies fell about 5 percentage points a year. By Fall 2025 instructors most commonly restricted drafting and revising (79%) and reasoning and [[problem-solving]] (65%), and most commonly permitted editing and proofreading (83%), study support and synthesis (80%) and coding (75%); ideation was the most contested use (46% permit, 54% restrict). Academic-integrity mentions fell from 63% to 49%, while references to AI's impact on learning rose from 1% to 29% and attribution requirements from 16% to 43%. Instructors are shifting from "don't cheat" language toward "here is why this boundary protects your learning" language. [[nash-preservice-teachers-classroom-ai-policies-2026|Nash and Burriss (2026)]] analyzed 27 classroom AI policies written by preservice English language arts teachers: **26 of 27 permitted some AI use**, almost always on teacher-specified terms, times and places, and one prohibited it entirely. Ideation was the most broadly permitted use (22 of 27) and the most ambiguous; 22 disallowed or left unclear the AI composition of sentences, paragraphs or papers; 16 permitted grammar checking and proofreading while 16 prohibited producing large AI-generated text; and **22 of 27 did not address reading at all**. The recurring failure was operationalization: one policy allowed AI "to get your thinking started" and declared "this is where the line should be drawn" without saying where. That ambiguity is a defect, not a compromise — it leaves students unable to comply and you unable to apply a consistent standard. ## What to write, and where to put it Start from tasks, not from philosophy. The clearest worked method is [[mccorkle-aligned-genai-course-policy-2025|McCorkle (2025)]], who replaced a blanket prohibition students did not believe applied to them. The problem was misalignment, not defiance: students did not see themselves as behaving dishonestly. The method was to inventory every task in the semester project, ask of each "what, specifically, am I assessing?", pair each task with a plausible professional GenAI use, and decide task by task, weighing the need to assess a capability against the value of building a workforce competency. The resulting policy is deliberately uneven: brainstorming permitted but composing specific and measurable learning objectives not, image curation permitted but slide-level message design not, scripts and narration permitted with required evaluation of the AI output. **On the syllabus:** the short summary — the tasks, the permitted use for each, how disclosure works, and the consequence of a crossed boundary. **On the assignment sheet:** the call-out, because one summary is too coarse for a complex project — McCorkle's writing worked because it gave a rationale addressed to students in the second person, naming the assessment that justifies each restriction. **In class:** say the rule out loud once, at the moment it first applies, and tie it to the work in front of them — a rule a student first meets alone is one they interpret alone. Disclosure mechanics deserve their own syllabus paragraph: students who co-designed a course policy through guided inquiry prioritized training for students and instructors, **standardized disclosure procedures**, stronger institutional support and greater involvement in decisions about AI ([[guided-inquiry-genai-course-policy-2026|Hingle and Johri (2026)]]). Say how, where and in what form a student declares AI use, and say it once. [[teaching-intro-ai-course-redesign-bill-of-rights-2026|Pisan (2026)]]'s equivalent move was to make the prompt log the graded artifact: when a model does the production, the prompt and interaction log are the thing worth versioning. ### Write only rules you can enforce A rule requiring per-task monitoring, or a judgment call about when an assessment "begins", will be applied inconsistently, and inconsistency is itself an equity problem. Detector-based regimes bring a surveillance cost that lands unevenly and raise [[remote-proctoring|proctoring]] and data-retention questions. Chirikov's finding of substantial disciplinary and task variation argues for frameworks that grant instructors autonomy within their domain rather than one-size-fits-all mandates. There is an [[equity-in-ai-education|equity]] dimension here: opaque policies assume shared background knowledge about authorship, attribution and the norms of academic work, so they fall hardest on students who arrive without it. McCorkle's argument is that transparency dismantles part of that hidden curriculum and reduces the chance that a policy failure becomes a disciplinary matter. The cheapest version of all this is fewer and more specific rules — the tasks you assess, the uses permitted for each, the disclosure mechanics, and the consequence of a crossed boundary — with the reason for each boundary stated. ## How to say it so students comply Disclosure rules do not work as compliance mechanisms on their own; assuming they do is the most common design error. In a mixed-methods study of 409 undergraduates, [[qu-wang-disclose-or-not-genai-2026|Qu and Wang (2026)]] found non-disclosure was **strategic adaptation to perceived peer norms and low interpretive [[trust]] in instructors**, not moral negligence: perceived peer disclosure and comfort with instructors were the strongest predictors of students' own disclosure, while moral disengagement had weaker effects. Relational climate is a designable variable, and mandates alone are insufficient: what students believe their peers are doing, and whether they trust you to read a disclosure fairly, predicts their honesty better than the severity of your wording. [[student-rationalization-ai-writing|Kim et al. (2026)]] identified **five distinct sites** where a course's real policy lives — faculty intention, formal policy, student interpretation, student self-policy and student practice — with systemic gaps between them, plus **23 rationalizations in six classes** students used to justify AI use in academic writing, from "no human victim" to "instructor indifference." Those rationalizations were ad hoc and post hoc: students explained behavior that had already happened rather than reasoning beforehand, which is why clearer wording and harsher detection do not close the gap by themselves. One participant who wanted to obey a strict no-AI rule kept using AI and reported the prohibition "is creating conflict for me, because I'm breaking the rules" — the ban intensified moral conflict rather than preventing use. If a rule produces guilt without compliance, it is not doing work; it is eroding the relationship you need for disclosure. Watch what you model: because most assignments in Pisan's course disclosed the prompt used to generate them, one student read transparency as **permission**: "the dependency on ai from the teacher for grading and creating assignments also made it difficult to not use ai for assignments in the same way." Modeling a norm is part of the policy as students experience it, whatever the syllabus says. A model statement you can adapt: > **AI use in this course.** I want you to leave this course able to do X yourself, so I assess X directly. You may use AI for [permitted uses] on [task list]; you may not use it for [restricted uses], because those are what I am grading. If you use AI anywhere in your work, tell me in the [single named place — e.g. a disclosure line on each submission] what you used it for and paste the prompts. Declaring it will never lower your grade; not declaring it will. If you are unsure whether a use is allowed, ask me before you submit, not after. That last sentence moves the decision point to before the work. ## The four conversations that generate conflict **The student who used AI and did not say so.** Handle this as a process question, not a character question. [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi (2026)]] model assessment as a problem of imperfect information: the student knows how the work was produced and the institution sees only the artifact and partial traces. Their response-region model compares three responses — no AI use, [[ai-use-disclosure|disclosed]] use, and hidden use — and asks which each mechanism makes most attractive. Prohibition defines the formal boundary but leaves hidden use attractive when students perceive detection as weak; monitoring raises the expected cost of hidden use without increasing disclosure; students become cautious without becoming transparent. Permission with disclosure makes honest reporting viable only when the cost of honesty is low. Practically: ask what the work was meant to assess, ask the student to account for their process, and treat a late declaration as information, not a confession. **The student who cites AI for everything.** This is usually a skills gap wearing a compliance costume. Attribution requirements in syllabi rose from 16% to 43% while academic-integrity mentions fell from 63% to 49% — the field is asking for citation, not confession. Specify the format once, in one place, on the assignment sheet rather than in a syllabus footnote. State what counts as adequate attribution for your discipline, and give one worked example. Where the real problem is that the student cannot yet do the underlying work without the model, the answer is the task design, not the rule. **Group work.** Group tasks are where the five policy sites diverge fastest: five students can hold five different interpretations of one sentence and only one submission is graded. Decide and publish, for each group deliverable, whether AI use is permitted, who declares it, and what happens when one member goes outside the rule — before the project starts, not during the dispute. If you cannot state the individual contribution a member is accountable for, grade the task as group-only work rather than adjudicating authorship after the fact. **Disclosure you cannot verify.** Some declarations will be incomplete and some omissions unfalsifiable. Do not build the rule around resolution you cannot achieve: [[bassett-ai-detectors-education-2026|Bassett et al. (2026)]] argue AI detection should not be used in education at all, on three grounds that bear on policy wording directly. Detector output is a probabilistic estimate that cannot be independently verified, because real-world text origin is unknown, so validation runs on circular reasoning. Detector scores do not meet the balance-of-probabilities standard an [[academic-integrity|integrity]] investigation requires. And the human-versus-AI dichotomy is meaningless for work created *with* rather than *by* AI. Their conclusion is that detection "does not safeguard academic integrity; it undermines it." They also flag a drafting problem: rules restricting AI use "in assessment" fail to specify when an assessment begins, so AI-assisted research, planning or editing may or may not be a violation depending on who is reading. Instead, name process evidence you can observe — a submitted prompt and interaction log, an in-class exercise produced in the room, an oral explanation — and treat a declared AI use as context. Where your institution runs a detector anyway, the same authors' warnings about data storage, retention and commercial use of student work apply; the [[privacy]] objection two preservice teachers raised is the objection your students will raise. ## "But..." — the three objections you will hear **"Students will just hide it."** Some will, and the design question is which students your rule tempts. Deterrence runs through a detector's **discrimination** between hidden use and legitimate work, not its raw catch rate, so when extra sensitivity produces more new false positives than new true positives, **stronger monitoring can make concealment relatively more attractive** — honest students are penalized faster than hidden users are identified. And permission and disclosure are different levers: permission moves the boundary of acceptable use, disclosure changes visibility. Diagnose where your task pulls the student most tempted to conceal, and design for that student, not the most conscientious one. **"I do not want to ban tools I cannot detect."** You do not have to. Prohibition is only one of four design responses, alongside monitoring, permitted use with disclosure, and redesign. Redesign changes what the task rewards, and stays cosmetic if the rubric still grades mainly the final product. [[teaching-intro-ai-course-redesign-bill-of-rights-2026|Pisan (2026)]] shows the options made explicit and graduated by level: AI use is barred in the introductory programming course, because outsourcing the first loops removes the thing being taught; encouraged on projects in data structures (Copilot permitted) but barred from pen-and-paper examinations; a study and review aid in the upper-division systems course; and **required** in the AI course itself, where one boundary carried most of the weight — the model may write code, but reflections must be the student's own voice. Assessment moved onto work a model cannot quietly ghost-write, with examinations removed entirely. **"My institution's policy already covers this."** Less than you think. That top-down versus bottom-up gap runs both ways: 92% of syllabi give explicit guidelines while local rules drift from institutional guidance, and only 47 institutions had both levels detectable. Procurement, data and review cycles belong to [[institutional-ai-policy|How Do We Write and Implement an Institutional AI Policy?]], but the classroom translation is yours. Nash and Burriss conclude that institutions must equip teachers to resist as well as adopt, providing the guidance and [[educational-development|professional development]] that let an instructor decline a specific AI use without being framed as behind the times — their participants' technodeterminism contradicted their own [[pedagogy|pedagogical]] commitments. Having an institutional policy is not the same as having cover for your judgment call. ## Your first week: an action list - Inventory the graded tasks and ask of each, "what, specifically, am I assessing?" before deciding anything about AI. - Decide task by task, and check which response that decision makes most attractive to the student most tempted to conceal. - Do not make a detector the enforcement mechanism; name process evidence you can observe instead — logs, in-class work, an oral check. - Say how and where students declare AI use, in one standard place, and treat a declaration as context rather than a confession. - State each rule as a rationale, naming the assessment it serves, with reminders on complex projects. - Write only rules you can apply consistently across a marking cycle, and cut the ones you cannot. ## Where to go next Governance, data policy, procurement and review cycles belong to [[institutional-ai-policy|How Do We Write and Implement an Institutional AI Policy?]]; mandated institutional and district rules sit with [[educational-policy-ai]] and [[governance]]; rebuilding the task itself so a grade still supports a defensible inference sits with [[redesign-assessment-ai-era]]. --- ## [How Should AI Be Designed Into the Learning Experience?](https://edtechdev.github.io/aied/faqs/designing-ai-into-learning/) # How Should AI Be Designed Into the Learning Experience? **Start with the learning goal and the learning process — not the AI feature.** The knowledge base's [[pedagogy|Pedagogies and Teaching Strategies]] concept emphasizes that the same AI can function as a scaffold, [[socratic-method|Socratic]] interlocutor, feedback partner, [[simulation]], or answer generator depending on the instructional design. What matters is whether the configuration preserves the activity that produces the intended learning. ## A strong default pattern A strong default pattern is: **learner attempts → AI supports → learner evaluates or revises → learner demonstrates understanding.** More concretely: - Preserve [[productive-failure|productive failure]] where it serves learning. - Ask for an initial prediction or solution before displaying AI assistance. - Favor questions, hints, examples, counterarguments, and feedback over immediate completion. - Require verification of consequential claims. - Incorporate opportunities for explanation and [[learning-by-teaching|teach-back]]. - Gradually fade support as competence grows. - Retain some AI-free opportunities for learners to calibrate what they can do independently. The [[reducing-ai-misuse|Reducing AI Misuse]] synthesis specifically recommends think-first/AI-second/reflect sequences and deliberate evaluation checkpoints. ## Match the AI role to the level of cognitive engagement [[thermomix-genai-education-analogy-2026|Rummel, Nachtigall and Panadero's kitchen-machine analogy]] reframes the design question from *whether* learners use [[generative-ai|generative AI]] to *how* that use shapes what they become. Mapping four uses of a smart kitchen appliance onto learning cases via the [[icap-framework|ICAP]] and [[samr-model|SAMR]] frameworks gives a design ladder: fully outsourcing an assignment, without revision or [[critical-thinking|critical engagement]], is Passive substitution and risks [[cognitive-offloading|skill loss and over-reliance]]; [[prompt-engineering|refining prompts]] and cross-verifying outputs requires [[prior-knowledge|prior knowledge]] and [[self-regulated-learning|self-regulation]] (Active/Augmentation); using AI to brainstorm, outline and evaluate original work is Constructive/Modification; and AI as a genuine dialogue partner for co-construction and adaptive [[feedback]] is Interactive/Redefinition. The design implication is direct: the same tool is a bypass at one rung and a scaffold at the next, so specify the intended mode rather than granting blanket access. ## Sequence the design, don't just permit the tool [[learning-paths-patterns-learning-design-2026|Divjak, Svetec and Horvat]] analyzed the planned sequence of 29,064 teaching and learning activities across 554 courses and found a visible design grammar: Acquisition-type activities are the most common entry point and the largest single type (above 20%), learning type tracks the intended Bloom level (Acquisition falling from ~50% at level 1 to ~20% at level 6, Production rising above 20% at levels 5–6), and the strongest transition is Assessment → Discussion (0.332). Two lessons for AI design: AI belongs where the sequence intends a specific activity type rather than bolted on at the end, and because feedback clustered with [[collaborative-learning|collaboration]], [[group-work|group work]] and [[teacher-role|teacher]] presence, peer and synchronous structures create the [[feedback]] moments AI support should plug into rather than replace. [[refrain-amplify-genai-curriculum-2026|Torres-Sahli and colleagues' "refrain, then amplify" framework]] pushes this to program level: withhold a generative tool while a capacity is forming, then restore it once the student can direct it, judge what it returns, and answer for it, with a hard-to-fake checkpoint at each hinge. Devices are governed by a forming-versus-[[cognitive-offloading|offloading]] criterion — allowed where they support engaged work, excluded where they drain attention. This turns offloading decisions into a [[curriculum-design|curriculum]] and [[governance]] question that precedes, rather than follows, course-level design. ## Constructive alignment comes first [[mcinnes-salvaging-constructive-alignment-genai-2026|McInnes and colleagues' discourse analysis]] of 14 pieces of higher-education guidance warns that efficiency-framed advice — using generative AI to draft outcomes, rubrics and course outlines — produces alignment that *looks* aligned while neglecting the "constructive" half: outcomes, activities and [[assessment]] generated as discrete items rather than interdependent ones. Their remedy is re-sequencing, not prohibition: educators should understand constructive alignment well enough to direct, interrogate and reject AI output before delegating any part of it, because a surface-acceptance habit based on plausibility is the same evaluative failure instructors warn students against. Where AI is used, they argue for institutionally bounded, [[rag|retrieval-augmented]] systems configured around local policy and quality standards rather than generic internet-trained defaults. ## The broader principle The broader principle in [[finkelstein-principled-ai-education-2025|the Principled AI Education Framework]] is that technology should augment rather than displace human capabilities that education intends to develop. See also [[learning-design|Instructional Design]], [[active-learning|Active Learning]] and [[scaffolding]]. For the pedagogical defaults that determine whether a designed interaction preserves learning, see [[reduce-ai-cheating]] and [[redesign-assessment-ai-era]]; for how the same principles constrain the software itself, see [[designing-educational-ai-software]], and for their translation into a tutor's architecture, see [[developing-ai-tutor]]. --- ## [What Are Best Practices and Tips for Designing Effective Educational AI Software?](https://edtechdev.github.io/aied/faqs/designing-educational-ai-software/) # What Are Best Practices and Tips for Designing Effective Educational AI Software? **Educational AI should be designed as an instructional system, not merely a general-purpose model with an educational interface.** The recent design research in the knowledge base sharpens that claim into something more specific: the strongest systems are built as **bounded experts under human supervision**, co-designed with the teachers and learners who will use them, and grounded in verifiable content rather than model memory. A practical set of design rules: - Align the system to explicit learning goals. - Scaffold rather than complete target cognitive work. - Ground responses in instructor-approved or authoritative content when factual reliability matters. - Communicate uncertainty. - Provide a [[human-in-the-loop-ai|human escalation path]]. - Design for "kind-but-correct" responses rather than agreement with the user. - Give instructors meaningful configuration and [[teacher-role|oversight]]. - Minimize unnecessary learner data and collect only what is pedagogically necessary (see [[privacy]]). - Design [[accessibility]] from the beginning. - Test for unequal performance across learner populations (see [[equity-in-ai-education|Equity]]). - Evaluate sustained, multi-turn interaction rather than isolated demonstration prompts. ## Pedagogical safety The [[pedagogical-safety|Pedagogical Safety]] page stresses that conventional safety testing is insufficient for education. A system can avoid toxic content and still cause educational harm by over-disclosing answers, reinforcing [[misconceptions]], suppressing reflection, promoting dependence, or drifting from instructional goals. It recommends discipline-aware, multi-turn safety evaluation, human-in-the-loop quality assurance, grounding, and alignment toward guidance rather than answer provision. The [[hazra-safetutors-pedagogical-safety-2026|SafeTutors]] [[benchmark]] turns that warning into numbers. Across every model tested — from 3.8B open-weight models to GPT-5-mini — [[pedagogy|pedagogical]] harm was universal, model scale did not reliably improve safety, and failure rates escalated from 17.7% in single-turn interactions to 77.8% in multi-turn conversations, while violation patterns varied by subject. The benchmark's 11-dimension, 48-sub-risk taxonomy (cognitive, epistemic, [[metacognition|metacognitive]], [[motivation|motivational]]-affective, developmental and equity, instructional alignment, and others) is a usable design checklist. The practical lesson for anyone specifying educational software: a system can be accurate and "safe" by conventional metrics while quietly eroding learning, so multi-turn, discipline-aware evaluation is a requirement rather than a final gate. ## Accessibility and equity Accessibility should include concrete operational requirements such as keyboard operability, screen-reader compatibility, captions and transcripts, appropriate contrast, usable text alternatives, and compatibility with assistive [[ai-technologies|technologies]]; AI-generated accessibility features still require quality checking. See [[accessibility]]. Equity testing should examine the whole pipeline and disaggregate behavior across language, disability, culture, and other relevant learner characteristics rather than relying only on aggregate accuracy. See the knowledge base's [[bias-mitigation]] guidance summarized alongside [[equity-in-ai-education|Equity]]. ## Design for bounded authority, not autonomy [[reichert-human-centered-llm-chatbot-design-teachers-2026|Reichert and colleagues' participatory design study]] with six secondary teachers is a useful corrective to the assumption that educational AI should be an autonomous agent. Asked to prototype classroom [[conversational-ai|chatbots]] on paper, every teacher described a **bounded expert** — specialized capability confined to a strictly defined domain and operating under human supervision — along two dimensions. *Authority boundaries* kept teachers in ultimate control because professional responsibility for student learning and safety cannot be delegated; *expertise boundaries* reflected AI's lack of contextual knowledge of individual students, classroom dynamics, and institutional norms. The architecture they sketched had four interconnected components — content scoping, content presentation, student adaptation, and [[teacher-role|teacher]] oversight — resting on three protective layers: domain boundaries that restrict scope, content filtering that enables safe [[personalized-learning|personalization]], and teacher override for ambiguous cases. Delegation was selective: mapped onto Gagné's nine events of instruction, teachers welcomed AI for presenting content, supplying practice problems, and offering [[formative-assessment|formative]] [[feedback]], but refused it for setting objectives or conducting [[summative-assessment|summative]] [[assessment]]. Notably, they prioritized behavioral transparency — visible limits and uncertainty cues — over model explanations, and all six asked for complete conversation logging, real-time alerts, and override capability as an expression of [[teacher-role|professional responsibility]] rather than distrust. ## Design with stakeholders, not just for them Two further studies extend this. [[ko-hughes-vsd-student-centered-its-2026|Ko and Hughes]] applied value-sensitive design to an [[intelligent-tutoring|intelligent tutoring system]] with community college students and instructors — a stakeholder group historically left out of learning-platform design — and found persistent value tensions to manage rather than solve: transparency versus interpretability, privacy versus instructional insight, and [[agency|student agency]] versus system-guided [[scaffolding]]. Students preferred collaborative, humanized explanations to raw model transparency, and the resulting prototype encoded 16 value-aligned features across [[explainable-ai]], human-in-the-loop, and [[privacy]] controls. [[wang-teacher-ai-co-design-review-2026|Wang, Liu and Islam's review]] of 28 empirical studies of teacher–AI co-design adds a design vocabulary: [[generative-ai|generative AI]] is used mainly for lesson planning, prompt generation, and creative ideation, with AI acting as assistant or content generator far more often than as co-designer, and four recurrent affordances — efficiency, responsiveness, [[creativity]], and [[equity-in-ai-education|equity]] — that teachers can use to judge which tool fits which design problem. Both studies treat design as [[human-in-the-loop-ai|human-in-the-loop]] [[human-ai-collaboration|collaboration]] — [[usability-research|usability research]] rather than outreach — and both found that the stakeholders consulted surfaced requirements no accuracy benchmark would capture. ## Ground and verify, don't trust the model Grounding is an architectural decision, not a prompt. [[eduguard-safe-rag-llm-tutor|EduGuard]], a safe [[rag|retrieval-augmented]] tutor for [[cs-education|introductory programming]], pairs instructor-approved course retrieval with an architecturally separate claim verifier, explicit [[cognitive-offloading|over-reliance]] control, and a 600-query instructor-authored benchmark spanning misconceptions, debugging, code-mixed queries, and adversarial direct-answer prompts — improving on GPT-4o-mini and Llama [[socratic-method|Socratic]] tutor baselines. For designers this is the concrete shape of "ground responses in instructor-approved content": separate the components that verify from the components that converse, and test against cases that actively try to extract answers. See [[hallucination-risk|hallucination risk]]. For how these design principles translate into a built tutor — diagnosis, hint ladders, feedback, and evaluation — see [[developing-ai-tutor]]; for the pedagogical defaults that decide whether a well-built tool is used well, see [[designing-ai-into-learning]]. **Treat the student's submission as untrusted input to any AI grader.** The threat model that most design advice omits is adversarial content inside the artifact being assessed. [[humble-prompt-injection-ai-grading-red-team-2026|Humble's (2026) adversarial red-team evaluation]] tested whether students could manipulate an LLM-based grading system through prompt injection embedded in their submissions, and found the manipulation works: injections that instruct, reframe or role-play the grader shift the score without changing the work. The design consequences follow from the same separation principle as the verification architecture above — keep the grading rubric and instructions outside the student-controlled context window, strip or flag instruction-like content in submissions, never let a submission establish its own criteria, and keep a human decision on any consequential grade. An AI grader that reads its instructions from the same text it is judging has handed the rubric to the candidate. --- ## [What Are Best Practices for Developing an Effective AI Tutor?](https://edtechdev.github.io/aied/faqs/developing-ai-tutor/) An effective AI tutor should be designed as a **learning system, not an answer-generation [[conversational-ai|chatbot]]**. The strongest theme across the knowledge base is that [[pedagogy|pedagogical]] structure—diagnosis, scaffolding, feedback, learner agency, and evaluation—matters at least as much as the underlying model. The two worked examples below (a calculus tutor and a writing coach) show how the same core architecture must be shaped by what the discipline requires of the learner. ## 1. Start with explicit learning objectives and define the learner's job Before choosing a model, specify: - What learners should know or be able to do afterward. - What cognitive work they must perform themselves. - What the tutor may assist with. A tutor optimized for "finish the problem" can easily undermine a tutor optimized for "learn to solve the problem." The [[intelligent-tutoring]] concept emphasizes that effectiveness depends on pedagogical design rather than model capability alone. ## 2. Diagnose before you prescribe Maintain a learner model based on evidence such as demonstrated knowledge, [[misconceptions]], recent attempts, help-seeking behavior, and confidence where appropriate. Adapt difficulty and assistance from this evidence rather than simply reacting to the learner's latest prompt. Be cautious about allowing an [[llm]] to perform diagnosis by itself: [[benchmark|benchmarking]] found that LLM tutors could recognize clearly correct reasoning while sometimes rejecting valid alternatives or accepting incorrect reasoning. For consequential domains, a useful architecture is **structured diagnosis + flexible LLM dialogue**. See [[yasir-llm-tutoring-agents-2026|Confirming Correct, Missing the Rest]]. ## 3. Use a hint ladder rather than giving the solution immediately A useful tutoring sequence is: ask for an attempt, probe the learner's reasoning, give a small clue, give a stronger conceptual hint, demonstrate a partial step, provide a worked solution only when warranted, then ask the learner to explain or apply the idea independently. Support should **fade as competence increases**. This is central to the [[scaffolding]] concept. A key field experiment found that an unguarded GPT interface increased assisted mathematics performance but reduced subsequent unassisted exam performance, while a hint-giving tutor largely removed that learning penalty — see [[generative-ai-guardrails-harm-learning|Generative AI without guardrails can harm learning]]. A larger randomized field experiment with more than 6,000 middle-school students on a mastery-based practice platform found the same signature in finer detail: students assigned to AI support progressed more slowly and attempted fewer questions but answered more accurately and — the clearest mechanism — improved their next-attempt correctness after mistakes, needing fewer attempts to return to a correct answer. That is a **productive slowdown**, not answer-grabbing, and it is the behavior a hint ladder is supposed to produce. The same study supplies a caution about proxies: requiring three correct answers in a row sharply raised platform-defined mastery without producing detectable gains on a delayed test a week later, and the strongest delayed-test evidence appeared only where the AI sat inside the mastery workflow (coefficient 0.085) rather than as standalone access. See [[making-ai-tutoring-productive-mastery-math-2026|Making AI Tutoring Productive]]. ## 4. Make feedback specific, immediate, actionable, and connected to reasoning Avoid feedback that merely says "Correct," "Incorrect," or "Good job." Instead the tutor should identify the reasoning step involved, explain what needs reconsideration, give the learner something concrete to do next, and ask the learner to predict or explain before revealing feedback when appropriate. The knowledge base treats [[feedback]] as a complete **provision–uptake loop**: feedback only supports learning when students understand it and act on it. ## 5. Ground factual content rather than trusting the LLM's memory Use retrieval-augmented generation against trusted materials such as instructor-approved textbooks, course notes, worked examples, policies, and [[curriculum-design|curricular]] resources, and expose citations or provenance where useful. For domains with formally checkable answers, add deterministic tools such as calculators, symbolic mathematics systems, code execution, knowledge graphs, rule-based validators, and [[discipline-specific-aied|domain-specific]] solvers. RAG can reduce [[hallucination-risk|hallucination risk]], although it does not eliminate it — see [[rag|Retrieval-Augmented Generation]]. ## 6. Design for metacognition and learner agency Regularly require the learner to generate, choose, justify, evaluate, or reflect. A useful design principle is **learner first → AI second → learner again**. The long-term goal is for learners to internalize the tutor's questioning and [[problem-solving]] strategies rather than becoming dependent on the tutor. See [[agency|learner agency]] and [[scaffolding]]. ## 7. Treat pedagogical safety as different from ordinary chatbot safety Safety testing for an educational tutor should include more than toxicity and jailbreak resistance. Test for answer leakage, misconception reinforcement, excessive agreement or [[ai-sycophancy|sycophancy]], inappropriate difficulty, [[cognitive-offloading|cognitive offloading]], biased treatment, loss of learner agency, instructional drift, and overconfidence in incorrect explanations. A tutor should be **kind but correct**, including when the learner insists on a misconception, and testing should involve extended conversations because pedagogical failures can accumulate over multiple interactions. See [[pedagogical-safety]] and [[hazra-safetutors-pedagogical-safety-2026|AI Tutor Safety and Pedagogical Harms]]. ## 8. Build privacy, accessibility, and equity into the architecture Collect only learner data that is pedagogically necessary. Where persistent memory or [[student-modeling|learner modeling]] is used, make its purpose transparent, give learners appropriate control, protect sensitive information, define retention policies, and provide instructor or [[human-in-the-loop-ai|human oversight]] for consequential situations. Audit tutor behavior across language backgrounds, ability levels, cultural contexts, [[accessibility]] needs, and different levels of [[prior-knowledge|prior knowledge]] and AI experience. Do not make sophisticated [[prompt-engineering|prompting]] a prerequisite for good instruction — the tutor itself should help learners formulate productive questions. ## 9. Measure learning, not just chatbot quality Metrics such as response accuracy, conversation length, student preference, satisfaction, task completion, and [[student-engagement|engagement]] are insufficient by themselves. Instead evaluate unassisted performance, delayed retention, transfer to new problems, misconception correction, learner independence, feedback uptake, and [[differential-effects-across-learner-groups|differential effects across learner groups]]. The critical question is whether learners can perform successfully after the tutor is removed. See [[ai-ed-evaluation]] and [[ai-tutor-behavioral-evaluation|The Missing Evaluation Axis]]. Two recent studies sharpen that rule against [[self-report-measures|self-report]] and short horizons. A pilot with 38 novice programming students found a strong association between [[generative-ai|generative AI]] usage and *perceived* learning (rs=0.802, p<0.001), while the indicators of autonomous progress without instructor support scored lowest — the gap the authors warn produces an illusion of competence and epistemic debt, and exactly the gap that satisfaction metrics reward. See [[genai-cognitive-tutor-programming-2026|Generative AI as an Informal Cognitive Tutor]]. More fundamentally, most evaluations stop at the moment assistance ends. [[cognitive-washout-ai-skill-decay-2026|Cognitive washout dynamics]] names the unmeasured post-withdrawal interval and formalizes four possible outcomes — elastic rebound, partial plateau, latent scaffold, and over-recovery — with a washout curve model whose parameters include recovery time constant, recovery completeness, and a hysteresis index comparing relearning effort to original effort. Because reversibility determines severity, the framework argues that scheduled, unassisted practice should be dosed to the recovery curve rather than argued about morally. A tutor evaluation plan should therefore include a withdrawal phase, not only an immediate post-test. See [[wang-tutor-copilot-human-ai-live-tutoring-rct-2024|the randomized evidence that brief assistance depresses later unassisted performance]]. ## 10. Keep teachers or domain experts in the quality-assurance loop Before deployment, have educators test realistic learner profiles, common misconceptions, edge cases, adversarial prompts, ambiguous responses, and extended tutoring conversations. Log pedagogical failures and use them to revise system prompts, tutoring policies, knowledge sources, [[guardrails]], learner-model rules, and model selection. Human oversight remains important because a fluent tutoring response can still be pedagogically inappropriate or incorrect. [[teacher-intervention-k12-ai-based-instruction-2026|Lee's systematic review of 29 K-12 studies]] shows what that oversight actually consists of, and where it breaks. [[teacher-role|Teacher]] intervention is a repeating four-phase cycle of monitoring, judgment, intervention and orchestration; AI alerts, [[visualization|dashboards]] and automated scores "do not automatically lead to pedagogical action"; and the leading teacher strategy is *pedagogical translation* — selecting, revising, supplementing, summarizing or deleting chatbot feedback rather than passing it through unchanged. Two design warnings follow. More AI information is not better: systems that continuously emit diagnostics overloaded teachers and pulled attention away from their own observation, so prioritize what is worth acting on and make it interpretable, and offer recommendations in a form teachers can accept, modify, defer or reject. More teacher support is not better either: delaying intervention so students can work independently is itself expertise, and structural conditions — time to review data, class size, ability to physically reach the groups needing help — are part of the intervention rather than background logistics. ## A useful AI tutor architecture A strong production architecture can be represented as: learning objective → learner evidence/learner model → pedagogical policy → grounded and validated content → conversational generation → learner response → updated learner model (looping back). Safety, privacy, accessibility, teacher oversight, and evaluation should surround the entire loop. ## The most important success criterion The most important development metric is not "did the AI solve the problem?" but **"after interacting with the AI, can the learner solve a comparable problem independently?"** The evidence for this principle is strongest in structured learning domains such as mathematics and programming; generalization to more open-ended domains remains less certain, making domain-specific evaluation essential. For where a tutor sits inside a course sequence rather than standing alone, see [[designing-ai-into-learning]]; for the software-level design rules a tutor must satisfy, see [[designing-educational-ai-software]]; and for how to evaluate the resulting intervention, see [[evaluating-ai-interventions-methods]]. --- ## Example 1: Designing a Calculus AI Tutor Consider a first-semester college calculus tutor. Its goal should be to increase what students can solve and explain **independently after the tutor is removed**, not to maximize correctly completed problems. Math tutoring is especially vulnerable to over-scaffolding, premature hint use, and incorrect diagnosis of student reasoning. See [[math-education]], [[zhang-tutormoments-2026|When Help is Unhelpful]], [[correct-answer-trap-ai-tutor|Catching the Correct Answer Trap]], and [[scaffolding]]. **Learning objectives.** The tutor might maintain a concept map (functions and graphs → rates of change → limits → derivative as a limit → derivative rules → applications → antiderivatives → definite integrals → fundamental theorem). For each concept it should distinguish several kinds of mastery. For example, "derivative mastery" should not simply mean producing the correct derivative; it could include recognizing when a derivative is appropriate, interpreting it as an instantaneous rate of change, selecting the right rule, carrying out the procedure, explaining why it is appropriate, checking reasonableness, and applying it to an unfamiliar problem. This helps prevent the [[correct-answer-trap-misconceptions|correct-answer trap]], where a learner reaches the right answer through faulty reasoning. **System architecture.** A practical six-layer design: course materials + instructor policies → retrieval/RAG layer → (problem engine → learner model) and (symbolic verifier → diagnostic engine) → pedagogical policy → conversational LLM → student → updated learner model. 1. **Course-grounding layer:** retrieves from instructor-approved materials (textbook sections, lecture notes, worked examples, terminology, approved methods, notation, assignment rules) so the tutor never introduces techniques that are mathematically valid but inappropriate for the course. 2. **Mathematical verification layer:** uses a computer algebra system to verify algebraic equivalence, derivatives, integrals, equation solutions, critical points, and numerical approximations — the LLM handles explanation and dialogue while the deterministic system handles mathematical checking. 3. **Learner-model layer:** maintains per-concept estimates (e.g. limit: developing, power rule: mastered, product rule: developing, chain rule: not demonstrated) plus misconception hypotheses with evidence and confidence. The AI should treat a misconception as a **hypothesis**, not established fact, because LLMs can hallucinate evidence or infer misconceptions incorrectly — a **detect → verify → respond** process is needed. **A tutoring interaction.** For differentiating $f(x)=(x^2+1)\sin x$, a conventional chatbot might immediately reveal the answer. A learning-oriented tutor instead reasons internally: the student differentiated both components but appears to have multiplied their derivatives (a possible product-rule-as-$f'g'$ misconception), so it asks a diagnostic question first ("what rule do you use when two functions are multiplied?"), then has the student write the product rule symbolically, then sets up $u$ and $v$, and only verifies the final expression after the student reconstructs it. **A graduated help policy.** Assistance can adapt via a level ladder: independent attempt → [[metacognition|metacognitive]] question → conceptual cue → identify the relevant rule → set up part of the problem → worked intermediate step → worked solution → student explains → student solves a transfer problem independently. Seeing a worked solution does not demonstrate mastery, so after substantial help the tutor should have the learner attempt a comparable problem unaided. **Avoiding unproductive hint use.** The interface should not make unlimited hints a frictionless shortcut, since premature hint requests and superficial hint reading are associated with lower [[learning-gains|learning gains]]. Instead of `[Hint][Hint][Hint][Show Answer]`, the system might ask "what have you tried?" and "what part is blocking you?" (choosing a rule, setting up the equation, doing the algebra, understanding the concept, something else) and provide targeted assistance. **Supporting conceptual calculus.** The tutor should connect symbolic procedures to multiple representations (formula, graph, table, verbal interpretation, physical rate-of-change context) to distinguish procedural fluency from conceptual understanding. **Teacher dashboard.** The system should expose aggregated evidence rather than opaque AI judgments — e.g. "product rule — 62% demonstrated mastery; common patterns: 18% omit one term, 11% multiply derivatives" — with individual diagnoses presented as hypotheses supported by evidence. **Evaluation plan.** Measure performance while using the tutor, performance on comparable problems without it, delayed retention, transfer to unfamiliar problems, conceptual [[explainable-ai|explanation quality]], misconception correction, appropriate vs premature [[help-seeking|help seeking]], answer leakage, diagnostic false-positive/negative rates, and [[differential-effects-across-learner-groups|differential outcomes]]. The key comparison is performance **with** the tutor versus performance **without** it afterward — a student moving from 60% to 95% while assisted but staying at 60% independently has not received effective tutoring. --- ## Example 2: Designing an AI Writing Coach An AI writing coach requires a different design because writing does not have one objectively correct answer. The goal is to help the learner become better at planning, drafting, evaluating, and revising their own writing. The knowledge base frames writing as a **cognitive, social, and rhetorical process**, meaning an AI writing system can support learning but can also eliminate exactly the thinking the assignment was intended to develop. See [[writing-education]], [[ai-writing-support-stage-ownership-2026|From Planning to Revision]], [[coach-not-crutch-ai-writing|Coach not Crutch]], and [[feedback]]. **Learning objectives.** The coach's learner model might track argument (thesis specificity, claim-evidence alignment, counterargument), organization (paragraph focus, logical progression, transitions), evidence (source relevance, evidence integration, interpretation), revision (global and sentence-level revision, feedback evaluation), and style (sentence clarity, grammar, authorial voice) — tracking writing **capabilities**, not just an essay score. **Ground the coach in the assignment.** Retrieve the assignment instructions, instructor rubric, course readings, citation requirements, genre conventions, instructor examples, and AI-use policy so feedback can reference the actual assignment ("your instructor's rubric asks you to connect every major claim to evidence from at least two course readings") rather than inventing generic expectations. **Treat writing stages differently.** AI involvement at different stages affects perceived ownership differently — planning support reduces ownership less than drafting support, and AI-generated drafting produces the largest ownership decrease. So a coach can give different permissions per stage: at planning it can ask questions, compare positions, challenge assumptions, and critique outlines but avoid generating the whole argument; at drafting the learner produces prose first (the coach helps develop, not take over); at revision the coach can identify unclear claims, point out missing evidence, check whether evidence supports a claim, detect organizational problems, and compare a draft against the rubric — **diagnosing before rewriting**; at editing (after revision) it can support grammar, punctuation, concision, and citation formatting. **Example interaction.** For an essay on requiring [[online-teaching-and-learning|online courses]], a generic system might rewrite the student's paragraph into polished prose, doing the intellectual work. A writing coach instead says what is working, names the main issue (the paragraph gives reasons but does not explain why they justify a university-wide mandate), poses a revision question, and asks the student to complete a sentence in their own words — leaving the argument construction to the learner. **Feedback should be prioritized.** Each feedback round might contain one strength to preserve, one high-impact issue, one question requiring writer judgment, and one concrete revision goal — rather than overwhelming the learner with dozens of comments. **Make the student evaluate [[ai-feedback-quality|AI feedback]].** [[feedback-literacy|Feedback literacy]] is itself a learning objective; the coach should periodically ask whether the learner agrees with a suggestion and why, and allow the learner to reject AI feedback — developing **[[evaluative-judgment|evaluative judgment]]**, not obedience. **Preserve authorial voice.** The coach should distinguish errors, clarity issues, rhetorical choices, and style preferences, and should not automatically "correct" the latter two — otherwise it risks homogenizing writing toward whatever style the model prefers, especially for [[multilingual-learning|multilingual]] writers and non-standard rhetorical styles. **A revision-history learner model.** Rather than storing only final essays, the system can learn from the student's revisions (e.g. a repeated "evidence introduced but not interpreted" pattern that improves across essays), adapting based on evidence of learning. **Teacher involvement.** The instructor controls the rubric, assignment objectives, allowed forms of AI assistance, source collection, citation expectations, whether generative drafting is permitted, and when human review is required. A teacher dashboard might show class-level patterns (e.g. 41% need support on claim-evidence connection) as a [[formative-assessment]] signal. **Evaluating the writing coach.** Measure quality of AI-assisted and later unassisted writing, ability to identify weaknesses in unfamiliar writing, revision quality, feedback uptake, ability to explain revisions, student ownership, dependence on AI prose, preservation of voice, bias across dialects/multilingual writers/groups, alignment with instructor judgment, and delayed transfer. A revealing experiment compares a group writing independently, a group where AI generates/revises text, and a coach group — all then completing a new essay without AI: if the AI-generated group performs best in practice but poorly without AI, the system improved performance rather than learning. --- ## Comparing the two designs | Design question | Calculus tutor | Writing coach | |---|---|---| | Primary learning object | Mathematical concepts and problem solving | Argumentation and writing process | | Verification | Often objectively checkable | Usually requires contextual judgment | | Deterministic tools | Symbolic math engine / calculator | Grammar, citation, rubric checks | | Main AI role | Diagnose and scaffold reasoning | Diagnose and scaffold revision | | Major risk | Giving away the solution | Writing the text for the learner | | Important learner action | Solve and explain | Draft, evaluate, and revise | | Learner model | Concepts, procedures, misconceptions | Argument, evidence, organization, revision | | Key guardrail | Attempt before solution | Student prose before AI rewriting | | Transfer test | New no-AI calculus problems | New no-AI writing task | | Success criterion | Independent mathematical reasoning | Independent writing and evaluative judgment | The two systems use many of the same AI [[ai-technologies|technologies]] but embody **different pedagogical policies because the disciplines require different kinds of thinking**. The common principle: **identify the cognitive activity that produces learning, and design the AI to support that activity without taking it away from the learner** — preserving mathematical reasoning for calculus, and authorship, rhetorical decision-making, evaluation, and revision for writing. That principle is more fundamental than any particular model, prompt, agent framework, or user interface. --- ## [Does Using AI Actually Help My Students Learn?](https://edtechdev.github.io/aied/faqs/does-ai-help-students-learn/) # Does Using AI Actually Help My Students Learn? **Yes, it can—but better work produced with AI is not necessarily evidence of better learning.** Research documents genuine learning benefits, negligible effects, and learning harms. The important question is not simply whether students use AI, but **what the AI helps them do, what thinking remains their responsibility, and what they can do afterward**. The knowledge base calls the distinction between successful AI-assisted work and acquired capability the **performance–learning gap**. A student might submit a stronger essay or solve more practice problems with AI while becoming no better—and sometimes worse—at doing comparable work independently. Conversely, well-designed AI feedback, examples, and tutoring can improve subsequent performance without AI. See [[learning-gains|Learning Gains]] and [[cognitive-offloading|Cognitive Offloading]]. ## What does the research actually show? ### Some AI-supported interventions improve learning A broad starting point is [[burneo-can-edtech-close-learning-gaps-2026|Can EdTech Close Learning Gaps? Global Evidence from Digital Interventions]]. This World Bank synthesis of 14 randomized studies across ten economies estimated a positive average learning effect of **0.125 standard deviations** for adaptive and AI-enabled educational technology. However, it combined earlier [[adaptive-learning|adaptive systems]] with [[generative-ai]] tools. It found no statistically established advantage for the newer generative tools, while acknowledging considerable uncertainty in that comparison. The result supports the potential of these interventions—not a claim that any [[conversational-ai|chatbot]] will improve learning. More specific evidence shows why [[learning-design|instructional design]] matters. In [[genai-feedback-design-multisite-experiment|Human-centered GenAI feedback design in higher education]], a multisite randomized study involving 1,176 first-year undergraduates compared [[peer-assessment|peer feedback]], direct AI feedback, self-evaluation followed by AI feedback, and a hybrid sequence combining self-evaluation, peer feedback, and AI critique. The reflective and hybrid designs produced stronger **delayed AI-free transfer** than direct AI feedback. Students subsequently performed better on a new scientific argumentation task, not merely on the assignment being revised. However, those conditions also required additional evaluative activity and potentially more time. The study supports the **complete feedback design**, rather than proving that one component or the AI itself caused the advantage. ### AI can improve practice performance while harming learning The clearest caution comes from [[generative-ai-guardrails-harm-learning|Generative AI without guardrails can harm learning: Evidence from high school mathematics]]. In this randomized field experiment with nearly 1,000 students, a general-purpose-style AI interface increased assisted practice scores by **48% relative to the control group**, but students subsequently scored **17% lower on unassisted exams**. A tutor configured with teacher-designed guidance and hints largely avoided that penalty. Crucially, it **did not produce a statistically significant improvement in unassisted exam performance over the control group**. Avoiding harm is not the same as demonstrating additional learning. These percentages describe this particular intervention and setting, not universal effects of AI use. The implication is practical: judging an AI tool by completed homework, correct practice answers, or student satisfaction can give a misleading picture of its educational value. Satisfaction and perceived learning are [[self-report-measures|self-report measures]], and the knowledge base documents how far they can drift from measured learning. A larger, longer field study points the same way at scale. [[stromberg-generative-ai-learning-penalty-secondary-2026|The Generative AI Learning Penalty]] followed 26,811 Chinese secondary students (grades 7–12) over 30 months using staggered AI adoption. Homework scores rose **18%** and completion time fell **30%** (from 64 to 45 minutes), while closed-book monthly exam scores fell **20%** within six months and high-stakes entrance-exam scores fell **18–24%** of baseline — but only after about two years. The losses were concentrated among the roughly **81%** of AI users whose behavior indicated homework outsourcing; users who kept homework time comparable to non-users learned about as efficiently. The divergence between homework and exam performance is the performance–learning gap written across a national cohort, and the two-year lag means short evaluations systematically underestimate the cost. ### Access and safeguards are not enough A two-year randomized school experiment, [[one-click-away-khanmigo-two-year-school-experiment-2026|One Click Away: AI Tutoring with Khanmigo]], found modest [[math-education|mathematics]] achievement gains across 18 middle schools. Yet students rarely engaged in substantive tutoring conversations. The authors noted that the gains resembled those associated with structured practice without AI. Because the intervention combined individualized practice and [[intelligent-tutoring|AI tutoring]], it did not cleanly isolate the AI component’s additional contribution. A capable tutor being available is different from students using it productively. The newer [[making-ai-tutoring-productive-mastery-math-2026|Making AI Tutoring Productive]] experiment offers a related lesson. Among more than 6,000 middle-school students using NUMI, AI support improved recovery after mistakes but slowed progress through questions. A three-correct-in-a-row mastery rule increased platform-defined success without, by itself, improving learning one week later. The strongest delayed-learning signal appeared when AI was embedded in the mastery workflow, but the gains were **marginally statistically significant and concentrated on particular practiced material**. This working paper provides suggestive evidence for a carefully structured approach, not a broadly proven recipe. ### How students use the tool, not just access to it, determines the outcome The same technology produces different learning depending on how the interaction is structured. In [[yan-cognitive-outsourcing-genai-assessments-2026|a qualitative study of 38 undergraduates]] in unsupervised essay assessments, engagement spanned a spectrum from [[cognitive-offloading|cognitive outsourcing]] to **cognitive reallocation** — shifting effort from low-level retrieval to [[critical-thinking|critical evaluation]]. Most students (n = 31) intended to use generative AI as a learning assistant, yet **76.32%** relied on single-turn ask–get answer–stop dialogue and **78.94%** used the tool before or after drafting rather than through the task, producing an efficiency paradox: convenience gained at the cost of the cognitive work that builds schemas ("the speed at which you forget it is also very fast"). Only 8 students worked as cognitive partners through sustained, iterative dialogue. The [[pedagogy|pedagogical]] lesson is that blanket permissions or prohibitions both leave students to guess. What changed behavior was task-specific guidance about which cognitive work students must retain and which AI assistance was appropriate — the design direction developed in [[reduce-ai-cheating]]. ### Less effort does not automatically mean less learning It would also be a mistake to conclude that AI helps only when it makes students work harder or never shows a complete example. In the preregistered experiments reported in [[coach-not-crutch-ai-writing|Coach not crutch]], adults who practiced revising cover letters with AI subsequently produced better no-AI writing than those who practiced alone, despite expending less effort. Benefits persisted at a one-day follow-up. Another experiment found that viewing an AI-revised example produced comparable benefits to practicing with the tool. These were brief, bounded writing tasks—not evidence of lasting improvement across all kinds of writing—but they show that examples can support learning rather than necessarily replace it. **The goal is therefore not maximum difficulty. It is to preserve or improve the learning activity that develops the intended capability.** ### The pattern repeats in computing education [[kumar-genai-computing-education-systematic-review-2026|A systematic review of 72 studies in computing education]] finds the same structure in its most robust result. Generative AI reliably raises short-term completion and reduces time-on-task (36 studies), and not a single study in the corpus documents a negative effect on immediate performance — yet those gains "do not transfer to independent performance" (21 studies). Codex-assisted students completed twice as many tasks during learning but performed no better than controls on post-tests without AI. [[prior-knowledge|Prior knowledge]] moderates everything: well-prepared students convert assistance into durable skill, while under-prepared students risk using it as a crutch that removes the practice they need. The review's central design requirement is **verification** — reading, testing, modifying, explaining, and critiquing AI output — made a graded, observable component of the work rather than an aspiration left to discretion. ## A useful design principle: scaffold, do not substitute The knowledge base’s [[scaffolding]] and [[active-learning|Active Learning]] syntheses emphasize support that helps learners understand, practice, evaluate, and eventually perform with less assistance. “Scaffold, do not substitute” is a useful principle, but it needs to be applied to the **learning objective**, not mechanically to every AI feature. A complete worked example can be something students learn from; a sequence of hints can still become something they click through without thinking. The important distinction is what the learner does with the assistance. See [[cognitive-offloading|Cognitive Offloading]] and [[help-seeking]]. For an activity intended to develop independent capability, a reasonable starting routine is: 1. **Establish the learner’s thinking.** Ask for an initial attempt, prediction, draft, explanation, or interpretation of an appropriate example. 2. **Provide targeted assistance.** Use AI for a hint, explanation, contrasting example, or focused feedback on the difficulty. 3. **Require a response to the assistance.** Have the student explain, revise, verify, or justify accepting or rejecting the suggestion. 4. **Check a new application with less support.** Ask the student to solve a comparable problem or apply the idea in a different context. This is an instructional starting point, not a universally validated sequence. The amount and timing of support should reflect students’ prior knowledge and the task. See [[reducing-ai-misuse|Reducing AI Misuse]] and [[prior-knowledge|Prior Knowledge]]. ### Make the retained thinking explicit For mathematics, an instructor might ask students to submit their attempted solution before requesting help, then use a prompt such as: “Identify the first step I should reconsider, give me one useful hint, and ask me to try again.” The subsequent check should require a new solution and an explanation—not reproduction of the AI’s answer. This applies the guidance in [[scaffolding]] and [[help-seeking]]. For writing, an instructor might have students assess their draft against a rubric before receiving AI critique, then explain which suggestions they accepted, modified, or rejected. If argument construction is the objective, AI-generated prose should not substitute for evidence that the student can construct an argument. If evaluating alternative revisions is the objective, comparing complete examples may be appropriate. See [[writing-education|Writing Education]] and [[feedback-literacy|Feedback Literacy]]. These are design applications of the evidence. They should still be evaluated in the course rather than assumed to work because they sound pedagogically sensible. ## How can I tell whether AI is helping in my course? The [[ai-ed-evaluation|AI Ed Evaluation]] page recommends separating the quality of the AI’s output, students’ experience using it, and students’ actual learning. A practical evaluation needs more than a satisfaction survey or comparison of assignment grades. **Specify the capability first.** Decide what students should understand or be able to do after the activity. “Produce a polished report” is different from “select appropriate evidence,” “explain a causal relationship,” or “detect an unsupported conclusion.” The assessment should reveal the capability you intend to develop. See [[educational-measurement|Educational Measurement]]. **Measure before, after, and later.** Use a brief baseline task, an immediate learning check, and a delayed application. Include both a comparable task and, where appropriate, one that changes the context or requires a different application. When the objective is independent competence, remove the AI assistance that could perform that competence for the student. A successful check immediately after practice is useful, but does not establish retention over months or transfer across domains. See [[learning-gains|Learning Gains]] and [[transfer-of-learning|Transfer of Learning]]. **Use a meaningful comparison.** Where feasible, compare the AI-supported activity with a well-designed non-AI alternative using similar content, instructional time, and practice opportunities. A simple before-and-after improvement cannot establish that AI caused the gain; students might improve through the rest of the [[teacher-role|teaching]]. For stronger causal claims, consult [[research-methods-aied|Efficacy Research Methods]] when designing the comparison. **Observe how students use the help.** Look for explanations, attempts to correct errors, justified revisions, and verification—not just logins, message counts, or completed questions. These observations can help explain a result, but should not replace a learning measure. See [[help-seeking]] and [[student-ai-interaction|Student–AI Interaction]]. **Check what a dashboard actually measures.** In [[zhang-platform-scores-miss-ai-teaching-agents-2026|an evaluation of eight AI teaching agents]] in [[medical-education|medical education]], agent rankings by the platform's own score diverged from an expert-validated teaching-quality rubric — the agent ranked third by the platform ranked last on rubric quality — because the platform score tracked student performance during the interaction, not the agent's teaching behavior. A built-in metric is a hypothesis to validate, not evidence of learning. See [[evaluating-ai-interventions-methods]] for the measures and comparison designs that make such a check credible. An important qualification is that **not every legitimate learning outcome must be demonstrated without AI**. A course may deliberately teach effective AI-supported work. In that case, assess students’ ability to select, verify, revise, and defend their use of AI, alongside whatever independent foundations the discipline requires. [[human-capability-test-learning-outcomes-ai-2026|A Human Capability Test for Learning Outcomes in the AI Era]] proposes this distinction between independent capability, AI-augmented performance, and verification responsibility. It is a conceptual assessment framework, not a validated solution for every course. ## Check who benefits—and what the benefit costs An average improvement can conceal students who receive little benefit or encounter new barriers. The [[digital-divide|Digital Divide]] synthesis distinguishes access to a tool from the skills needed to use it and the outcomes ultimately obtained. For a classroom evaluation, examine results by relevant starting points such as prior knowledge, language needs, and accessibility requirements. Provide guidance rather than assuming students already know how to obtain and evaluate useful feedback. When assessing independent learning, remove assistance that supplies the target thinking—not accommodations needed to access the task. See [[ai-literacy|AI Literacy]], [[equity-in-ai-education|Equity in AI Education]], and [[accessibility]]. Also inspect the feedback students actually receive. A fluent response can misdiagnose their difficulty, reinforce an error, or offer an answer when a different kind of support was needed. Use course-aligned materials, retain a route to human help, and avoid collecting more student information than the activity requires. These concerns connect directly to [[ai-feedback-quality|AI Feedback Quality]], [[pedagogical-safety|Pedagogical Safety]], and [[privacy]]. ## Bottom line **AI can help students learn, but neither access, polished output, nor a “tutor” label establishes that it does.** The evidence is strongest when a specified instructional design is evaluated against a meaningful alternative using measures of the capabilities students are supposed to develop. Long-term retention, broad transfer, and generalization across learners and settings remain important uncertainties. See [[limitations-in-aied-research|Limitations in AIEd Research]]. The most useful question for an instructor is: > **After this AI-supported activity, what can my students understand, explain, judge, or do that they could not do before—and what evidence shows that improvement?** --- ## [How Should AI in Education Research Incorporate Equity, Accessibility, Privacy, Ethics, and Pedagogical Safety?](https://edtechdev.github.io/aied/faqs/equity-ethics-pedagogical-safety-research/) # How Should AI in Education Research Incorporate Equity, Accessibility, Privacy, Ethics, and Pedagogical Safety? **These should be treated as design requirements and evaluation outcomes from the beginning — not as limitations added to the discussion section after an efficacy study is complete.** ## Equity For equity, examine not only who has access to AI but also who has the skills to use it effectively and who ultimately receives its benefits. The knowledge base's [[digital-divide|Digital Divide]] concept distinguishes access, skills, and outcome divides, which means equal access to a [[conversational-ai|chatbot]] is not equivalent to [[equity-in-ai-education|equitable]] educational benefit. Researchers should therefore report relevant subgroup outcomes and investigate differential effectiveness rather than relying only on overall averages (see [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]]). Newly added research shows how structural inequities accumulate. A [[adeniranye-ai-integration-nigerian-higher-education-2026|comparative study of 45 Nigerian universities]] (15 federal, 15 state, 15 private) found only moderate [[ai-education|AI]] integration (M = 4.79 on a 10-point scale, range 1.83–7.83), and [[governance]] type did not predict it (F(2,42) = 1.01, p = 0.372); institution age (β = 0.43, p = 0.016) and South-West location (β = 0.31, p = 0.029) did. Internal capabilities were mutually reinforcing (r = 0.79–0.80), and international collaborations and industry partnerships compounded one another (r = 0.74)—"connections beget connections." Policy frameworks were the weakest dimension (M = 4.09; only 27% of institutions scored 6 or higher), exposing a gap between formal strategy and operational [[curriculum-design|curriculum]] activity. Equity interventions should therefore target newer institutions and underserved regions rather than assume that institutional category determines capacity, and should account for the [[global-south|Global South]] contexts from which much of the evidence base is still missing. ## Accessibility For accessibility, include learners with disabilities in design and evaluation, test actual interfaces against accessibility requirements, provide equivalent ways of participating and demonstrating learning, and distinguish technical accessibility from genuinely inclusive pedagogy. See [[accessibility|Accessibility]]. A [[ko-hughes-vsd-student-centered-its-2026|value-sensitive design study]] with community college students, instructors, and developers shows what this looks like in practice: it produced 16 value-aligned features spanning [[explainable-ai|explainability]] (E1–E5), [[human-in-the-loop-ai|human-in-the-loop]] control (H1–H9), and [[privacy]] (P1–P4)—and surfaced value *tensions* rather than tidy solutions, including transparency versus interpretability, privacy versus instructional insight, and learner [[agency]] versus system-guided [[scaffolding]]. ## Privacy and ethics For privacy and ethics, collect only data needed for the educational purpose, make data use and system limitations transparent, maintain meaningful human accountability, and examine fairness, consent, bias, explainability, learner autonomy, and the consequences of AI-mediated decisions. These are core dimensions of [[ethics|Ethics]]. The design study above shows these are trade-offs to be managed, not independent requirements to be checked off: students wanted control over [[learning-analytics|learning analytics]] and affective data while instructors wanted visibility to support learning. As systems gain [[agency]] of their own, the governance question sharpens. [[beyond-agent-label-agentic-ai-governance-2026|A critical review of agentic AI governance]] argues that autonomy should not exceed the maturity of the evidence or the strength of accountable human control, and proposes a shared reporting language—autonomy levels A0–A4, oversight levels O0–O4, and evidence-maturity stages M0–M5—so that consequential uses (A4: admissions, grading, progression) require replicated, context-relevant validation and continuous human authority rather than a high benchmark score. ## Pedagogical safety For [[pedagogy|pedagogical]] safety, measure harms that conventional AI [[benchmark|benchmarks]] miss: [[cognitive-offloading|overreliance]], answer over-disclosure, [[misconceptions|misconception]] reinforcement, loss of agency, suppression of [[metacognition]], inequitable treatment, [[motivation|motivational]] harm, and instructional misalignment. Safety testing should include realistic multi-turn interactions and [[discipline-specific-aied|discipline-specific]] scenarios rather than only single prompts. The [[pedagogical-safety|Pedagogical Safety]] synthesis and the [[hazra-safetutors-pedagogical-safety-2026|SafeTutors]] benchmark show why technically "helpful" or accurate systems can still undermine learning. Practitioners converge on a compatible design answer: in a [[reichert-human-centered-llm-chatbot-design-teachers-2026|participatory design study with six secondary teachers]], teachers independently designed "bounded experts" rather than [[agentic-ai|autonomous agents]]—systems with a narrowly scoped domain operating under human supervision, with domain boundaries, content filtering, and [[teacher-role|teacher]] override forming three protective layers. They welcomed AI for presenting content, supplying practice, and giving [[formative-assessment|formative]] [[feedback]], but refused to delegate objective-setting or [[summative-assessment|summative assessment]], framing oversight as professional responsibility rather than distrust of the technology. ## Methodological triangulation Finally, combine [[quantitative-research|quantitative]] and [[qualitative-research|qualitative]] evidence. Disaggregated quantitative outcomes can reveal [[differential-effects-across-learner-groups|differential effects]]; interviews, observations, focus groups, and participatory or co-design methods can surface barriers, harms, cultural assumptions, and learner experiences that aggregate scores miss. The [[research-methods-aied|Research Methods in AIED]] synthesis explicitly treats methodological triangulation as important because no single method simultaneously maximizes causal inference, ecological validity, contextual understanding, and generalizability. Transparency about how AI-assisted analysis itself produced its findings is part of that obligation. [[chain-behind-claim-warrantability-2026|The warrantability proposal]] argues that an AI-assisted interpretation should remain inspectable, contestable, and revisable, supported by artifacts such as source-linked topic tables, lens stacks, and evidence rivers—so that a fluent summary cannot conceal the analytical pathway that produced it. The [[beyond-agent-label-agentic-ai-governance-2026|agentic-AI review]] adds a discipline of separating outcome levels: artifact outcomes (accuracy, [[ai-feedback-quality|feedback quality]]) can be necessary for a learner benefit but are never sufficient evidence of one, and equity and institutional outcomes are precisely where the evidence is thinnest. ## Where this fits in the knowledge base These requirements connect directly to the flagship [[top-10-findings-ai-education-instructors|Top 10 Findings for Instructors]] (findings 9 and 10), to the stakeholder refutations in [[addressing-common-misconceptions-ai-education|Addressing Common Misconceptions About AI in Education?]], and to the unmet-evidence agenda in [[research-gaps-aied|Research Gaps in AIED]]. For methods that take these constraints seriously, see [[evaluating-ai-interventions-methods|Evaluating AI Interventions: Methods]]. --- ## [What Measures and Research Methods Can an Instructor Use to Evaluate AI-Related Interventions?](https://edtechdev.github.io/aied/faqs/evaluating-ai-interventions-methods/) # What Measures and Research Methods Can an Instructor Use to Evaluate AI-Related Interventions? **Match the method to the claim.** If you want to know whether students *liked* an AI activity, a survey can help. If you want to know whether they *learned*, use performance measures. A survey is a [[self-report-measures|self-report measure]] — the right instrument for attitudes and the wrong one for learning, for the reasons collected on that page. If you want to know whether the AI *caused* an improvement, you need a credible comparison condition and preferably random assignment. ## Method options The [[research-methods-aied|Research Methods in AIED]] page distinguishes several useful options: - **Randomized experiments** provide the strongest causal inference. - **Quasi-experimental** pre/post or matched-group designs are often more practical in intact classes but support weaker causal claims. - **[[qualitative-research|Qualitative]]** interviews, focus groups, observations, and artifact analysis reveal mechanisms and unexpected experiences. - **AI-assisted qualitative analysis** can reorganize interview or observation corpora in minutes, so [[chain-behind-claim-warrantability-2026|warrantability]] matters as much as accuracy or disclosure: record the documented reorganizations that let a reader inspect, contest, and revise the pathway from data to claim. - **[[mixed-methods-research|Mixed methods]]** combine outcome evidence with explanations of why effects occurred. - **[[design-based-research|Design-based research]]** is useful when instructors are iteratively developing and refining an intervention in an authentic course. Whatever channel you use, only three conditions make a causal claim interpretable: a **precisely described treatment**, a **well-defined comparison condition**, and a **valid measure of durable learning**. [[weidlich-chatgpt-effect-search-cause-2025|Weidlich et al. (2025)]] audit 19 ChatGPT-in-education comparisons against exactly these criteria and find that only 4 (21%) satisfy all three — 74% had a well-defined treatment, 42% a well-defined control group, and 53% an outcome that qualified as learning. A [[generative-ai|general-purpose tool]] introduced alongside new activities, feedback, or interface design confounds the medium with the method, so a significant result cannot be attributed to the AI. ## A manageable classroom evaluation For a manageable classroom evaluation, a useful minimum is a **baseline measure, the intervention, an immediate post-measure, and a later unassisted measure**. Wherever possible, include a comparison condition such as existing practice, no AI, unrestricted AI versus scaffolded AI, or two alternative designs. Measure assisted performance and independent learning separately. The [[ai-ed-evaluation|AI Ed Evaluation]] synthesis recommends outcomes such as unassisted learning gain, delayed retention, transfer to a new task, quality of reasoning, [[misconceptions]], feedback uptake, and subgroup performance. Engagement, satisfaction, AI-use logs, self-efficacy, and [[technology-acceptance-model|perceived usefulness]] can be valuable secondary measures but should not be treated as substitutes for learning. Instructor workload and time savings are also legitimate implementation outcomes. Two cautions apply to how results are read. First, an averaged effect from the literature is a weak guide to a single classroom: [[oneill-presumed-effective-meta-analysis-2026|O'Neill's (2026) audit]] of 14 high-impact [[meta-analysis-systematic-review|meta-analyses]] found that none provided a valid basis for its claims — pooled [[learning-gains|"academic achievement"]] mixed test scores, motivation, [[self-efficacy]], and attitudes into one estimate, reported heterogeneity was extreme (I² ranged from 77.2% to 94.4% in the 13 analyses that reported it, with 12 of those 13 above 80%), all 14 assessed [[limitations-in-aied-research|publication bias]] invalidly, and 61% of randomly vetted primary studies carried validity problems, while the statistics that would show how far individual results actually spread were largely missing (only four analyses reported between-study variance and only two reported a prediction interval, both of which included zero). Second, a convenient score can measure the wrong construct: in [[zhang-platform-scores-miss-ai-teaching-agents-2026|an evaluation of eight AI teaching agents]], the agent ranked third by the platform's own score ranked last on an expert-validated rubric, because platform scores indexed student performance during the interaction rather than the agent's teaching quality. Treat any single dashboard metric as a hypothesis to validate against a measure tied to the capability you intend to develop. For what the current evidence does and does not show, and where it remains thin, see [[does-ai-help-students-learn]] and [[research-gaps-aied]]. --- ## [What Competencies Do Faculty Need in Regard to AI?](https://edtechdev.github.io/aied/faqs/faculty-ai-competencies/) # What Competencies Do Faculty Need in Regard to AI? **Faculty competency should extend well beyond prompt writing.** The knowledge base's [[ai-literacy|AI Literacy]] synthesis identifies four broad dimensions: - **Foundational understanding** of AI capabilities and limitations. - **Practical competence** using AI tools. - **[[critical-thinking|Critical evaluation]]** of accuracy, bias, appropriateness, and uncertainty. - **[[ethics|Ethical]] and institutional awareness** involving issues such as integrity, privacy, [[equity-in-ai-education|equity]], and [[governance]]. ## The pedagogical layer For instructors, those dimensions need an additional pedagogical layer. Faculty should be able to: - Decide what cognitive work students must retain. - Select AI uses that align with learning objectives. - Design valid AI-era assessment. - Teach students to verify and regulate AI use. - Recognize [[hallucination-risk|hallucination]], overreliance, and [[ai-sycophancy|sycophancy]]. - Know when [[human-in-the-loop-ai|human judgment]] should override automation. Faculty also need enough systems understanding to evaluate how [[generative-ai|generative AI]]-supported workflows fit together rather than viewing AI as an isolated tool. See [[teacher-ai-competency|Teacher AI Competency]] and [[educational-development|Faculty Development]]. ## Confidence is not competence Research summarized in the knowledge base also cautions against equating confidence with competence: [[self-report-measures|self-reported]] AI literacy can diverge substantially from demonstrated capability, making **performance-based professional development and authentic practice** preferable to confidence surveys alone. Teacher [[agentic-ai|multi-agent]]-workflow research further suggests that systems thinking, pedagogical beliefs, and [[self-efficacy]] interact in AI integration, so professional development should be differentiated rather than a one-size-fits-all tool workshop. ## What professional development actually moves — and what it does not Recent higher-education research puts concrete numbers on where faculty competency is strong and where it breaks down. [[sutedjo-faculty-genai-tpack-21-2026|A survey of 127 faculty at a large U.S. research university]] using the [[tpack]]-21 instrument found self-perceived [[pedagogy|pedagogical]] content knowledge high (M = 4.70) and [[tpack|content knowledge]] near-ceiling (M = 5.15), but the technology-integrated domains markedly lower — technological pedagogical knowledge (M = 2.62), technological content knowledge (M = 2.75) and holistic TPACK (M = 2.55, the lowest of seven). Content expertise did not predict GenAI-integration knowledge (CK correlations r = .11–.15, all non-significant), so faculty development **cannot assume subject-matter mastery transfers** and should build [[generative-ai|generative AI]] competence through deliberate, [[discipline-specific-aied|discipline-specific]] programming. Because the three technology-integrated domains correlated r = .81–.91, they may function as a single GenAI-integration capacity worth teaching as a shared foundation rather than as separate skills. What training can and cannot do is equally instructive. In a [[pishtari-teacher-ai-training-learning-design-2026|within-subjects study with 13 higher-education teachers]], giving teachers a [[conversational-ai|chatbot]] significantly improved the quality of their [[learning-design|learning design]] — the share attaining two or more higher-order Bloom tasks rose from **41.7% to 91.7%** (p = .01) — and sharply lowered perceived cognitive effort (median 6.33 to 3.33, p = .004); the pedagogical [[prompt-engineering|prompting]] training layered on top **did not further improve quality** and slightly raised effort. AI access, in other words, changes [[teacher-role|teaching behavior]] readily; deeper capability takes more than a workshop, and quality gains alone do not prove teachers internalized the practice. ## Standards, agency and identity Competency is also an institutional question. [[crompton-faculty-technology-integration-standards-2026|Design-based research with 114 participants across 28 U.S. institutions and 11 countries]] produced six faculty technology standards spanning the full academic role — Instructor, Coordinator, Leader, Researcher, Learner and Contributor — because existing frameworks ([[tpack]], [[samr-model|SAMR]]) and educator standards (ISTE, UNESCO, DigCompEdu) address [[k-12]] teachers or only the teaching dimension. [[dai-genai-frenemy-teaching-autonomy-2026|A survey of 287 teachers across 27 countries]] shows why adoption resists a purely technical framing: the refined model explained **79.8% of the variance** in behavioral intention, yet perceived artificial [[agency|autonomy]] influenced intention only indirectly through usefulness, and **risk aversion was a significant negative predictor** (−.163), with teachers worrying far more about students' [[cognitive-offloading|overreliance]] and [[academic-integrity|integrity]] risks than their own use. Adoption depended on context-sensitive judgment rather than the appeal of automation. Finally, [[ai-integrated-teaching-identity-tensions|interviews with two experienced academics]] suggest AI integration generates recurring **identity tensions** — pedagogy versus platform, educator versus facilitator, care versus compliance — navigated through *principled selectivity* grounded in [[pedagogy|pedagogical]] values. Competency therefore includes professional judgment about *when not to use AI*, not only skill in using it. For the student-facing counterpart, see [[incorporating-ai-literacy]] and [[ai-literacy-evidence]]; for the instructor-facing summary, see [[top-10-findings-ai-education-instructors]]. --- ## [How Do I Design Faculty Development for AI That Actually Changes Practice?](https://edtechdev.github.io/aied/faqs/faculty-development-ai/) # How Do I Design Faculty Development for AI That Actually Changes Practice? If you run faculty development, the pressure you are under is probably one of these: your provost wants an AI program this year, attendance is fine but nothing changes in classrooms, or the people who most need it will not come. All three point at the same design fault. What changes a teacher's practice is not learning what a tool can do; it is spending supported, structured time on a problem in *their own* course, long enough that a new habit replaces an old one, with the institution explicitly saying yes to the experiment instead of merely permitting it. **The short version.** Stop running tool demonstrations. Standardize a session cycle in which every participant brings a real course artifact and leaves with a revised one, wrap verification, disclosure and risk control into the working template rather than an ethics add-on, stretch the program past a single day and pair it with mentoring, run [[discipline-specific-aied|discipline-specific]] cohorts, and measure the artifacts people produced at a delay rather than how confident they felt on the last afternoon. The rest of this page is the evidence for each of those choices, the institutional conditions that decide whether they hold, and what to tell a dean who asks whether it worked. ## Why the workshop model underperforms The most direct comparison in this literature is [[faculty-development-centers-genai-training-optimization-2026|Bi et al. (2026)]], who surveyed 568 faculty across five universities and then ran a quasi-experiment with 160 of them. Both arms got four 90-minute workshops over consecutive weeks with identical contact time, the same platform and the same facilitator support. The difference was the content: one arm had conventional GenAI faculty development, the other worked a structured cycle of prompt formulation, output review, course-material revision and written reflection against a seven-component template covering objectives, task constraints, assessment criteria, verification, disclosure and risk control. The structured arm scored higher on all seven readiness dimensions, with its largest advantage in prompt design at roughly two standard deviations, and it was the only arm whose prompt-design gain survived delayed testing — the conventional arm's smaller gain had fallen back below baseline by follow-up. Two more results explain why demonstrations fail to travel. In [[ai-teaching-innovation-ai-tpack-2026|Bai and Hsieh (2026)]], who surveyed 898 university teachers on [[tpack|AI-TPACK]] scales, teachers who understood how generative tools work, or could operate them competently, were *no* more likely to report inventing and implementing new teaching practices than those who could not. The integrated AI-TPACK construct had no direct path to innovative behavior at all: its influence ran entirely through [[ai-literacy]], teaching [[self-efficacy]] and professional identity. In other words, capability is mediated by how the teacher sees themselves, which is not something a tool walkthrough addresses. There is also a ceiling. In [[pishtari-teacher-ai-training-learning-design-2026|a within-subjects study with 13 higher-education teachers]], chatbot access raised the share of activities reaching two or more higher-order Bloom tasks from 41.7% to 91.7% (p = .01) and cut perceived cognitive effort from a median of 6.33 to 3.33 (p = .004) — but the pedagogical prompting training layered on top did not improve design quality further and slightly increased effort. And in [[teachers-collaborative-evaluation-ai-content-2026|a workshop with 60 middle-school science teachers]], the productive element was collaborative review of ChatGPT-generated assessment questions, which is what surfaced conceptual-precision problems and the risk of reinforcing [[misconceptions]]. What those two studies share is that the learning came from critiquing and revising material, not from being shown material. ## What to build instead The design architecture that most closely fits the evidence above is [[designing-ai-professional-development-itpack-2026|Dogan's (2026) i-TPACK framework]], built from a PRISMA review of AI professional development plus a search that surfaced 17 further studies. It aligns five knowledge domains — intelligent technological knowledge, technological content knowledge, technological pedagogical knowledge, integrated i-TPACK and AI ethics — with four evidence-based pathways: [[active-learning]], models and examples, coaching and expert support, and [[feedback]] and reflection, governed by five principles running from Collective Synergy Over Silos to Flexibility within Structure. Its criticism of the field is worth taking seriously before you copy anyone's program: in many existing offerings, ethics is addressed superficially or absent entirely. The framework's limit is that it is conceptual, with no program-level outcome data. Translated into a program you could run next term, that means: - **Every session ends with a revised artifact.** Participants bring an objective, a rubric or an assessment task and leave with it changed. [[faculty-development-centers-genai-training-optimization-2026|Bi et al. (2026)]] measured exactly this cycle, and it is the version that held up at follow-up. - **The template carries the hard parts.** Verification, disclosure and risk control belong in the working document participants use all week, not in a closing segment. It is also how nervousness gets somewhere useful to go, which matters below. - **Treat assessment design as the highest-yield target.** In [[genai-pd-ai-pck-learning-gain-2026|Talebzadeh's (2026)]] eight-hour program across four two-hour sessions with 163 teachers and pre-service teachers, total AI-PCK rose from 151.65 to 220.14 (*d* = 2.36, *p* < .001) and the largest component effect was Assessment Rubrics (*d* = 2.19) — also the weakest area at pretest. The gap was narrower in teaching method than in assessment. - **Budget more than a day, and add mentoring.** Pre-service teachers in that same program gained significantly more than experienced ones (*p* = .033), mostly in scenario-based task design and rubric design. The experienced teachers said eight hours was too short to shift deep-seated habits and asked for long-term, mentor-based support. - **Segment your audience instead of running one universal workshop.** [[sutedjo-faculty-genai-tpack-21-2026|Sutedjo, Chowdhury and Liu (2026)]] found among 127 respondents that self-perceived pedagogical content knowledge was high (M = 4.70) and content knowledge near ceiling (M = 5.15), while technological pedagogical knowledge (M = 2.62), technological content knowledge (M = 2.75) and holistic TPACK (M = 2.55) were markedly lower. Content knowledge showed no significant correlation with any technology-integrated domain (r = .11–.15), so subject-matter mastery does not carry over, and those three low domains correlated r = .81–.91 with each other — plausibly one capacity to teach as a shared foundation rather than three separate modules. - **Decide whether relationships are in scope.** [[ai-emotional-intelligence-teacher-development-2026|Aponte et al. (2026)]] argue AI contributes to teachers' socio-emotional development only when it acts as relational infrastructure — protected time for mentoring, peer dialogue, communities of practice — and propose *relational densification* as the standard, counting trust, continuity of support and co-regulation rather than contact hours. Two of their claims translate directly: freed time is reabsorbed as extra demand unless the institution formally requires it to be reinvested, and conversational agents can quietly become a low-friction substitute for the difficult human negotiations that produce growth. Their framework is a heuristic, not a tested causal claim. ## The institutional conditions decide more than the curriculum does This is the finding to bring to a leadership meeting. In the same study, Kazakhstani faculty reported higher baseline readiness than their Chinese counterparts on all seven dimensions, with the widest gaps in disciplinary transfer, prompt design and basic understanding. Yet in hierarchical regression, the Kazakhstan coefficient — clearly significant on its own — fell to non-significance, a reduction of about four fifths, once prior GenAI use, recent AI training, institutional support, perceived permission, multilingual resource access, policy clarity and risk sensitivity were entered. Institutional support and perceived permission carried the most positive weight. Risk sensitivity was the only predictor working against readiness. More than twice as many Kazakhstani as Chinese faculty reported AI-related training in the previous six months. For a program lead the levers are therefore local and proximate: hands-on practice, explicit permission to experiment, materials in the languages your faculty actually work in, and visible backing from the unit. National policy framing does not carry readiness on its own. Where risk sensitivity bites, the answer is not to talk caution down but to give it a place to be exercised — rubric design and verification routines are how nervousness turns into judgment. Two practical frictions belong in the same conversation. Continued-use intention was the highest-rated satisfaction dimension in [[faculty-development-centers-genai-training-optimization-2026|Bi et al. (2026)]] while tool usability was the lowest, so some of what looks like reluctance is platform friction. And [[ai-emotional-intelligence-teacher-development-2026|Aponte et al. (2026)]] locate the hard limit: AI cannot compensate for chronic overload, punitive accountability cultures, weak leadership, or the absence of protected time. If those are your conditions, a better workshop will not fix them. ## The skeptics and the anxious are the point, not the obstacle [[laidlaw-genai-identity-crisis-faculty-2026|Laidlaw (2026)]] argues from an autoethnographic account that resistance is not a skills deficit. Faculty asking what the point of teaching is in a GenAI world are not asking how to use a tool safely; the question is ontological. Reading GenAI as a threshold concept — transformative, troublesome, irreversible, integrative, bounded — she concludes that anxiety, resistance and confusion are necessary parts of crossing it rather than problems to be corrected, and her recommendations invert the standard workshop: open with identity questions rather than demonstrations, run discipline-specific cohorts, allow different timelines, and treat principled non-adoption grounded in disciplinary values as legitimate. She also warns that communicating rules can slide into an "enforcement illusion" that substitutes compliance messaging for support. There is a design argument for adaptation here as well. In [[ai-teaching-innovation-ai-tpack-2026|Bai and Hsieh (2026)]], professional identity predicted innovative behavior nearly twice as strongly among infrequent AI users as among daily users — a borderline result, but it suggests identity work pays off most with the people who barely touch the tools. [[sutedjo-faculty-genai-tpack-21-2026|Sutedjo and colleagues]] argue for deliberate discipline-specific programming because expertise does not transfer; [[laidlaw-genai-identity-crisis-faculty-2026|Laidlaw]] wants discipline-specific cohorts as the setting for identity conversations; and [[ai-teaching-innovation-ai-tpack-2026|Bai and Hsieh]] suggest differentiating by experience, with limited users needing foundational guidance and frequent users getting more from interdisciplinary projects. For the anxious participant, [[ai-emotional-intelligence-teacher-development-2026|Aponte et al. (2026)]] add a governance condition rather than a motivational one: separate well-being support from managerial evaluation, minimize data, limit purpose, keep participation voluntary, and guarantee human oversight. Those are what make it safe to admit uncertainty in an AI-mediated space. **The affective layer has its own instrument now.** [[vassallo-ai-guilt-complex-faculty-2026|Vassallo (2026)]] surveyed 109 academics and identified an **"AI guilt complex"**: 35% worried AI use undermines their credibility and 26% reported feeling they are "cheating", with **anticipatory guilt about credibility exceeding remorse after use**, so the distress is socio-professional rather than private. Four profiles emerged — Comfortable Adopters (27%), Guilty Non-Users (29%), Cautious Users (28%) and Morally Distressed Avoiders (16%) — which means a single program is addressing four different problems, and the 29% who feel guilt while not using the tools need something other than a demonstration. The scale of the institutional task is visible in [[watson-rainie-ai-challenge-faculty-survey-2026|the AAC&U/Elon survey of 1,057 faculty (Watson and Rainie 2026)]]: 68% said their schools have not prepared faculty to use generative AI for teaching and mentoring, and a similar share for scholarship, while 26% of respondents do not use the tools at all — including 40% of arts and humanities faculty — and 82% name colleagues' resistance as a barrier to departmental adoption. ## Proving it worked to someone who funds it The institutional scaffold worth aligning to is [[crompton-faculty-technology-integration-standards-2026|Crompton, Burke and Nickel (2026)]], whose design-based research across two macro cycles and 114 participants — faculty from 28 U.S. institutions and 11 countries, plus centers-for-teaching directors and accreditation coordinators — produced six standards spanning the whole academic role rather than the teaching slice of it: Instructor, Coordinator, Leader, Researcher, Learner and Contributor. The gap is real: existing frameworks and educator standards target [[k-12]] teachers or address only teaching. Standards of this kind let a development center align its offer with appraisal and accreditation instead of running disconnected workshops, which is also the argument that survives a budget review. Be careful how you evaluate, because most of this literature measures self-report. [[faculty-development-centers-genai-training-optimization-2026|Bi et al. (2026)]] measured [[self-report-measures|self-reported]] readiness, not enacted teaching or student learning, and its groups were formed by voluntary sign-up and institutional scheduling rather than randomization, so its results are associations rather than causal estimates. [[ai-teaching-innovation-ai-tpack-2026|Bai and Hsieh (2026)]] is cross-sectional and self-selected, drew on eight universities in one country, and measured every construct from the same respondents at one time; their innovation scale also kept the wording of a general innovation measure. [[sutedjo-faculty-genai-tpack-21-2026|Sutedjo and colleagues]] cannot establish direction of causation, which matters because their strongest claim — that technological knowledge acts as a gateway — is exactly what a cross-sectional design cannot settle. [[genai-pd-ai-pck-learning-gain-2026|Talebzadeh (2026)]] reports very large effects without randomization, follow-up or discipline-level breakdown, so *d* = 2.36 should read as short-term rather than as proof of durable practice. [[laidlaw-genai-identity-crisis-faculty-2026|Laidlaw (2026)]] is an autoethnographic reflection, and [[ai-emotional-intelligence-teacher-development-2026|Aponte et al. (2026)]] state that their framework is conceptual. A defensible evaluation therefore measures artifacts rather than opinions: the course materials a participant revised, the rubrics written, the tasks designed. It tests at a delay, because the gap in [[faculty-development-centers-genai-training-optimization-2026|Bi et al. (2026)]] appeared only at follow-up, where the conventional group had decayed and the structured group had not. It treats satisfaction and confidence as weak signals — [[pishtari-teacher-ai-training-learning-design-2026|Pishtari and colleagues (2026)]] note that quality gains alone do not prove internalization, since rising design quality can equally reflect delegating routine work or accepting AI-generated structure. And it adds at least one relational indicator, following [[ai-emotional-intelligence-teacher-development-2026|Aponte et al. (2026)]], because improving artifacts while isolating participants is not what the program claimed. ## Three objections, answered **"We do not have the hours."** Then spend them differently rather than adding them. The comparison in [[faculty-development-centers-genai-training-optimization-2026|Bi et al. (2026)]] held contact time constant at four 90-minute sessions; the structured arm outperformed on the same budget. Six hours arranged as an artifact cycle beats six hours of demonstrations, and a shorter format spent on assessment rubrics targets the weakest area in [[genai-pd-ai-pck-learning-gain-2026|Talebzadeh (2026)]]. **"Our faculty will not come, or the skeptics will derail it."** Do not recruit the skeptics into a demonstration. Open with identity questions, run [[discipline-specific-aied|discipline-specific]] cohorts, allow different timelines, and say out loud that principled non-adoption is a legitimate position — that is the recommendation from [[laidlaw-genai-identity-crisis-faculty-2026|Laidlaw (2026)]] and it removes the argument about compliance from the room. Expect identity work to matter most with infrequent users, per [[ai-teaching-innovation-ai-tpack-2026|Bai and Hsieh (2026)]]. **"We cannot make people act on it."** You mostly cannot, but you can change what the institution signals. Perceived permission and institutional support carried the most positive weight in [[faculty-development-centers-genai-training-optimization-2026|Bi et al. (2026)]] after the national-background effect collapsed, so the useful moves are the cheap ones: state that experimentation is sanctioned, protect the time formally so it is not reabsorbed, provide materials in faculty's working languages, and remove platform friction, which was the lowest-rated dimension in that same study. ## Checklist for program designers - Anchor every session in a real course artifact — objectives, rubric or assessment task — that participants bring with them and revise before they leave. - Put verification, disclosure and risk control inside the working template rather than adding ethics as a closing unit. - Run long enough, and pair the program with mentor-based follow-up; experienced staff report that short formats cannot shift deep-seated habits. - Build [[discipline-specific-aied|discipline-specific]] cohorts and differentiate by experience level instead of running one universal workshop. - Open with identity questions before demonstrations, and treat principled non-adoption as a legitimate professional position. - Secure institutional support, perceived permission, protected time and access in faculty's working languages; these carry more weight than national policy framing. - Evaluate with artifacts, delayed measurement and at least one relational indicator; report effect sizes honestly against self-report and cross-sectional designs. For the competency target itself see [[faculty-ai-competencies]]; for structuring AI into a course rather than into a workshop, see [[designing-ai-into-learning]]; and for the wider institutional picture, see [[top-10-findings-ai-education-instructors]]. --- ## [How Should I Handle AI in Group and Collaborative Assignments?](https://edtechdev.github.io/aied/faqs/group-work-ai/) # How Should I Handle AI in Group and Collaborative Assignments? You set a group assignment, and two failure modes arrive with the submissions: an artifact that is polished and coherent but that no member can explain, and a team where one student ran the work through [[generative-ai|generative AI]] while three others coasted on the shared grade. Nothing in the file tells you which. Both are grading problems before they are integrity problems, and design problems before either: a group assignment makes two claims at once — what the team produced, and what each member learned — and AI presses on both, changing how easily work can be partitioned, how smoothly contributions fuse, and how hard it is to tell whose thinking is in the submission. The bottom line: decide AI's place **per task, not per course**; make the group's negotiation about AI use something they submit and you grade; design the work so it cannot be split into parallel chatbot sessions; and keep at least one component each member answers for alone. None of that requires policing tools — it requires designing the task. ## The short version 1. **Permit deliberately, and say what the permission covers.** Name which parts AI may touch — idea generation, language polishing, layout, scripting — and which are meant to be worked out without it. 2. **Protect the segments where the group builds shared understanding** — the framing, the disagreement, the reconciliation — because AI-mediated efficiency compresses exactly those. 3. **Require a short, jointly authored AI-use statement, and grade the reasoning.** This turns a peer norm into reviewable judgment. 4. **Make the process visible**: intermediate deliverables, shared planning documents, per-member reflections, [[peer-assessment|peer assessment]] of contribution. 5. **Keep individual accountability beside the group mark** — an oral defense, or an individual explanation of any section a member did not draft. 6. **Choose the access configuration on purpose.** A shared interface produces negotiated, visible AI use; private prompting produces opaque flows and quietly de-labeled outputs. 7. **Read a group's restraint carefully** — declining AI can be deliberate [[self-regulated-learning|self-regulation]] or simple unfamiliarity, and the two call for different responses. ## Decide what AI is for in this assignment The corpus does not settle the permission question as a permission question, and it argues against treating it as one. [[chen-zou-genai-group-assessment-agency-2026|Chen and Zou (2026)]] studied groups of three to five delivering a presentation worth 30% of the grade under an institutional "use only with explicit acknowledgment" policy, with GenAI explicitly permitted for idea generation, language polishing, visual layout and scripting. Inside that one permissive policy, students reached opposite conclusions with defensible reasoning on each side, and the design shaped which conclusion a group reached. A course-level rule does not determine what a group does; the task does. So write the permission at the level of the artifact: say which phases are AI-open and which are not, and tell students why. The scoping review's design logic points the same way — if AI-mediated communication compresses negotiation and collective sensemaking, the segments where a group builds shared understanding are the ones to protect. Be honest about the warrant, though: neither review claims that withholding AI improves learning outcomes. That case rests on the structure of the trade-off, not a measured comparison. ## Build the task so the group cannot just divide it [[wei-perkins-genai-student-collaboration-scoping-2026|Wei and Perkins (2026)]] — a PRISMA-guided scoping review of 18 English-language studies published between January 2023 and March 2025, analyzed with reflexive thematic analysis and mapped across eight themes — gives both sides of the ledger. The benefits: group knowledge development, idea generation, support for reflective thinking, communication efficiency, task coordination, and feedback. The risks: reduced peer interaction and [[student-engagement|engagement]] under over-reliance, plus privacy, transparency and accuracy concerns. Their assessment is blunt — benefits are empirically supported, while risks around privacy, transparency, bias and accuracy are largely discussed conceptually — and much of the research is short-term or conceptual, with little longitudinal work on group dynamics or cognitive development. The trade-off governs design. The review reports that group work became more efficient with reduced demand for communication, negotiation and collective sensemaking — Lin et al.'s finding, cited within the review — so the efficiency gain and the interaction loss can be the same mechanism. You cannot take the speed without the loss unless you build the negotiation back in as required work. The review concludes that [[group-work|group-based assessment]] should shift focus from product to process. Chen and Zou show how hard that is to do by accident: their three patterns of [[agency]] ran in different directions at once. Cooperation-oriented agency (five groups) intensified use to hold the work together, often feeding peers' contributions into a [[conversational-ai|chatbot]] to decode them and align their own part, with one group rebuilding its workflow as "discussion → externalisation to GenAI → collective review → re-discussion." Normative agency (seven groups) deliberately restrained use, judging the task to demand situated knowledge AI could not reach — "AI only knows that moment when you type" — or treating AI-generated sections as unfair to the groupmates who would carry them. Non-enacted agency (three groups) changed nothing, partitioning work into discrete subtasks on "individual platforms" while individual students used GenAI in ways they never contributed to the team. The rubric already required integration and coherence under a multicultural theme, each member's reflective insight tied to their own classroom contribution, and demonstrated group collaboration — and three groups still divided and conquered. A coherence criterion does not produce collaboration on its own. The access configuration is part of this. [[xu-genai-collaborative-space-2026|Xu et al. (2026)]] observed 18 students in six groups of three on a single shared ChatGPT-4 interface for roughly 45 minutes, and interviewed nine further students about asynchronous teamwork. With a shared, visible interface, teams co-constructed "collective prompts," ran a recurring surface–evaluate–embed cycle, treated the chat as shared external memory, and openly negotiated AI's role. In asynchronous work, private prompting and output "de-labeling" produced opaque, privatized information flows that raised the cost of maintaining a shared cognitive model. GenAI's role also shifted, from subordinate assistant to contested teammate, and the system is cognitively involved but contextually unaware of the team's shifting focus. ## Make each member's contribution visible [[genai-group-writing-strategies-2026|Korchak, Costley and Fanguy (2026)]] add the detail that should worry an assessor. In a scientific writing course, from interviews with 10 postgraduate students aged 23 to 39 (mean 27.7), groups used GenAI through pre-planned strategies or through open, individually driven interactions coordinated in shared documents; a distinctive use was GenAI as an editorial integrator, merging parallel sections into coherent text. Their most pointed observation: students' accounts sometimes contradicted their own written reflections — one student denied having a group strategy while describing one in writing. Group strategies are not always equally visible to every member, so they are not reliably visible to you either. Capable individual AI use that never becomes team practice is the same phenomenon from the other direction. Both studies imply one countermeasure: engineer visibility into the deliverables, with a decision log, shared documents you can open, and a short note from each member on what they contributed and which parts they still cannot explain. ## Assess the group without losing the individual Chen and Zou's central design argument is that the negotiation of acceptable GenAI use should itself become an explicit, assessable learning outcome: teams should collectively justify and document how GenAI will and will not be used, rather than leaving the norm to peer pressure or perceived risk. That is a gradeable artifact that exists only if the group actually talked, and it converts an unenforceable rule into reviewable student judgment. Pair it with an individual component — an oral defense of their own understanding, or a written explanation of any section they did not draft. That is also the honest answer to the free-rider problem: it grades comprehension rather than authorship, and comprehension is what a group mark was always a proxy for. [[llm-critical-thinking-teamwork-review|Martínez-Peláez et al. (2025)]] screened 203 studies for 2023–2024 in the Web of Science Core Collection and included 22 under a PRISMA 2020 protocol registered in PROSPERO. They report LLMs acting as catalysts for [[collaborative-learning|collaboration]] — idea generation, organization, [[peer-assessment|peer feedback]], simulated rubric-based evaluations and expert reviews, lower participation barriers — and, for [[problem-solving|problem solving]], help exploring alternative solutions, interdisciplinary perspectives and authentic scenarios. They also note a useful inversion: LLMs often produce incomplete or incorrect responses, which prompts students to question, verify and improve the information — the AI use to look for inside an individual checkpoint. Their limitations travel with any citation: a single database, journal articles only, no focus on ethics or privacy, only the first two years of the technology, and no risk-of-bias instrument applied. Treat "catalysts" as a description of observed practices, not a measured outcome. Two things not to do. Do not reach for detection to sort individual contributions: none of these six studies tests detectors in group settings, and the individual-level problems with detection are documented elsewhere in this knowledge base. And do not read a group's restraint as either endorsement or a problem without checking which it is — Chen and Zou caution that restraint can reflect unfamiliarity or risk avoidance rather than deliberate [[self-regulated-learning|self-regulation]]. ## Handle disagreement about AI inside the team Disclosure is not only a question you ask at the end; it is a norm the group has to settle, and the settling is where the interesting failures live. Chen and Zou identified two rationales for intensified use. Performance: students read the criteria and peer benchmarks and used AI to protect their part of a shared grade — a rational response to a shared mark. Perceived safety through norms: a permissive collective climate lowered the felt [[ai-misuse-learning-harm|risk of misuse]], so groups used the tool more openly and, in the cooperation-oriented case, more dependently. If you want transparency, a stated classroom norm is more effective than a warning. A group that brings you a disagreement about AI use has usually done the hard part. Give them a structure: require the AI-use statement to record what the group decided *not* to do as well as what it did, and grade the justification. This is where Chen and Zou's point about [[agency|collective agency]] lands: individual capability does not become collective capability on its own, which is why the negotiation has to be required rather than hoped for. ## The three objections you will actually hear **"Group work is already unfair — AI just makes it worse."** AI changes the shape of the problem, not its existence. The fairness complaint is already inside the group: Chen and Zou's restrained students treated generating their own section through AI as unfair to the groupmates who would carry it, and the performance rationale shows students protecting their share of a shared grade. No study here shows that AI has made free-riding worse, or that any intervention reduces it. Individual accountability and visible process are defensible on the design argument, not as proven fixes. **"AI makes collaboration better, so I should encourage it."** Partly, and the partly matters. Martínez-Peláez et al. describe genuine collaboration supports — idea generation, organization, [[peer-assessment|peer feedback]], simulated rubric evaluations, lower participation barriers — but that is a review of observed practices with the limitations above, and Wei and Perkins found the benefits empirically supported while the risk side is largely conceptual. The randomized evidence they cite is the sharpest caution: AI produced more innovative suggestions without significantly boosting participants' overall innovativeness, and could homogenize ideas and invite [[cognitive-offloading|cognitive offloading]]. Simulation work does not close the gap. [[llm-agents-collaborative-problem-solving-simulation-2026|Fang (2026)]] fine-tuned LLM agents on real participant dialogue (LoRA/QLoRA adapters on LLaMA 3.2–3B) to model 48 participants across 3,824 turns and six thematic codes, then compared real and simulated discourse with Epistemic Network Analysis; the simulated network reached a distance of 0.17 from the empirical network, below the 0.30 threshold, with a permutation p-value of 0.65, statistically indistinguishable, though the simulation slightly overemphasized Technical Constraints–Design links and under-represented Data and Performance Parameters. This is a framework paper with medium confidence: it supports one claim only — participant-specific agents can reproduce realistic collaborative discourse — and does not show that students learn more from simulated teammates. **"I cannot grade process."** You are already grading a proxy for it and calling it product quality. The move is not to grade more; it is to grade the artifact that exists only if the process happened — the negotiated AI-use statement, the decision log, the individual defense. All three are bounded and checkable. What you should not claim is that any of this is validated: the agency patterns come from one qualitative study, 52 pre-service teachers in one course, and the authors say to read them as possible responses to design, not a distribution you can expect in your class. The writing study's 10 participants were AI-expert postgraduates, and its authors call the findings context-specific rather than a taxonomy. ## What the evidence does not support GenAI does not improve unaided collaboration: the scoping review's benefits concern the process and the product, it explicitly calls for longitudinal work on group dynamics and cognitive development, and gains measured during AI-supported group work have not been shown to [[transfer-of-learning|transfer]] to collaboration without the tool. Detection does not sort group contributions. And no study here compares a group task with AI against the same task without it, so "AI-free group work produces more learning" is a design argument, not a finding you can cite. ## Do this week **In ten minutes, before the next assignment goes out:** require a jointly authored half-page stating how this team will and will not use GenAI, and why, and add an individual component to the rubric. **In one class period:** run the task by Xu et al.'s test — can the work be cut into independent pieces and reassembled? If it can, change the sequencing so at least one stage requires the whole team, and put a shared interface or document at the center of it. Tell the class which segments are meant to be AI-free, and why. **This term:** decide AI permission per task, protect the framing and reconciliation stages, and read any group's restraint as a signal to interpret rather than a position to correct. ## A practical checklist for group assignments - **Write the negotiation into the task.** Require each team to submit a short, jointly authored statement of how GenAI will and will not be used, and grade the reasoning. - **Make the process visible, not just the product.** Intermediate deliverables, [[peer-assessment|peer assessment]] of contribution, shared planning documents and reflective contributions create the interactions through which norms get negotiated. - **Protect individual accountability alongside the group mark.** Keep a component each member must answer for alone — an oral defense of their own understanding, or an individual explanation of any section they did not draft. - **Choose the access configuration deliberately.** One shared interface produces collective prompts, shared memory and negotiated roles; private prompting produces opaque flows and de-labeled outputs. - **Design the task so it cannot be partitioned.** A coherence rubric line does not produce collaboration on its own; the workflow has to require joint work. - **Say which parts are AI-free.** If negotiation and collective sensemaking are the point, name the segments where the group works without the tool, and tell students why. - **Assign roles, especially for learners who need them.** Explicit role definitions, structured assignments and small consistent teams are requirements AI collaboration tools routinely fail to accommodate. - **Read restraint carefully.** A group that declines GenAI may be exercising deliberate [[self-regulated-learning|self-regulation]] or guarding against unfamiliarity and risk; the two call for different responses. - **Keep the evidence standard honest.** Engagement, enjoyment and satisfaction are weak indicators; group GenAI use under a permissive policy with reduced peer interaction is not evidence of learning. The underlying evidence sits in the [[group-work]] and [[collaborative-learning]] concept pages. --- ## [How Is AI Impacting Students?](https://edtechdev.github.io/aied/faqs/how-ai-impacts-students/) # How Is AI Impacting Students? **AI impacts students in both positive and negative directions, and usually at the same time.** The same tool can scaffold a student's learning while inviting over-reliance, raise motivation while eroding agency, or support belonging while threatening authorship. Research on the [[student-experience|Student Experience]] points to recurring positive and negative impacts across several dimensions — and the direction depends heavily on how the AI is designed and how students use it. ## Positive impacts **1. Support for learning.** AI can [[scaffolding|scaffold]] understanding with [[feedback]], hints, and explanations, giving students on-demand help, practice, and adaptive support that keeps the learner doing the cognitively important work. Done well, this supports learning and [[help-seeking]] (see [[does-ai-help-students-learn|Does using AI actually help students learn?]]). **2. Motivation and engagement.** Personalized, immediate, and low-stakes support can raise [[motivation]] and [[student-engagement|engagement]], helping students persist and feel competent (see [[self-determination-theory|self-determination theory]] and its needs for autonomy, competence, and relatedness). **3. Access and equity in some dimensions.** AI can give students who are reluctant to ask questions a low-pressure way to get help, and can support [[well-being]] by reducing anxiety about seeking assistance. **4. Identity and future-readiness.** AI can help students build transferable skills for an AI-integrated world — [[ai-literacy]], [[distributed-cognition]], and [[metacognition]] (see [[lodge-adaptive-capabilities-genai-future-2026|adaptive capabilities]]) — and support [[learner-identity|identity formation]] as students come to see themselves as capable, AI-fluent practitioners. **5. Psychosocial support during stress.** Beyond cognitive help, students describe [[generative-ai|generative AI]] as a source of reassurance and a confidential space for reflection in high-stress training — a relational, emotional dimension of usefulness that adoption models do not capture. In a [[qualitative-research|qualitative]] study of 28 nursing students and faculty, several students described AI as a "companion" that provided comfort and non-judgmental space during [[medical-education|clinical]] stress ([[akbaba-nursing-ai-experiences-tam-2026|Akbaba & Calik Kus, nursing education]]). Read this alongside the dependence evidence below: the comfort is real, and it can shade into reliance. ## Negative impacts **1. Over-reliance and the performance–learning gap.** AI invites [[cognitive-offloading|cognitive offloading]] — students delegate the reasoning they need to practice. This can raise performance *with* AI while *lowering* later unassisted performance: the knowledge base's [[ai-misuse-learning-harm|AI misuse and learning harm]] synthesis documents a field [[rct]] where an unguarded [[intelligent-tutoring|AI tutor]] improved practice but reduced later exam performance. Overuse can also erode [[metacognition]] and [[self-regulated-learning]]. A 2026 synthesis of **72 studies** in [[cs-education|computing education]] quantifies the pattern: generative AI reliably raised short-term completion and reduced time-on-task (36 studies), yet the gains "do not transfer to independent performance" (21 studies), and benefit depended on [[prior-knowledge|prior knowledge]] — well-prepared students converted help into durable skill while under-prepared students used it as a crutch that removed [[productive-failure|productive failure]], widening competence gaps *within* a single classroom ([[kumar-genai-computing-education-systematic-review-2026|Kumar et al., systematic review]]). **2. Reduced effort and agency.** Knowing AI is available can reduce students' willingness to engage in [[productive-failure|productive failure]] ([[ai-availability-student-motivation|AI availability and motivation]]), and passive acceptance of AI output can erode [[agency]] and the sense of accomplishment that comes from doing work oneself (see the [[wang-safety-gap-productive-struggle-2026|safety gap]]). **3. Anxiety, shame, and well-being costs.** AI use is associated with [[anxiety-and-stress|anxiety and stress]] — about being replaced, uncertain assessment, or keeping up (see [[kim-ai-anxiety-comprehensive-analysis|comprehensive analysis of AI anxiety]]). [[shame-guilt-ai-regulation-computing-education|Shame and guilt]] around AI use can drive hiding and selective disclosure, harming honest engagement and [[social-emotional-learning|social-emotional]] well-being. The motivational consequences are measurable: Zhang, Shi, and Lu's survey of **1,484 undergraduates** found AI anxiety negatively correlated with academic motivation (r = −0.175) and with emotion [[regulation]] (r = −0.172), with emotion regulation partially mediating the AI-anxiety → motivation path (indirect effect = −0.076), while daily usage duration predicted almost nothing (r = 0.053). Their conclusion is that literacy instruction should pair with emotion-regulation and [[metacognition|metacognitive]] support, not simply expand tool access ([[zhang-ai-anxiety-academic-motivation-emotion-2026|Zhang et al., AI anxiety and academic motivation]]). A critical synthesis of 51 records on [[conversational-ai|conversational AI]] adds a measurement caveat: trust, reliance, over-reliance, attachment, and dependence are routinely conflated, and the cross-sectional correlates with loneliness, anxiety, fatigue, procrastination, and weaker [[critical-thinking|critical thinking]] do not establish causation — frequent use is not dependence without impaired control or functional harm ([[yan-conversational-ai-engagement-dependence-synthesis-2026|Yan, conversational AI engagement and dependence synthesis]]). **4. Threats to learner identity.** When AI produces the work, students may stop feeling the result is "theirs" — an authorship threat to [[learner-identity|learner identity]]. The [[t2i-competence-paradox-2026|competence paradox]] in creative fields shows ease-of-use undermining the craft-based identity students derive from authorship. **5. Integrity and equity risks.** AI enables new forms of [[academic-integrity|academic dishonesty]], and unequal access to (and understanding of) AI tools can widen gaps between students — a core concern of [[equity-in-ai-education|Equity]]. ## The bottom line The [[guardrails|guardrail]] is to **use AI as a scaffold, not a substitute** — keep the learner doing the cognitively important work while AI provides support — and to watch the full range of impacts (cognitive, motivational, affective, identity, social, and [[equity-in-ai-education|equity]]), not just performance. For a deeper treatment organized by dimension, see the [[student-experience|Student Experience]] concept page. For the intervention-side response — how to design AI into a course so the scaffold does not become the substitute — see [[incorporating-ai-literacy|How should I incorporate AI literacy into my course?]] and [[ai-literacy-evidence|What is the evidence on AI literacy interventions in higher education?]]. --- ## [How Should I Incorporate AI Literacy into My Course?](https://edtechdev.github.io/aied/faqs/incorporating-ai-literacy/) # How Should I Incorporate AI Literacy into My Course? The strongest approach is to treat **[[ai-literacy|AI literacy]] as part of disciplinary learning**, not as a standalone lesson on how to use ChatGPT. Students develop more useful AI literacy when they repeatedly **use, question, verify, critique, and make decisions about AI in authentic course tasks** rather than simply learning prompting techniques or attending a one-off orientation. Effective instruction combines practical competence with critical evaluation, ethical awareness, [[metacognition]], and attention to when AI should *not* be trusted or used. A useful design principle is **think first, AI second**. Ask students to form an initial interpretation, solution, argument, prediction, or draft before consulting AI. They can then compare their thinking with the AI's output, identify agreements and discrepancies, verify claims, and decide what—if anything—to adopt. This keeps AI from substituting for the cognitive work the assignment is intended to develop. This approach connects directly to concerns about [[cognitive-offloading|cognitive offloading]] and [[ai-misuse-learning-harm|AI misuse]], where immediate performance can improve even while independent learning suffers. A 2026 synthesis of **72 empirical studies** of [[generative-ai|generative AI]] in [[cs-education|computing education]] makes the point sharper and gives it numbers. [[generative-ai|Generative AI]] reliably raises short-term completion and reduces time-on-task (36 studies — the corpus's strongest, most replicated result), yet those efficiency gains "do not transfer to independent performance" (21 studies): students complete more while understanding less unless critical engagement with the output is structurally required. Benefit also depends on [[prior-knowledge|prior knowledge]] — well-prepared students convert AI help into durable skill, while under-prepared students risk using it as a crutch that removes [[productive-failure|productive failure]]. The authors recommend **graduated access** as a default: AI-free foundations first, then guided use with mandatory explanation tasks, then reflective critique. See [[kumar-genai-computing-education-systematic-review-2026|Generative AI in computing education: a systematic review and a framework for responsible integration]]. ## Require evaluation, not just generation AI literacy activities should require **evaluation, not just generation**. Students might annotate an AI response for factual errors, unsupported claims, bias, missing perspectives, inappropriate confidence, or weak disciplinary reasoning. They can compare AI output with readings, empirical evidence, professional standards, or a course [[assessment|rubric]]. One particularly useful pattern is to make the AI produce something plausible but imperfect and ask students to serve as the critic or editor. Identifying AI mistakes is an important route to stronger calibration and [[critical-thinking|higher-order thinking]]. Treat evaluation as a **teachable competence in its own right**, not a by-product of subject knowledge. In the computing-education synthesis, students performed significantly *worse* on correcting AI-generated code than on traditional programming [[summative-assessment|examination]] tasks — direct evidence that evaluating and repairing AI output does not transfer automatically from general disciplinary skill, and that without instruction students risk resubmitting incorrect work uncritically. The review consolidates this into its **VIE framework** (Verification, Implementation, Equity) and argues that critical engagement should be a **graded, observable component** of student work rather than an aspiration left to discretion — "the common mechanism linking every effective intervention in the corpus". ## Make AI literacy disciplinary and recurring Newer work also pushes back on **skill-list conceptions** of AI literacy — and shows what that looks like in practice. Burriss et al. had 22 eleventh-grade students compose video public service announcements about the AI-[[ethics]] issues they actually lived with: school surveillance, electronic "hall passes", and algorithmic [[academic-integrity|plagiarism]] accusation. Grounded in Critical Posthumanist Literacy and [[multimodal|multimodal composition]], the unit asked youth to translate abstract principles into emotionally resonant films; students overwhelmingly portrayed harm as emerging from tangled human–machine responsibility rather than a villainous tool, flipped the cheating narrative to adults' overreliance on faulty AI, and still closed on [[agency]] ("the solution is in our reach"). The authors argue that [[storytelling-in-education|storytelling]] and collaborative, creative production belong at the center of [[critical-pedagogy|critical AI literacy pedagogy]] — and that requiring students to compose *about* AI, for a real audience, develops critical competence that conventional competency scales, which assume individually measurable performance, cannot see. See [[burriss-multimodal-composition-critical-ai-literacy-2026|"Young Scholar[s] on the Beat": multimodal composition as a form of critical AI literacy pedagogy]]. Where possible, make these activities **[[discipline-specific-aied|discipline-specific]] and recurring**. Generic rules such as "check AI for [[hallucination-risk|hallucinations]]" are less useful than showing students what unreliable AI output looks like in *your* field: fabricated citations in history, invalid assumptions in economics, misleading interpretations of experimental evidence in [[biology-education|biology]], poorly justified design decisions in engineering, or superficially fluent but methodologically weak writing in psychology. The knowledge base's AI literacy synthesis argues that movement toward critical evaluation is most evident when AI-literacy experiences are sustained and embedded in authentic coursework rather than isolated workshops. See [[ai-literacy-continuum-higher-education|A Practical Five-Stage Continuum for AI Literacy in Higher Education]]. This is consistent with the [[quantitative-research|quantitative]] evidence: a [[liu-ai-literacy-interventions-meta-analysis-2026|meta-analysis of 59 AI-literacy interventions]] found the largest effects for integrated and reflective pedagogies, while cautioning that knowledge-focused interventions show larger measured effects than those targeting skills, attitudes, or ethics. ## Include collaborative learning Collaborative activities can strengthen this work. The [[icap-framework|ICAP framework]] suggests moving beyond passive exposure toward active, constructive, and interactive engagement. Students can compare prompts in pairs, debate whether an AI answer is trustworthy, jointly revise an AI-generated product, or explain to peers why they accepted or rejected particular suggestions. Importantly, the evidence supports cognitively [[active-learning|active]] and [[collaborative-learning|collaborative]] designs. See [[hingle-collaborative-ai-literacy-2025|Systematic Review of Collaborative Learning Activities for Promoting AI Literacy]]. ## Assess demonstrated judgment, not just confidence [[assessment]] should focus on **demonstrated judgment rather than confidence or [[self-report-measures|self-reported]] skill**. Students often overestimate how well they can evaluate AI. Instead of asking whether they "feel confident using AI," give them tasks requiring them to detect errors, verify sources, improve an output, explain limitations, or justify why a particular use of AI is appropriate. The knowledge base highlights a significant mismatch between self-reported and performance-based AI literacy and recommends performance-based assessment. See [[ai-literacy-assessment-misalignment|AI Literacy Assessment: Self-Reported vs Performance Misalignment]] and [[jin-glat-genai-literacy-assessment|GLAT: The Generative AI Literacy Assessment Test]]. This matters because unsupervised use defaults to the weakest form. Yan et al.'s presage–process–product study of 38 undergraduates found that **76.32%** relied on a single-turn ask–get answer–stop pattern, **78.94%** integrated generative AI only before starting or after drafting, and even mastery-oriented students defaulted to surface processes — an *efficiency paradox* in which speed is gained at the cost of the cognitive work that builds schemas ("the speed at which you forget it is also very fast"). The authors locate the cause in an "institutional vacuum": instructors prohibit copying but rarely teach productive use. Requiring process artifacts — the dialogue history, a reflection note on how AI was used, a revision log — makes the workflow visible and turns [[self-regulated-learning|self-regulation]] into something gradeable. See [[yan-cognitive-outsourcing-genai-assessments-2026|From cognitive outsourcing to reallocation: a 3P analysis of student–generative AI engagement in unsupervised assessments]]. **Scenario-based measurement gives you a judgment measure rather than a confidence measure.** [[reed-ai-literacy-ethical-judgment-scenarios-2026|Reed et al. (2026)]] put vignettes to undergraduates — using AI to improve grammar in one's own draft, generating search terms before doing the reading oneself, submitting an AI-generated reflection journal without personal input, minimally editing a generated assignment — and asked whether each was ethical. The pattern they report is that students converge on the clear cases and diverge most where the tool did part of the intellectual work, which is exactly the boundary a course policy has to state. That makes vignettes a usable classroom instrument: they surface the disagreements your policy is silently relying on, and they measure judgment about a case rather than self-reported proficiency. ## Issues and cautions **Do not equate AI literacy with [[prompt-engineering|prompt engineering]].** Students can become technically proficient with AI while remaining poor judges of its output. Some skill-oriented dimensions of AI literacy are positively associated with reported AI dependency, suggesting that operational training without self-[[regulation]], academic confidence, and [[self-regulated-learning|metacognitive support]] can inadvertently encourage greater reliance. **AI literacy has an interactional dimension — and it presupposes the very skills it is meant to build.** Brunnström and Palmqvist's eight-round demonstration of a "naive student" using a [[conversational-ai|chatbot]] on a take-home examination found the default output "polished but pedagogically thin", and that reaching a usable learning loop required eight rounds of learner meta-interventions ("simplify", "this is overwhelming, can you condense it?"). They name the capacity **AI-interaction literacy** — steering, evaluating, and learning from iterative interaction with generative AI — and draw a sharp equity conclusion: because unguided use imposes a competence that is unevenly distributed, generative AI "may be most beneficial to already advantaged students". Teaching the *interaction*, not just the tool, is therefore an equity measure as much as a [[pedagogy|pedagogical]] one. See [[brunnstrom-ai-interaction-literacy-srl-2026|AI-interaction literacy: reflections on how generative AI might be used to support self-regulated learning in higher education]] and [[how-ai-impacts-students|How is AI impacting students?]]. Likewise, **literacy instruction cannot solve model-level problems by itself**. Training can reduce some errors without eliminating them. In research on contextual sycophancy, AI-literacy and prompting instruction reduced direct mirroring of users' errors but did not completely prevent those errors from propagating into later AI advice. Students therefore need to understand that "better prompting" does not make an AI system inherently reliable. See [[contextual-sycophancy-ai-literacy|The Hidden Cost of Contextual Sycophancy]] and [[ai-sycophancy|AI Sycophancy]]. **[[equity-in-ai-education|Equity]] also matters.** Students enter courses with different prior access to paid tools, different experience prompting models, different language backgrounds, and different levels of digital confidence. Requiring sophisticated AI use without providing equitable access and foundational support can turn prior exposure into an academic advantage. Access, skills, and outcomes are separate dimensions of the [[digital-divide|digital divide]]. ## Examples from the AI in Education knowledge base - **Economics:** Students solve an economics problem first, obtain a GenAI response, compare it with their reasoning, critique the AI, reflect, and discuss with peers. See [[beck-genai-literacy-economics-hands-on|Fostering Generative AI Literacy in Economics]]. - **Psychology:** Students evaluate a ChatGPT-generated media release against a course rubric, identify weaknesses, and revise it. This simultaneously develops disciplinary, feedback, and AI literacy. See [[richmond-nicholls-genai-psych-feedback-ai-literacies|Using Generative AI to Promote Psychological, Feedback, and AI Literacies]]. - **Biology:** AI-literacy activities are integrated into disciplinary learning rather than taught as a separate technical topic. See [[zha-ai-literacy-biology-case-study|Integrating AI Literacy Education in a Biology Class]]. - **Database/design coursework:** Students analyze AI-generated errors and iteratively improve prompts and solutions, using failure analysis as the learning activity. See [[pedagogy-ai-mistakes|The Pedagogy of AI Mistakes]]. - **Cross-disciplinary university courses:** The **LearnAI** model uses just-in-time AI co-creation embedded across disciplines rather than a single generic AI-literacy module. See [[learnai-just-in-time-ai-cocreation-university-2026|LearnAI: Just-in-Time AI Co-Creation Across Disciplines]]. - **[[physics-education|Physics]]:** A redesigned introductory particle-physics course allowed generative AI as an explicit partner on research-shaped problems, but an unaided written examination (mean **20.6 out of 80**, only two students reaching 40) warned that assisted performance and independently retrievable knowledge are distinct achievements — so the authors argue AI literacy, including how to verify generated answers and treat AI as a teacher rather than a copier, should be taught **early**, alongside dedicated unaided practice. See [[ai-particle-physics-education-redesign-2026|AI in particle physics education: research problems and foundational skills]]. - **Program-level [[curriculum-design|curriculum design]]:** The **AI Literacy Continuum** provides a developmental model moving students from non-use or uncritical use toward informed use, critical evaluation, and improvement. See [[ai-literacy-continuum-higher-education|A Practical Five-Stage Continuum for AI Literacy in Higher Education]]. ## Practical takeaway A practical [[learning-design|course design]] is not simply **"Week 2: How to use ChatGPT."** Instead, identify several places where AI intersects with important disciplinary judgments and build repeated cycles of: **independent thinking → AI use → verification → critique → revision → reflection** Students should eventually be able to explain not merely *how* they used AI, but **why they trusted some outputs, rejected others, what they verified, what cognitive work remained their responsibility, and when using AI would undermine the purpose of the task**. That combination is much closer to the AI-literacy construct supported by the current evidence base than operational proficiency alone. For how strong that evidence base actually is — and where it thins out — see [[ai-literacy-evidence|What is the evidence on AI literacy interventions in higher education?]]. --- ## [How Do We Write and Implement an Institutional AI Policy?](https://edtechdev.github.io/aied/faqs/institutional-ai-policy/) # How Do We Write and Implement an Institutional AI Policy? You are the provost, dean, governance lead, teaching center director or committee chair who has to produce or defend this policy, and to answer a colleague who calls it too strict or too vague. The direct bottom line: publishing a policy is the easy part, and a document alone changes nothing. Most institutions have already answered whether to have one by publishing one, and the resulting estate is broad but shallow — advisory rather than binding, clustered around academic integrity, and thin on the data, equity and capability problems that decide whether AI use goes well. What separates a defensible policy from a document that gathers dust is a named owner, a review calendar, a translation path into courses, and an assessment design that rewards the behavior you are asking for. ## The short version Write short on aspiration and explicit about obligation. Audit your draft for every "may" and "can" where you mean "must," and define plagiarism and prohibited uses in words a first-year student can apply. Name an executive owner and a review cycle before you publish. Say where FERPA, GDPR, HIPAA, IRB and research-integrity obligations live instead of folding them in — one document absorbing data protection, research ethics, clinical regulation and academic conduct will be long, unread and unenforceable. Fund staff capability and student access together, because they fail together. Ask how many courses actually operationalized the policy, and treat a low number as a support failure rather than a compliance problem. Build reflection on AI use into graded work at a level proportionate to the effort it takes, and attach an evaluation plan, because nobody can claim the policy improved anything until they measure it. ## Decide first: what your document governs [[institutional-governance-ai-universities|Manikonda and Outlaw (2026)]] crawled AI policies from 149 R1 and R2 United States universities and verified 130 university-level policies spanning 34 states. The register is advisory rather than directive: 95% of policies carried a negative clarity-strength score, so weak language ("may," "can") dominates over directives ("must," "prohibited"). University-level policies prioritize data security, risk mitigation, procurement and legal compliance with little pedagogical guidance, while the school-level policies that exist focus on teaching and [[ai-literacy]]. Only eight business schools had a school-specific policy, and those differed from their host framework in six of eight cases. The authors' remedy is a layered structure: a university-wide risk-management layer, department-level policies, and inter-departmental committees with faculty and students. That gives you a scope decision to state in writing: declare the document a risk-management instrument, and place operational detail at department level, which also answers the charge of vagueness. ## Decide second: whether the provisions bind Look at what institutions publish when nothing forces the question. [[institutional-ai-policy-health-informatics-2026|Eldredge et al. (2026)]] scanned all 48 CAHIIM-accredited health informatics and health information management master's programs in the US. Forty programs (83%) published at least one AI-related document; of the analytic sample, 21 documents (53%) were guidelines, 9 (23%) informational and only 7 (18%) formal policies. Academic integrity was the most frequent keyword (n = 139), ahead of citation (59), [[assessment]] (50) and plagiarism (38); [[ai-use-disclosure|disclosure of AI]] and contract cheating each appeared once; HIPAA appeared five times, FERPA eleven and electronic health records not at all. Equity vocabulary was near-absent — inclusion (n = 9), accessibility (n = 9), equitable access (n = 2). That is attention without commensurate binding commitment. The lesson costs nothing: count your modal verbs and write a consequence where you mean one. ## Decide third: who writes it, and who is bound Two reviews place the weakness in capacity, not intent. [[ai-uk-higher-education-policy-2026|Ashiq (2026)]] synthesized more than seventy UK sources and found that only a minority of institutions maintain official AI governance plans, that strategic planning gaps produce policy drift and performative compliance, that national guidance from the Office for Students and Jisc has been criticized for lacking specificity and enforcement power, and that staff AI-training and infrastructure investment concentrate in research-intensive institutions while teaching-led institutions face capacity constraints. [[baroudi-anticipatory-governance-ai-higher-ed-2026|Baroudi (2026)]] reviewed 19 sources and found that only 7% of institutions had created senior AI leadership roles despite 49% treating AI as a strategic priority, with a theory-implementation gap driven by weak policy frameworks and limited digital infrastructure, particularly in the Global South. Writing a policy is not the bottleneck; owning, resourcing and revising it is. So decide who chairs this and who sits on it before you draft. [[crompton-governing-genai-higher-ed-delphi-2026|Crompton et al. (2026)]] convened a Delphi panel from 22 countries and locations across six continents and produced a consensus framework with eight core areas — academic integrity, ethical and responsible use, privacy and protection, equitable access, GenAI literacy, integration strategy, human oversight and accountability, and institutional support and infrastructure — plus a six-part mechanism to keep policy current: a multidisciplinary governance committee (more than half the panel), scheduled review cycles (half the panel), professional development, communication with all [[stakeholders]], evaluation of effectiveness and monitoring of external developments. Use the framework as your coverage checklist and the mechanism as your operating model. Note where the consensus is lopsided, because your committee will drift the same way: ethical and responsible use was referenced by two-thirds of panelists and privacy by over half, but equitable access by only one-third — [[equity-in-ai-education|equity]] is what a consensus process under-weights first. Half the panel insisted significant GenAI outputs be reviewed by a human, and favored process-focused and oral assessment. One participation decision is worth making deliberately. [[guided-inquiry-genai-course-policy-2026|Hingle and Johri (2026)]] had students co-design a GenAI course policy through guided inquiry; the priorities that emerged were training for students and instructors, standardized disclosure procedures, stronger institutional support rather than reliance on individual instructors, and a role in decisions about the rules governing their own learning — which argues for student membership on the committee rather than a consultation round. Baroudi found empowering and distributive leadership styles associated with higher faculty engagement and openness to change, and documented [[change-management]] machinery (Valente's contagion model, Rieber and Welliver's five-stage framework) as the support for adoption. Leadership style, on this evidence, is part of the policy. [[tan-aigem-ai-educational-management-2026|Tan et al. (2026)]] offer a six-dimension organizational framework whose propositions are explicitly left for future empirical validation. It is a checklist of what a policy should cover, not evidence of what works — useful for structure, over-claiming if cited as proof. **The gap in practice is between individual rules and institutional ones.** [[watson-rainie-ai-challenge-faculty-survey-2026|The AAC&U/Elon survey of 1,057 US faculty (Watson and Rainie 2026)]] found **87%** had created their own policies for students on generative AI use, while only **48%** said their institution had written such guidelines and **35%** said their department had. Structurally, 55% reported a task force or oversight group, 37% new AI-focused classes, 17% an AI major or minor, 16% new academic leadership offices, and only **13%** had adopted AI literacy as a general education learning outcome. The sample is non-scientific and the authors say it is not generalizable, but the asymmetry is the point: the rules students actually meet are written by the instructor in front of them, which is why translation into courses matters more than publication. [[coates-governing-academic-integrity-indicators-2025|Coates, Croucher and Calderon (2025)]] supply the governance instrument for the opposite end — an academic integrity indicator framework built for institutional governors through research reviews, multi-institutional case studies, prototyping and expert confirmation, together with reforms to governance architectures, people and technologies. ## Decide fourth: where data, privacy and consent obligations live Compliance with baseline law is not ethical data governance. [[league-ethical-governance-student-data-2026|Varadaraju and Vijayakumar (2026)]] argue that [[learning-analytics]] governance in higher education has remained compliance-first, centered on meeting FERPA and GDPR, and propose the LEAGUE framework — Lawfulness, Equity, Agency, Governance, Utility and Ethics by Design — demonstrated against an early-alert case study. The Delphi panel independently put privacy and data protection as a critical theme for over half its members, and the [[oecd-digital-education-outlook-2026|OECD Digital Education Outlook (2026)]]'s third policy pillar is an enabling environment for trustworthy GenAI covering [[privacy]], safety, bias testing and transparency. The lesson from the health informatics corpus is scope discipline: its AI documents barely mentioned HIPAA, FERPA or IRB guidance and never mentioned electronic health records, because those obligations live in broader compliance policy. Name where each obligation sits, and require students to encounter it in coursework. ## Decide fifth: how the policy reaches the classroom This is the decision most policies fail, and it is measurable. [[genai-policies-higher-ed-computing|Ganguly et al. (2026)]] analyzed 116 institutional GenAI policies from 131 R1 universities alongside 98 computer-science course syllabi from 54 R1 institutions. Institutions mostly encouraged GenAI use (73, 63%) while 31 (27%) discouraged it; at course level, 90 syllabi (92%) gave explicit guidelines but 49 (50%) outright prohibited GenAI use, 40 (41%) permitted partial use, only 7 (7%) encouraged it. Only 47 of 131 institutions had both an institutional policy and detectable course-level translation. Governance, on this evidence, means translating policy into concrete instructor support, or instructors will improvise local rules that may not align with institutional guidance. Ask your committee how many courses detectably translate the policy, and treat the answer as a support metric, not a compliance metric. ## Decide sixth: capability, access and the review you promised Two frequently omitted provisions are staff development and equitable access. The Delphi panel placed institutional support and infrastructure as the foundational enabler; Baroudi found senior AI roles rare and emphasized AI literacy and hands-on training for faculty and staff. Ashiq found staff AI-training concentrated in research-intensive institutions, widening a [[digital-divide]]. [[adarkwah-genai-unesco-policy-2026|Adarkwah et al. (2026)]] analyzed 159 documents from 30 highly ranked universities in the ten countries best placed on the IMF AI Preparedness Index, scored against UNESCO's eight-component GenAI framework. Four universities were excluded for lacking publicly accessible policies, leaving 26, and no public policies were found for German universities or Tallinn University of Technology. Core ethical and governance principles were widely adopted while inclusion, equity, internet access, gender parity and environmental impact were frequently overlooked, and many provisions remained declarative rather than operationally assured — the same pattern the health informatics corpus found, where equity vocabulary was near-absent. The OECD's fourth pillar is equitable digital infrastructure, including offline small language models for low-connectivity settings. Tan et al. extend the competency dimension past [[ai-literacy]] to data literacy, digital ethics, strategic thinking, change leadership and lifelong learning. Budget these together and name an owner: a policy that mandates AI use without funding devices, connectivity, accommodations or training has assigned an obligation it cannot support. **Two recent policy studies give you an audit instrument and a warning about support.** [[gutowski-hurley-genai-policy-legal-education-2025|Gutowski and Hurley (2025)]] scored institutional policies on five dimensions with explicit rubrics — prohibitiveness, permissiveness, educational integration, transparency and accountability, and depth — and found most institutions in their sector taking generally prohibitive positions while reserving discretion to individual instructors, no single accepted approach, and clarity itself treated as the precondition for defensible enforcement. Their recommendation is deliberately procedural: comprehensive guidelines whatever your stance, stakeholders involved in drafting, proactive training, governance designed to be flexible, and review cycles, on the view that policy generation is not a one-time event. [[qian-governing-genai-higher-ed-policy-2026|Qian (2026)]] reaches the same structural point from the support side: what distinguishes institutions is not the rule alone but the support ecosystem published alongside it — faculty development, student guidance, assessment support and governance structures that make the policy operational. Read your draft as a support commitment with rules attached, not the reverse. ## Four objections you will hear ### "We already have an academic integrity policy." It is necessary and not sufficient, and the corpus shows why: integrity is the easiest provision to write, which is why it dominates the keyword counts while equity vocabulary, access and data obligations sit near-absent. What the evidence supports is division of labor, not merger — conduct rules stay where the academic-integrity process can enforce them, and the AI policy carries what an integrity code cannot reach: procurement and data security, disclosure procedure, accessibility and access, staff training, and the assessment conditions under which integrity is achievable. The Delphi panel's insistence on human review of significant GenAI outputs, and its preference for process-focused and oral assessment, has no home in an integrity code. ### "Faculty will ignore it." Some will, and it is rarely about will. Ashiq's evidence is capacity: governance plans are uncommon, planning gaps produce policy drift and performative compliance, and training and infrastructure investment concentrate in research-intensive institutions while teaching-led ones face constraints. Baroudi's is structure — 7% of institutions had created senior AI leadership roles against 49% treating AI as a strategic priority. Ganguly et al.'s is translation: only 47 of 131 institutions had both a policy and detectable course-level implementation. Faculty are not ignoring the policy so much as being left alone with it. The counter-measure is professional development, worked examples and distributive leadership, which Baroudi associates with higher faculty engagement and openness to change. ### "Students will not read it." They will respond to it anyway, which is the more useful fact. [[ai-adaptation-gap-higher-education-2026|Braun and Khafizov (2026)]] surveyed 1,809 students, 250 faculty and 62 administrative staff at one teacher-education university and found students reporting higher AI-use intensity and usefulness while faculty and staff reported stronger integrity concerns. In their pooled model, perceived usefulness had the strongest standardized association with [[trust]] (β = 0.402) and institutional policy clarity was positive but weaker (β = 0.223); students reported higher perceived policy clarity than faculty did. These are cross-sectional, self-reported associations, but the ordering warns against treating a well-written policy as a trust-building instrument on its own. [[ethical-ai-higher-ed-game-theory|Ogbo et al. (2026)]] model student AI use as a coordination problem in which cohort norms, not stated rules, govern behavior, and show threshold-driven transitions: reflective assessment must be rewarded above a critical level before responsible use displaces opportunistic practice, the reward must be proportionate to the effort reflection demands, and peer sensitivity determines cascade speed. Students respond to assessment structures rather than pronouncements, which supports pedagogy-led governance over surveillance. The OECD adds that general-purpose chatbots improve output quality but that advantage disappears and sometimes reverses in exams when AI access is removed, with [[metacognition|metacognitive]] engagement dropping as work is offloaded. Put the requirement in the graded assessment. ### "We cannot police it." You cannot, and the evidence says you do not have to. Ogbo et al. offer the one formal result in this corpus — policy pronouncements alone leave opportunistic practice intact when assessment incentives are misaligned — so the lever is design, not detection. Braun and Khafizov's strongest measured driver is perceived usefulness, not policy clarity, so a policy that makes responsible use useful to students travels further than one that threatens detection. Where you do enforce, keep it narrow and procedural: defined prohibited uses, disclosure rules students can follow, and the human-review and oral-assessment practices half the Delphi panel endorsed. Your real exposure is the translation gap, not student evasion — 92% of syllabi gave explicit guidelines while 50% prohibited GenAI outright, against 63% of institutional policies that encouraged it. The inconsistency students meet is between your policy and your own courses. Where enforcement fails, it is more often an evidence problem than a detection one, and that is where institutional exposure sits. [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] documenting real allegation files found the evidentiary base thinner than policy implies: system-recorded traces exist only in supervised assessment, process evidence (drafts, supervision meetings, presentations) exists only where those practices were already required, and natural justice requires the student to be told the allegation and given a chance to respond before any determination. Detector output cannot substitute for that: [[hadra-ai-detector-accuracy-efl-2026|Hadra et al. (2026)]] measured accuracy of 0.69 and 0.61 for two widely used commercial tools, both failing on hybrid human-AI text and with a borderline misclassification risk for EFL student writing, and [[bassett-ai-detectors-education-2026|Bassett et al.]] show no threshold resolves the tradeoff between flagging honest work and missing concealed use. A finding that rests on a score nobody can interrogate is the case that generates the appeal, the complaint and occasionally the claim. Write the evidence standard into the policy, ban standalone detector evidence, and link the procedure to the obligations on [[legal-issues-and-risks]]. ## What the evidence cannot yet tell you This corpus contains no study measuring whether an institution's AI policy changed student behavior, staff practice or learning outcomes. What exists is descriptive: Manikonda and Outlaw find policy language weak and misaligned; Ashiq documents policy drift; Baroudi names the absence of longitudinal and causal evidence as a gap. Eldredge et al. caution that analyzing public documents only creates a visibility bias, so the absence of privacy or equity language is a finding about published guidance rather than practice. [[policy-deficit-ai-sel-2026|Tran, Liu and Nguyen (2026)]] reviewed 65 peer-reviewed papers on AI and social-emotional learning and found nearly three-quarters mentioned no policy implications at all, and that the minority which did mostly lacked the actor-specific detail — who should act, on what, why, when and how — that real policymaking needs. Treat the gap as a reason to instrument your own policy, not to wait. The defensible stance is that the policy is not the deliverable; the mechanism is. Publish binding provisions where you mean them, give the document a named owner and a scheduled review, translate it into course-level support, put the compliance obligations where students will meet them, fund staff capability and access together, and attach an evaluation plan — nobody can claim the policy improved anything until they measure it. ## What to do this week - **Count your modal verbs.** Search the draft for "may" and "can," rewrite each instance where you mean "must" or "prohibited," and define plagiarism and prohibited uses explicitly. - **Name one owner and one date.** A standing governance committee, a scheduled review cycle, a named executive sponsor, and a line assigning responsibility for monitoring external developments. - **Book the layering conversation.** Convene department-level leads and a committee with faculty and student members, and decide which provisions are university-wide and which the department writes. - **Write the missing provisions.** Accommodations, accessibility, device and connectivity assumptions and language coverage belong with the rule that requires AI use; then name where FERPA, GDPR, HIPAA, IRB and research-integrity obligations live. - **Ask for the translation number.** Treat a weak answer as a support failure and pair restrictions with worked examples. - **Change one assessment and one evaluation question.** Build reflection on AI use into graded work at a level proportionate to the effort it takes, then state what you will measure, when, and who owns the finding. For the course-level counterpart of this work, see [[course-ai-policy|writing a course AI policy and communicating it to students]]. The research base sits in [[institutional-governance-ai-universities|institutional governance research]], [[crompton-governing-genai-higher-ed-delphi-2026|the global Delphi framework]] and [[institutional-ai-policy-health-informatics-2026|program-level policy audits]]; see [[administrator]] and [[change-management]] for the leadership and implementation dimensions. [[educational-policy-ai]] collects the sector-wide literature, and [[governance]] covers how authority and accountability are distributed. --- ## [How Do I Redesign Assessment So That a Grade Still Tells Me Something Defensible About What the Student Knows or Can Do?](https://edtechdev.github.io/aied/faqs/redesign-assessment-ai-era/) # How Do I Redesign Assessment So That a Grade Still Tells Me Something Defensible About What the Student Knows or Can Do? **Begin by stating explicitly what capability the grade is supposed to represent.** Are you assessing what a student can do independently, what they can accomplish appropriately with AI, their ability to evaluate and direct AI, or some combination of these? The knowledge base's [[assessment-validity|Assessment Validity]] treats this as a validity problem: a polished artifact is no longer sufficient evidence that the student possesses the capability apparently demonstrated by that artifact. This is the assessment-side half of the same argument; the enforcement-side half is in [[reduce-ai-cheating]], and the general validity framing runs through [[evaluating-ai-interventions-methods]]. ## The risk of construct substitution The article [[authentic-products-authenticated-processes-2026|From Authentic Products to Authenticated Processes]] calls the resulting risk **construct substitution**: the assessor attributes an AI-mediated product to the student and inadvertently measures the tool's capabilities rather than the student's. Its recommended response is to make the student's reasoning, judgment, verification, iteration, and responsibility visible through such approaches as staged submissions, annotated decision rationales, process records, feedback-use statements, oral defenses, and other forms of authenticated process evidence. [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi]] reframe the same problem as one of **imperfect information** rather than morality: the student knows how the work was produced and the institution sees only the artifact plus partial traces, so the design question becomes *which student response each assessment environment makes most attractive*. Their response-region model compares three choices — no AI use, [[ai-use-disclosure|disclosed]] use, and hidden use — and shows that rules, monitoring, disclosure, and redesign operate through different channels and succeed only when the most attractive response aligns with the assessment's purpose. Two results bear directly on architecture. Deterrence runs through a detector's **discrimination** between hidden use and legitimate work, not its raw catch rate, so when false positives rise faster than true positives stronger monitoring can make concealment relatively *more* attractive. And redesign moves students toward responsible use only when the rubric actually rewards process evidence; otherwise it stays cosmetic. Their practical rule is to design for the student most tempted to conceal rather than the most conscientious one. ## A strong assessment architecture In practice, a strong assessment architecture often combines an **AI-enabled authentic task** with some form of **independent verification**. Depending on the discipline, that might include a brief oral explanation, in-class application, live demonstration, annotated [[eportfolio|portfolio]], short unassisted component, or questioning about key decisions. The [[authentic-assessment|Authentic Assessment]] concept also stresses realistic intellectual work, cognitive challenge, student agency, feedback, and social or professional authenticity — not merely making conventional assignments harder for AI to complete. Oral and live components need careful attention to anxiety, disability accommodations, linguistic differences, and assessor bias. One popular no-surveillance architecture — **per-student task variation**, where each examinee gets a surface-distinct but construct-equivalent version of the same task so answers cannot usefully be shared — is capability-conditional rather than free. [[varia-construct-equivalent-assessment-variant-generation-2026|VARIA]], a [[benchmark]] of 600 generated variants across 60 condition cells, found frontier models clustering tightly on a joint integrity score (0.81–0.88) while non-frontier references collapse to 0.50–0.55, and that no single [[prompt-engineering|prompting]] strategy optimizes surface diversity and construct equivalence at once. Variation-at-scale cannot be assumed from prompting alone, so an institution should validate its own model–prompt pair before treating the guarantee as real. Redesign also rarely works assignment-by-assignment. In a [[qualitative-research|qualitative]] study of 12 academics and 17 students at a large Australian university, [[nicola-richmond-programwide-assessment-genai-2025|Nicola-Richmond et al.]] found both groups agreed assessment must change but stressed systemic friction — long lead times, large-cohort accreditation constraints, workload and cost — and concluded that redesign *takes a village*: a program-wide team combining assessment-design, [[generative-ai|generative AI]], subject-matter, industry and evidence expertise, with GenAI-literacy and assurance-of-learning points placed strategically across a qualification. The task-level method that makes this tractable is the one [[mccorkle-aligned-genai-course-policy-2025|McCorkle]] documents: inventory every step a student performs, ask of each step "what, specifically, am I assessing?", and derive the AI boundary from the answer — which is also how the expectations behind [[reduce-ai-cheating]] become explicit and enforceable. The purpose matters as much as the mechanics. [[ai-agents-joyful-assessment-third-space-2026|El Khoury and Ma]] argue that reform organized only around preventing misconduct is too defensive and propose **joyful assessment** — safe, emotionally responsive, empowering and supportive of student [[agency]] — in which integrity is a *consequence* of good design rather than its starting point. Their worked example uses an instructor-built [[agentic-ai|AI agent]] to rehearse students before judgment and to produce a rubric-aligned evidence report, but the division of labor is stated plainly: the AI organizes evidence, the instructor interprets it. --- ## [How Can I Reduce AI Cheating in My Course?](https://edtechdev.github.io/aied/faqs/reduce-ai-cheating/) # How Can I Reduce AI Cheating in My Course? **The strongest direction in the knowledge base is to rely less on detection and more on structural [[assessment|assessment design]], explicit expectations, learning verification, and [[ai-literacy]].** The [[academic-integrity|Academic Integrity]] synthesis reports substantial limitations in AI-text detection and argues that academic integrity in the [[generative-ai]] era is increasingly an assessment-design problem rather than simply a detection problem. Detection tools are unreliable and procedurally unfair. In a controlled study of 642 published English abstracts, two commercial detectors flagged guideline-compliant light AI *editing* at 38–80%, flagged unmodified 2023–25 originals at 9–15% (non-[[stem-education|STEM]] far above STEM, p<0.001), and — after AI "humanization" — caught fewer than 4% of AI-labeled rewrites, an integrity catch-22 that punishes honest assistance while enabling deliberate evasion.([[karr-ai-detection-humanization-2026]]) In a covert field study, 94% of wholly AI-generated submissions injected into live online examinations across five psychology modules went undetected, and the AI work on average outscored real students; Vanderbilt University disabled its licensed detector after failing to validate an advertised 1% false-positive rate that implied roughly 750 mislabelled students among 75,000 annual submissions.([[teichmann-detecting-undetectable-misconduct-2026]]) Detection alone is therefore a weak lever — the stronger lever is the assessment-design argument set out in full in [[redesign-assessment-ai-era]]. Below are concrete, actionable approaches, roughly ordered by strength of evidence. ## 1. Guardrailed AI tools: "hint, don't answer" Configure any [[simulating-students|AI students]] use so it [[scaffolding|scaffolds]] rather than reveals. The strongest causal finding in the knowledge base is a field [[rct]] where an **unguarded** ChatGPT-style tutor raised assisted-practice performance **+48%** but *reduced* unassisted exam scores **−17%**; a **guardrailed** tutor (hints instead of answers, plus [[teacher-role|teacher]]-authored problem information) eliminated the harm entirely.([[generative-ai-guardrails-harm-learning]]) A large study of 26,811 students found homework outsourcing raised homework scores 18% but *lowered* [[summative-assessment|closed-book exam]] scores 20% within six months — the exact harm [[guardrails]] and unassisted measures are designed to prevent.([[stromberg-generative-ai-learning-penalty-secondary-2026]]) **Concrete examples:** - Set the tutor to give incremental, [[socratic-method|Socratic]] hints rather than the next answer step. - Seed the AI with correct solutions *and* common [[misconceptions]] so it can target errors. - Require a student attempt *before* the AI reveals its output ("show your attempt first"). - Treat any tool that makes the task feel effortless as misplaced — the [[brcic-effortless-trap-productive-struggle-2026|"if letting AI in makes the task feel effortless, it's in the wrong place"]] rule. ## 2. Assessment redesign: make cheating surface (and deter) by design Because misuse harm is assessment-dependent, change what counts as achievement. The [[reducing-ai-misuse|Reducing AI Misuse]] synthesis ranks this Tier-1 because it works whether or not a student chooses the right behavior — it constrains the environment rather than depending on motivation. The [[ai-assessment-scale-reform|AI Assessment Scale (AIAS)]] is a structured framework for this: label each assignment by its AI-use level (e.g., "no AI," "AI for brainstorming only," "AI assistance with attribution," "full AI use") so expectations are explicit and enforceable.([[ai-assessment-scale-reform]]) A 72-study [[meta-analysis-systematic-review|systematic review]] of generative AI in [[cs-education|computing education]] reaches the same conclusion from a different evidence base: it rates **redesign** as both feasible and highest-leverage, singling out adding an oral or otherwise process-visible element to at least one high-stakes assessment per course as the single most effective intervention, and notes that detection is the *thinnest* area of the whole review — just 3 studies, despite dominating institutional debate ([[kumar-genai-computing-education-systematic-review-2026]]). The same review reports that 60–80% of computing students used generative AI in coursework, typically without explicit instructor sanction, and that 70% of one national faculty sample explicitly requested training on AI-resistant assessment design. **Concrete examples:** - **Unassisted, in-class, closed-book assessments** — proctored exams, quizzes, or timed written work where students perform without tools. Weight these more heavily, since homework is what AI inflates. - **Oral exams and defenses** — have students explain or defend their work aloud; real-time dialogue is inherently AI-resistant.([[fenton-oral-exams-ai-authentic-assessment-2025]]) - **Process artifacts** — require drafts, reasoning traces, annotated "show your thinking," or reflection logs so the *process* is visible, not just the product.([[authentic-products-authenticated-processes-2026]]) - **Authentic, contextual tasks** — use real-world, data-rich, or personal prompts that are hard to delegate and meaningful to the student (e.g., apply a concept to a local case, an internship, or the student's own data).([[kirsanov-beyond-detection-ai-online-assessments-2026]]) - **AI-free zones** — designate portions of the course (or specific assignments) where independent capability is genuinely the construct being assessed. - **Per-student task variation** — give each student a surface-distinct but construct-equivalent version of the same task so copying is structurally useless; treat this as capability-conditional, since [[varia-construct-equivalent-assessment-variant-generation-2026|VARIA]] found frontier generators reach only 0.81–0.88 on a joint integrity score while non-frontier models collapse to 0.50–0.55. ## 3. Learning verification: verify understanding, not provenance Rather than trying to prove *how* a submission was produced, occasionally ask students to *demonstrate* what they learned. [[best-response-student-ai-dialog-2026|"The Best Response to Student AI Use Is Not Detection, It Is Dialog"]] describes short verification conversations, early drafts, reflections, and student videos as practical mechanisms. **Concrete examples:** - A 2-minute one-on-one or recorded explanation of a submitted piece. - A follow-up quiz on the same material, taken without tools. - Ask students to revise a sample of their work and explain the changes. *Note:* this source is a practitioner account, so it is best treated as a promising practice rather than definitive causal evidence. **Verification is also what makes a misconduct process defensible.** [[munoz-misconduct-allegation-evidence-2026|Munoz et al. (2026)]] analyzed actual generative AI misconduct allegation files and found that principles of natural justice require a student to be informed of the allegation and given an opportunity to respond *before* any determination; the response opportunity is typically an investigative meeting or panel interview, and whatever the student says becomes part of the evidentiary record. Their evidence categories also explain why verification has to be built into the course rather than improvised during the investigation: system-recorded behavioral traces exist only in supervised assessment, and the weaker process evidence — drafts, supervision meetings, presentations — exists only where those practices were already in place. A verification routine is process evidence you can then rely on. The rule itself has to be precise too: [[wright-transcription-not-generation-2026|Wright (2026)]] shows that prohibitions which treat speech-to-text transcription and generative drafting as the same "AI use" are over-inclusive and risk sanctioning students who did not do the prohibited thing, which is a fairness problem before it is a legal one ([[legal-issues-and-risks]]). ## 4. Scaffolded use sequences: "think first, AI second, reflect third" Rather than banning AI, teach students a structured workflow that keeps them in the cognitive loop. The [[reducing-ai-misuse|Reducing AI Misuse]] synthesis outlines eight design principles: preserve [[desirable-difficulties|cognitive friction]], position AI as a *provisional* thinking partner (not an authority), embed evaluation checkpoints, and require [[metacognition|metacognitive]] journaling and prompt logs. **Concrete example sequence:** 1. **Think first** — students brainstorm, outline, or draft independently before any AI use. 2. **AI second** — they use AI to critique, extend, or generate alternatives against their own thinking. 3. **Reflect third** — they log what they used AI for, what they accepted/rejected, and why (a prompt + revision log). ## 5. Task-specific AI-use declarations Replace generic "I used AI ☐" checkboxes with **[[discipline-specific-aied|domain-specific]] declaration frameworks** that map AI use to cognitive stages (e.g., structural planning vs. content generation).([[genai-declaration-frameworks-higher-education]]) This forces students to reflect on *how* they used AI and clarifies the boundary between acceptable assistance and misconduct. Pair it with explicit expectations and assurance that honest disclosure will not be penalized — punitive or vague policies actively drive concealment.([[gonsalves-student-non-compliance-ai-declarations-2025]])([[chang-should-i-tell-my-teacher-ai-disclosure-2026]]) The assessment-design modeling of [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi]] explains the mechanism behind that advice. Disclosure becomes the attractive option only when the cost of honesty stays low; and because a detector's false positives fall on honest students too, **stronger monitoring can make concealment relatively more attractive** whenever extra sensitivity produces more new false positives than new true positives. Read a declared use as context rather than a confession, and design for the student most tempted to conceal rather than the average one. **Concrete example:** a coversheet that asks students to state, per assignment: *Did you use AI? For which stages (brainstorming / drafting / revising / checking)? What tool and prompts did you use? How did you evaluate the output?* ## 6. Build AI literacy and honest expectations The [[reducing-ai-misuse|Reducing AI Misuse]] synthesis ranks AI-literacy and [[prompt-engineering|prompting]] instruction as Tier-2: a [[k-12]] module using scenario-based prompt practice with an [[llm]] auto-grader improved actual prompting skills and raised confidence in using AI for learning **+10.4%**, with 87% reporting they learned to use AI responsibly.([[aaai2026-prompting-literacy-k12]]) Set clear expectations about what counts as cheating, *why* it harms learning (the [[ai-misuse-learning-harm|performance–learning gap]]), and how students can use AI productively — this addresses the "everyone is doing it" peer-norm and rationalization problems documented in [[ai-tools-academic-work-cheating-2026]] and [[student-rationalization-ai-writing]]. Student surveys support that framing. Among 504 sociology students, 65% had used generative AI for coursework but only 3% to generate assignment text and 2% to generate a full draft; meanwhile 81% had received some AI guidance but only 46% found it very clear ([[student-genai-use-views-writing]]). Ambiguity, not defiance, is the practical problem — which is why the instructional-capacity side of this FAQ connects to [[ai-literacy-evidence]] and to the ranking of interventions in [[top-10-findings-ai-education-instructors]]. Students do not describe the tools in the same terms the policy does: [[mulisa-students-genai-integrity-perspectives-2026|Mulisa and Mezgebu (2026)]] found the question students raise is whether generative AI is a tool that facilitates cheating or a partner that supports learning, and their account of the tension between institutional integrity rules and students' own learning needs is more useful for framing expectations than another warning about penalties. ## The bottom line Combine a **structural floor** (guardrails + assessment redesign that make cheating hard regardless of motivation) with **educative capacity-building** (AI literacy, declarations, "think-AI-reflect" sequences). Detection alone is the weakest lever; the goal is to make honest, productive AI use the path of least resistance. --- ## [How Do I Keep Students from Over-Relying on AI?](https://edtechdev.github.io/aied/faqs/reducing-over-reliance/) # How Do I Keep Students from Over-Relying on AI? You have graded work that reads better than the student can explain. They ace the take-home problem set and blank on the exam. Grades have stopped predicting what they understand, and you suspect the tool is doing the thinking. You are probably right, and this is a design problem, not a policing problem. **Over-reliance is not the same as frequent use, and the interventions that reduce it are mostly task designs rather than restrictions on access.** Durable learning is built by retrieval, elaboration and generation, and generative AI can supply the product of those processes without requiring them. The bottom line: put the student's own attempt before the tool, protect the moments where the tool is absent, make verification visible rather than merely available, and grade something the student does unaided. ## What over-reliance actually is, and how to tell you have a problem Cognitive offloading is the transfer of cognitive demands to external tools, freeing limited mental resources for higher-order processing. [[cognitive-offloading-metacognitive-review-2026|Guo and Ye (2026)]] review the construct through Nelson and Naren's dynamic [[metacognition|metacognitive]] model, in which monitoring of difficulty informs a decision to offload to an internal or external strategy. Offloading is a [[self-regulated-learning|self-regulatory]] choice rather than a defect; the failure is [[trust-calibration|miscalibration]], not volume. [[ai-overreliance-complex-adaptive-system-2026|Biswas (2026)]] models reliance as three actions — solve alone, accept the AI's answer unverified, or use it and verify — and defines the two calibration errors symmetrically: over-reliance is accepting wrong output, under-reliance discarding useful AI after it errs. Collective over-reliance is the population abandoning verification, and because raw over-reliance diverges from regret, high reliance is not automatically harmful. The student who checks every AI answer against your materials is not your problem; the one who submits a model's first draft is. Most of what you can measure is [[self-report-measures|self-report]]. [[gerlich-ai-tools-cognitive-offloading-critical-thinking|Gerlich (2025)]] surveyed 666 UK participants with 50 interviews and found AI use negatively correlated with [[critical-thinking|critical thinking]] (r = −0.68), with offloading partially mediating (total effect b = −0.42; indirect b = −0.25). [[genai-over-reliance-learning-2026|Gao, Sun and Khan (2026)]] used three-wave time-lagged survey data from 623 Chinese students plus educator interviews and found effective AI use raises sustainable learning performance *and* over-reliance at once. Both designs are correlational, and both say so. Behavioral instruments are thinner. [[pause-ai-cognitive-offloading-self-reflection-2026|PAUSE (Alam, 2026)]] is a browser-only self-check with no reliability or validity evidence, and its key warning concerns item validity: the items record when and how often AI enters a workflow, not whether the student's own reasoning stayed engaged, so a student who deliberately scaffolds with AI early will honestly score as offloading. It also cites Padmakumar et al.'s (2026) Offloading Score, which estimates the fraction of effort offloaded from behavioral logs (n = 40 developers). Usage frequency is not the construct; what the student can do unaided is. ## How AI displaces the work that produces learning [[lodge-loble-cognitive-offloading-2026|Lodge and Loble (2026)]] frame the risk as "fluency on demand": coherent, confident output that bypasses the [[desirable-difficulties]] — retrieval, elaboration, generation — through which knowledge is consolidated. Their **performance paradox**: AI-assisted work feels fluent and students perform well in the moment while retaining less — an illusion of competence. They name **metacognitive laziness** (after Fan et al. 2024): convenience lets learners abdicate self-regulatory processes they need to develop. [[cognitive-offloading-metacognitive-review-2026|Guo and Ye (2026)]] add a design boundary: **substitutive** offloading replaces internal processing while **duplicative** offloading supplements it; remove the external store and substitutive offloaders decline severely whereas duplicative offloaders hold accuracy through internal encoding. The causal evidence is sharpest there. [[brcic-effortless-trap-productive-struggle-2026|Brcic and Frljic (2026)]] report that an unguarded AI helper left high-school students roughly 17% worse on an unaided exam than peers with no tool, that the same model rebuilt to withhold answers erased the harm, and that a well-engineered tutor roughly doubled learning. Their diagnostic: if letting AI in makes the task feel effortless, it is in the wrong place. PAUSE adds matching findings: Bastani et al. (2025) found students given GPT-4 solved more problems with the tool but performed worse than controls once it was removed, and Liu et al. (2026) found assistance also reduced persistence in randomized controlled trials (N = 1,222). ## What actually moves student behavior, ordered by payoff **1. Withhold or ration what the task is meant to build (highest payoff, moderate cost).** The lever with the largest causal footprint is a tool that refuses to answer: guarded AI (hints, examples, practice) in the middle phases, the secured final check at the end. That makes an AI-use policy a per-skill placement rule rather than a prohibition list. [[zohar-bloom-inzlicht-against-frictionless-ai-2026|Zohar, Bloom and Inzlicht (2026)]] argue the effort–meaning link is an inverted U, so the target is a gradient: remove overwhelming obstacles while preserving the struggles that produce comprehension and ownership, with assistance as supplement rather than substitute. **2. Sequence assistance instead of restricting it (high payoff, high cost — redesign).** The strongest tested structure is "think first, ChatGPT later." [[think-first-chatgpt-later-2026|Wong and Qiu (2026)]] had N = 196 students work independently, with free ChatGPT, or in a regulated condition: generate your own ideas, collaborate with ChatGPT to improve and evaluate them, then independently refine and submit one solution. The free-use group produced more creative work on the assisted task but fell back to human-only levels on a later, harder task done without ChatGPT; the regulated group showed no immediate advantage yet outperformed both others on independent creativity afterward. Process analysis showed 88.6% of its prompts were collaborative, the only prompt type significantly correlated with later independent originality. **3. Make verification visible and required (high payoff, low cost).** [[ai-overreliance-complex-adaptive-system-2026|Biswas (2026)]] found that making verification visible triggered a counter-cascade to near-complete verification (over-reliance 0.00, regret down to 0.07), whereas reducing the friction of checking was the weakest lever because it does not counter the social pull toward unverified use. A required source-check sentence beats a link to the library. Checking output against course materials, peers or instructors when accuracy is uncertain is what [[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026|Viberg and colleagues (2026)]] found stronger students already doing. **4. Time reflection prompts to the phase they can influence (moderate payoff, low cost).** [[cognitive-offloading-metacognitive-review-2026|Guo and Ye (2026)]] derive a principle of timing-component matching: feedback targeting stable beliefs works before a task, while immediate task-specific correctness and difficulty feedback works during it. [[lodge-loble-cognitive-offloading-2026|Lodge and Loble (2026)]] recommend integrated metacognitive prompts that make learners pause, reflect and assess their understanding, alongside Load Reduction Instruction that manages cognitive burden while enabling progressive independence. A one-minute pre-task prediction costs nothing and lands in that window. **5. Require unaided retrieval, explanation and transfer (high payoff, moderate cost).** The assisted product is a poor proxy for capability; the graded moment must include one where the tool is absent. [[think-first-chatgpt-later-2026|Wong and Qiu's (2026)]] later unassisted task is that measurement — and even their human-only group declined on the harder follow-up, so unscaffolded solo work was not the answer either. Alternating modes has support: PAUSE reports Kosmyna et al.'s (2025) session-four result, in which brain-only participants who later used ChatGPT outperformed sustained LLM users. Accountability works through explanation: Makransky et al. (2025) found a tutoring chatbot that prompted students to connect ideas and explain their reasoning produced better assessment performance than traditional instruction. ## What does not work, and what backfires - **Friction reduction.** Making checking cheaper is the weakest lever in Biswas's model; the barrier is social, not mechanical. A better plagiarism checker does not make a student verify. - **Banning or blanket policing.** The evidence points to placement, not prohibition: the tool is not the variable; where it sits in the task is. - **Unscaffolded solo work.** [[think-first-chatgpt-later-2026|Wong and Qiu's (2026)]] human-only group also declined on the harder follow-up. Removing the tool without supporting the learner is not an intervention. - **Counting usage and survey attitudes.** PAUSE's items record when and how often AI enters a workflow rather than whether reasoning stayed engaged, so a careful scaffolder scores as an offloader. Acting on those numbers punishes the students you want. - **Satisfaction and fluency as evidence.** They are the illusion. Prefer unaided performance and delayed [[transfer-of-learning|transfer]]. - **Leaning on student-facing AI literacy instead of supporting teachers.** [[lodge-loble-cognitive-offloading-2026|Lodge and Loble (2026)]] caution that over-investing there may be the wrong allocation. ## Redesign the task so reliance is the harder path Four moves carry most of the weight, none requiring campus policy. **Task design.** Ask for the student's own ideas, hypotheses or a rough draft before the tool sees the task. In Wong and Qiu's regulated condition the sequence was fixed: generate your own ideas, collaborate with ChatGPT to improve and evaluate them, then independently refine and submit one solution. Configure tools to hint rather than answer wherever the target skill is what the task measures. **Verification requirements.** Make the check a deliverable: source-checking, peer comparison, instructor check, written into the assignment rather than assumed. Visible checking norms matter more than cheaper checking. **In-class demonstration of failure modes.** Run the demonstration live: give a task, let students solve it with a confident AI answer that is wrong, then have them check it against the course text. Add the evidence: an unguarded helper left students roughly 17% worse on an unaided exam, and students given GPT-4 solved more problems with the tool but performed worse than controls once it was removed. One demonstrated failure teaches more than ten warnings. **Assessment redesign.** Keep the first hard attempt and the final unaided check AI-free — the two moments [[brcic-effortless-trap-productive-struggle-2026|Brcic and Frljic (2026)]] identify as protected — and grade the unaided one. Ask students to explain their reasoning and to work on a parallel or transferred task, and read the assisted product as performance, not learning. Where the target skill is analysis rather than mechanics, offloading the lower-order part can serve the higher-order one — PAUSE cites Hong et al. (2025), where deliberately offloading lower-order writing tasks to free attention for analysis and revision produced larger critical-thinking gains. ## Healthy help-seeking versus harmful offloading The evidence distinguishes these clearly, and so should your rubric. [[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026|Viberg, Feldman-Maggor and Wong (2026)]] interviewed 20 STEM university students and found a four-stage process — deciding whether help is needed, choosing a source, choosing the type of help, judging the help received — in which stronger students favor *instrumental* help (hints, explanations) over *executive* help (direct solutions). They warn that using LLMs for debugging or cross-language programming can bypass independent [[problem-solving]] even when students avoid asking for answers. Instrumental use — hints, examples, explanations, debugging you then fix yourself — is the duplicative offloading that holds up when the tool goes away. Executive use — take the output and submit it — is the substitutive offloading that collapses. Free use tends toward the second: in [[think-first-chatgpt-later-2026|Wong and Qiu's (2026)]] experiment (N = 196), 70.9% of the free-use group's prompts were non-collaborative and 59.6% simply asked ChatGPT to generate ideas outright. The regulated group shows the opposite signature — 88.6% collaborative prompts, the only type correlated with later independent originality. Write the distinction into the assignment, require a collaborative prompt, and grade the reasoning students add. ## Who is most exposed, and when [[lodge-loble-cognitive-offloading-2026|Lodge and Loble (2026)]] locate the risk in prior knowledge and self-regulation, naming a **metacognitive equity gap**: leveraging AI well requires the resources novices lack, so the students who need the practice most are the likeliest to delegate the learning itself. They report 80% of Australian students already use AI and two-thirds of early secondary teachers do (OECD 2025). [[gerlich-ai-tools-cognitive-offloading-critical-thinking|Gerlich (2025)]] found participants aged 17–25 showed higher AI dependence and offloading and lower critical thinking than those aged 46 and over, and that attainment predicted better critical thinking regardless of AI use (r = +0.34), with a significant interaction indicating it mitigates the negative effect. [[genai-over-reliance-learning-2026|Gao, Sun and Khan (2026)]] found polychronicity — a multitasking tendency — moderates the pathway, with high-polychronicity students at greater risk. Context matters as much as person, which is where your course can act. [[ai-overreliance-complex-adaptive-system-2026|Biswas (2026)]] shows task difficulty and AI quality set the baseline (over-reliance rises from ≈0.02 to 0.38 with difficulty; on hard tasks, 0.38 with poor AI versus 0.16 with good AI) and that the highest regret comes from *high-quality* AI on hard tasks (0.441), because agents over-defer and rarely self-rely — a capable model on a demanding task is where checking stops. Peer exposure compounds it: as visible social proof rises from 0 to 0.6, verification collapses from 0.29 to 0.002. ## "But..." — three objections **"This is just good pedagogy."** Partly — these are scaffolding, formative feedback, and productive struggle. But visible verification beat cheaper verification, and the two protected moments are the first hard attempt and the final unaided check. Do the familiar things, in the new order. **"I cannot police it."** You cannot, and the evidence says you should not try. The strongest result here came from placement, not prohibition. Design the task so the tool's presence at the wrong moment shows up in the work itself — an unexplainable answer, a missing verification step — rather than relying on surveillance. **"My course is too large."** The cheapest levers scale. A required verification line, a one-minute pre-task prediction, and moving the graded check to an AI-free room cost minutes per section. The expensive one — a fully sequenced think-first unit — can start as a single assignment. Regenerate one problem set into assisted and unaided halves and see what the gap tells you. ## What is not yet established (read before you commit) No study in this corpus tests whether a specific over-reliance intervention holds up across settings or semesters. The think-first design rests on one experiment (N = 196), and the withholding result comes from [[brcic-effortless-trap-productive-struggle-2026|Brcic and Frljic's (2026)]] synthesis rather than a trial of its own. The friction argument is a conceptual Comment with no new data whose inverted-U relationship rests on one empirical anchor and is not quantified, so where the optimum sits for a given learner is unspecified. [[ai-overreliance-complex-adaptive-system-2026|Biswas (2026)]] states the model's limits — exogenous stationary AI quality, a fixed network, stylized verification — and names what future work would need to estimate from longitudinal traces: per-task verification rates, social-proof strength, and how trust updates after verified versus unverified use. Measurement remains the weakest link: the dominant designs are surveys, and both [[gerlich-ai-tools-cognitive-offloading-critical-thinking|Gerlich (2025)]] and [[genai-over-reliance-learning-2026|Gao, Sun and Khan (2026)]] call for longitudinal and experimental follow-up. [[pause-ai-cognitive-offloading-self-reflection-2026|PAUSE]] has no psychometrics and was built for adults. Whether the equity gap can be closed by instruction is theorized rather than demonstrated. ## Your action list for this week 1. Add a verification step to the next assignment: one sentence naming what the student checked and against what. 2. Move the graded check to an AI-free moment, and grade that one. 3. Put the first hard attempt before the tool: ideas, hypotheses or a draft before AI sees the task. 4. Configure the tool you recommend to hint rather than answer. 5. Demonstrate one confident AI failure live and have students catch it against the course text. 6. Add a one-minute pre-task prediction before a difficult unit. 7. Replace one quiz item with an "explain your reasoning" item on the same content. 8. Ask students to label their AI use as instrumental or executive, and grade the reasoning they added. 9. Read assisted work as performance, not learning, and compare it against the unaided check. 10. Treat engagement and satisfaction as weak indicators, and prefer unaided performance and delayed [[transfer-of-learning|transfer]]. ## How this page differs from the neighboring FAQs [[does-ai-help-students-learn]] asks whether AI produces learning at all and sets out the performance–learning gap; [[reduce-ai-cheating]] covers integrity, detection limits and assessment security. This page assumes students may be using AI honestly and asks which designs keep the learner's reasoning in the loop. For the surrounding research, see [[does-ai-help-students-learn]] and [[redesign-assessment-ai-era]], and the [[cognitive-offloading]], [[metacognition]] and [[desirable-difficulties]] concept pages. --- ## [What Are Best Practices for Reporting and Interpreting AI in Education Research?](https://edtechdev.github.io/aied/faqs/reporting-interpreting-aied-research/) # What Are Best Practices for Reporting and Interpreting AI in Education Research? You are writing up an AI in education study whose claim is already running ahead of its design, or reviewing one and deciding whether the headline number means anything. Either way the same six omissions decide it: which AI system was actually used, what it was configured to do, what pedagogical role it played, whether humans designed or reviewed its output, what the outcome measure truly captured, and what the authors did about bias, cost and limitations. An author who omits them submits a study that cannot be appraised; a reviewer who does not look for them cannot tell a designed intervention from a recycled tool demo. "Unassessable" — not "supported" and not "refuted" — is the honest verdict when they are missing. The stakes are not hypothetical. [[oneill-presumed-effective-meta-analysis-2026|O'Neill (2026)]] audited 14 peer-reviewed meta-analyses claiming AI improves education and found that *none* provided a valid basis for the claims it advanced. [[bartos-ai-learning-meta-meta-analysis-2026|Bartoš et al. (2026)]] pooled 1,840 effect sizes from 67 meta-analyses and found that once publication bias is modeled, the average effect falls to roughly **one-third** of the published median (SMD = 0.196 versus 0.67). [[citation-errors-hallucinations-computing-education-2026|Denny et al. (2026)]] verified **30 fabricated references across 14 computing-education papers**, all published in 2025 or 2026 — a defect that reaches readers precisely because reviewers cannot verify every entry in a reference list. Reporting discipline decides whether the field's evidence base is usable at all. For the design choices underneath it see [[research-methods-aied|Research Methods in AI in Education]]; for the catalog of failure modes, [[limitations-in-aied-research|Limitations in AIEd Research]]. For the reader's side — what a practitioner does with a paper once it exists — see [[interpreting-and-applying-aied-research|Interpreting and Applying AIEd Research]]. ## The short version Six moves decide whether your claim survives review: 1. **Describe the AI system as a treatment, not a vendor.** Model, version, configuration, prompts, role, and whether humans designed or reviewed the output. 2. **State the pedagogical rationale and the active comparison.** A product name is not a method; business-as-usual is not a fair control. 3. **Make the measure match the claim.** What the instrument captures, its validation for this population, and whether the outcome was delayed and unaided. 4. **Validate every automated judgment.** Name the gold standard, the calibration target and who adjudicated the disagreement. 5. **Report the analysis's uncertainty, not just its point estimate** — dependent effects, τ², prediction intervals, subgroup sizes, publication-bias assessment. 6. **Report the counterweights**: ethics and governance procedure, compute and environmental cost, your own AI use, null results, and only citations you have verified. ## Decide what you will say about the AI system — and write it so someone could rebuild it This is the check most submissions fail. The RAISE framework (*Reporting AI Studies in Education*, Allison 2026) is a 30-item checklist across ten thematic domains, and its diagnostic claim is specific rather than rhetorical: submissions routinely omit which model was used (GPT-4, Claude, or a custom algorithm), how it was configured (prompts, fine-tuning parameters), what role it played (feedback generator, co-author, tutor, evaluator), and whether human actors designed or reviewed its outputs. That leaves reviewers with four unanswerable questions — "What exactly was the AI doing?", "Was it necessary?", "Is this replicable?" and "Are the learning claims credible?" It runs from educational justification through API-level specification, learner–AI interaction, accessibility and cultural fit, participants and setting, human involvement, design, ethics, transparency and reproducibility, to limitations. Its companion **Ethics and Risk Matrix** covers risks to learner [[agency]], [[equity-in-ai-education|equity]], data [[governance]] and algorithmic transparency. [[tep-aied-model-reporting-2026|Hwang, Xie, Wah and Gašević (2026)]] offer a streamlined alternative, TEP-AIED, which folds Transparency, Ethics and Pedagogy into one interdependent structure, supplies a guideline table mapped to seven paper sections, and asks authors to add a Method subsection titled "Transparency, Ethics, and Pedagogy Considerations" — because, on their account, RAISE is comprehensive but too granular for routine empirical use. The disclosure sits under [[ai-use-disclosure]] and is what makes [[privacy]] and [[ethics]] claims checkable. ## Decide whether your treatment is a method or a product Educational justification is RAISE's first domain because AI is not a neutral tool: value depends on alignment between the learning problem, a theoretical rationale, and the mapping from objectives to outcome measures. [[oneill-presumed-effective-meta-analysis-2026|O'Neill (2026)]] found that all but two of the audited meta-analyses defined the treatment as a tool — ChatGPT, GenAI, "AI" — and that a product name is not a [[pedagogy]]; treating exposure to ChatGPT as one common intervention is comparable to meta-analyzing the effects of "paper." [[weidlich-chatgpt-effect-search-cause-2025|Weidlich et al. (2025)]] make the same argument for primary studies: auditing a subset of the comparisons behind a prominent meta-analysis, they found only **21% had a well-defined treatment, a control group, and a valid learning measure**, and the reported effect size (g = 0.7) exceeded that of purpose-built [[intelligent-tutoring|intelligent tutoring systems]] (0.66) — a red flag that the "treatment" was a heterogeneous secret sauce rather than a named method. The test: could a competent reader rebuild your intervention from the method section alone and get the same thing? ## Decide what your instrument measures and what the claim may therefore say In O'Neill's evidentiary audit of 46 randomly selected primary studies, **28 (61%) presented validity concerns**, and dependent-variable mismatch was the most common (n = 15), followed by independent-variable mismatch (n = 11), experimental design problems (n = 7), data extraction problems (n = 6), absence of a control group (n = 6) and nonrandom group assignment (n = 6). Multidimensional outcomes were routinely pooled as if interchangeable — test scores, homework quality, [[motivation]], [[self-efficacy]], attitudes and [[student-engagement|engagement]] collapsed into one "academic achievement" number. Two reporting habits prevent this. State what construct the instrument measures and that it was validated for that population: see [[self-report-measures]] for where perception-based measures diverge from behavior, and [[assessment-validity]] on construct validity. Then report an outcome that does not depend on the learner's belief about the aid, because immediate task performance under assistance is not [[learning-gains|learning gain]]: [[verification-quality-reliance-calibration-genai-2026|a 2026 mini review of 493 records and 14 priority studies]] found none measured verification success and the following reliance decision together against an independently adjudicated standard of output quality, and few looked past immediate performance to delayed retention or transfer. [[does-ai-help-students-learn|The evidence for durable learning]] remains mixed, and [[cognitive-offloading|cognitive offloading]] is the mechanism that most plausibly separates the two. ## Decide how your automated judgments were validated [[ai-ed-evaluation|AIED evaluation]] increasingly delegates scoring to models, which makes the validation procedure part of the result rather than an appendix. Khan Academy's account of its [[intelligent-tutoring|AI tutor]] metrics reports that cognitive engagement is scored by an [[llm]] judge calibrated against human pedagogical experts at F1 0.83, and that metric movements came from more than 40 live experiments in five months rather than offline evaluation ([[ai-tutoring-quality-k12-methodologies-2026|Udeshi et al. 2026]]). [[machines-misread-pedagogical-quality|Tseng et al. (2026)]] show that human–machine disagreement about [[formative-assessment|pretest question]] quality is systematic rather than random, and that rubric operationalization matters more than rationale-first prompting. Agreement with human coders is not quality, and reporting it as if it were is a reviewer target. [[agreement-not-quality-llm-coding-verification|Liu et al. (2026)]] had an independent expert judge 855 pairwise code sets blind to source, finding human–LLM agreement (mean Jaccard 0.30) well below human–human agreement (0.52) while the blind verifier preferred human and machine coding at indistinguishable rates (51.5% vs 48.5%, p = 0.537). Report the calibration target, the gold standard, and who adjudicated it. ## Decide whether the comparison is fair and how far the claim travels Fairness and scope fail the same way: a large effect that means less than it appears. 11 of the 14 audited meta-analyses set no population boundary, so "students" spanned children through medical trainees; a single course, institution, discipline or country does not license a general claim, and results for one tool rarely transfer to another. Novelty effects, extra time on task, and a comparison condition that received business-as-usual instruction rather than an active control are the usual explanations for a large effect — so report what the control actually did, how much time each arm spent, and how novel the tool was. State the population boundary as well, because effects differ across learner groups and an unstated boundary is what lets one pooled number stand in for all of them (see [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]]). Treat the tool as a moving target too: proprietary systems change without notice, so a finding is bound to a model version, and the [[benchmark]] that thrilled in one release may not hold in the next. ## Decide what your analysis can support — and report its uncertainty, not just its headline For syntheses, report dependent-effect handling, between-study variance, prediction intervals and publication-bias assessment, not just I². O'Neill found that reported heterogeneity was severe wherever it was given (I² from 77.2% to 94.4% across 13 meta-analyses, 12 of them above 80%) and never resolved (no moderator analysis met the minimum subgroup size of ten studies), that 12 meta-analyses treated dependent effect sizes from the same study as independent — inflating apparent evidence — and that only four reported between-study variance (τ²) and only two a prediction interval, both of which included zero. Publication bias was not validly assessed anywhere. Because I² is precision-dependent, a heterogeneity figure has to travel with its model: [[limitations-in-aied-research|Limitations in AIEd Research]] documents one synthesis whose heterogeneity is I² = 82.98% under a fixed-effect model and 15.75% under random effects, so a reader given only the first number cannot tell how inconsistent the corpus is. Report τ² and a prediction interval alongside it; see [[meta-analysis-systematic-review|meta-analysis and systematic review]] for the review-side conventions, and discount the headline accordingly — Bartoš et al.'s bias-adjusted average was SMD = 0.196 with a prediction interval running from −1.521 to +1.908, spanning substantial harm to substantial benefit for a hypothetical new study. ## Decide what you disclose about ethics, cost, governance — and your own AI use A review of all AIED 2025 conference papers found an "LLM adoption without disclosure" pattern: most projects used LLMs, but fewer than a handful reported resource consumption or carbon footprint. That paper supplies an open-source method with measuring tools for local and cloud hardware plus a formula for estimating the computational expense of frontier models whose parameter counts are undisclosed, and argues that not reporting these costs is itself an ethical concern. Add data governance, consent, and accessibility/cultural fit (RAISE items; [[universal-design-for-learning]], [[accessibility]]) — see [[equity-ethics-pedagogical-safety-research]] for the fuller treatment of these obligations. Your own use of AI in the research process needs the same disclosure. [[prisma-llm-ai-assisted-systematic-reviews-2026|Zabaleta and Lin's PRISMA-LLM analysis]] of 888 review-automation papers shows how uneven this reporting is: since 2023, **38.0% of software/product papers reported no evaluation at all** (against 9.3% of LLM papers), mean reporting richness was 6.3 for LLM papers versus 3.3 for software/product papers, and 84.1% of LLM usage relied on proprietary or hosted systems with only 4.9% open-weight. Their framework separates implementation disclosure from consequence-sensitive evaluation across five disclosure tiers — a workable template for describing what a tool did in a review. [[dai-chan-responsible-genai-research-ai-literacy-2026|Dai and Chan (2026)]] add the author-side picture: 27 of 28 postgraduate researchers used GenAI across ideation, literature review, explanation, data processing, programming, [[writing-education|academic writing]], editing and translation, calibrating use to stakes and disciplinary norms, while noting that institutional policies address teaching and assessment rather than research practice. ## Decide what to do with null results and citations Denny et al. verified 30 fabricated references across 14 papers, with the verified count at one technical symposium rising from 3 in 2025 to 17 in 2026 (2.3% of that year's proceedings papers) — and LLM-assisted drafting makes a plausible invented citation cheap to produce. Their counterweights matter when you cite that audit: most flagged references were benign — 229 were ACM metadata mismatches where the PDF was correct and 188 were valid bibliographic variants — so the fabricated count is a deliberate lower bound and the benign majority is not a rounding error. Verify every entry you cannot personally check. Reporting null and negative findings is the other side of the same coin. [[bartos-ai-learning-meta-meta-analysis-2026|Bartoš et al. (2026)]] found strong evidence of suppressed null results (every Egger test p < .0001) and extreme heterogeneity (τ = 0.869), with no outcome, field, level or AI-role subgroup showing consistent gains. They also found no difference between studies published before and after January 2023, which undercuts the claim that modern generative tools specifically produce gains. ## The field's documented weaknesses, in its own numbers The base rates against which a claim should be weighed: - **Validity:** 28 of 46 vetted primary studies (61%) had validity concerns; only 21% of audited comparisons had a well-defined treatment, control group and valid measure. - **Unresolved heterogeneity:** I² from 77.2% to 94.4% across 13 meta-analyses, 12 above 80%, no moderator analysis reaching the ten-study minimum, dependent effects treated as independent in 12, and only four syntheses giving τ² against only two a prediction interval. - **Publication bias:** unassessed across the 14 audited meta-analyses; every Egger test in the larger pool was significant (p < .0001), with τ = 0.869. - **References:** 30 fabrications across 14 papers, 17 in one 2026 symposium — 2.3% of that year's proceedings. ## The objections you will hear **"Reviewers want novelty, not method detail."** Journals are moving the other way: RAISE and TEP-AIED give editors a construct to require, and TEP-AIED's Method subsection is a low-friction entry point. Method detail is what converts an unassessable paper into a citable one. **"We do not have the space."** The five disclosure tiers are reporting levels rather than risk levels, so the pipeline's evaluation depth can be a sentence or two. Version, prompts and role fit in a short paragraph; cost and governance fit in a limitations sentence. **"Everyone reports it this way."** That is the finding, not a defense — 61% of vetted primary studies with validity concerns, 12 of 13 meta-analyses above 80% heterogeneity, 38.0% of software/product papers with no evaluation. The norm is the defect. **"Our tool is proprietary, so we cannot report the model."** Report version, date accessed, configuration, prompts you supplied, role and guardrails. If the vendor will not say, that constraint belongs in the limitations as a bound on the claim. **"Reporting nulls will hurt us."** Suppressed nulls are the documented mechanism behind the SMD = 0.196 bias-adjusted average; publishing them is what keeps your positive result readable. ## Pre-submission checklist - **Report:** the model, version, configuration and prompts, and the AI's role in the design. **Check:** could you rebuild the intervention from the method section alone? - **Report:** the pedagogical rationale and the active comparison condition. **Check:** is the treatment a named method rather than a product name? - **Report:** what each measure captures, its validation for this population, and the delayed or unaided outcome. **Check:** does the headline claim match the dependent variable? - **Report:** who adjudicated automated judgments, against what gold standard, with what agreement. **Check:** is accuracy or agreement being used as if it were quality? - **Report:** dependent-effect handling, τ², prediction intervals, moderator subgroup sizes and publication-bias assessment. **Check:** is heterogeneity quoted with its model and its precision? - **Report:** human involvement, data governance, consent, accessibility and cultural fit, plus compute and environmental cost. **Check:** are ethics claims backed by described procedure rather than asserted principles? - **Report:** AI use inside the research process, across all five disclosure layers. **Check:** was the reviewing or analysis pipeline evaluated at all? - **Report:** null, negative and disconfirming results alongside positive ones. **Check:** is the effect size discounted for likely publication bias? ## What would actually raise the standard The remedies in this literature are institutional rather than individual. O'Neill recommends mandatory data transparency for meta-analyses — full effect-size tables, dependency structures, τ² and prediction intervals — plus stronger reviewer and editorial gatekeeping and a functioning retraction practice, on the argument that these failures are the products of failed gatekeeping rather than isolated errors. PRISMA-LLM proposes five disclosure tiers as reporting levels rather than risk levels, so that a review pipeline's evaluation depth is stated plainly; RAISE and TEP-AIED give journals a construct to require in submissions. For related ground see [[research-gaps-aied|notable gaps in the research literature]], [[limitations-in-aied-research|Limitations in AIEd Research]], and [[evaluating-ai-interventions-methods|measures and methods for evaluating AI-related interventions]]. One low-cost remedy sits with authors: every article page in this knowledge base now pairs a study with a **What this means for practice** section and usually a **Limitations** section, so writing the paper such that both can be filled honestly — the action a practitioner could take, and the constraint on it — is a reporting habit as much as a writing one (see [[interpreting-and-applying-aied-research|Interpreting and Applying AIEd Research]]). --- ## [What Are Notable Gaps in the Research Literature on AI in Education?](https://edtechdev.github.io/aied/faqs/research-gaps-aied/) # What Are Notable Gaps in the Research Literature on AI in Education? The most consequential gaps in AI in Education (AIED) concern **whether particular educational designs produce durable benefits, for whom, through which mechanisms, and under what conditions—not simply whether AI can perform educational tasks**. The knowledge base documents promising interventions alongside persistent weaknesses in measurement, causal inference, generalizability, implementation, and reproducibility. These are often gaps in the *strength, specificity, or applicability* of evidence rather than a complete absence of research. See [[limitations-in-aied-research|Limitations in AIEd Research]]. The gaps also differ across the field. Evidence about established [[intelligent-tutoring|intelligent tutoring systems]], predictive [[learning-analytics|analytics]], [[generative-ai|generative AI]], and autonomous agents should not be treated as interchangeable. Likewise, using AI to support learning and teaching people *about* AI involve related but distinct research questions, as the [[ai-education|AI in Education]] overview explains. ## 1. Isolating what AI adds beyond good instruction Rigorous classroom experiments exist, so the gap is no longer adequately described as “we need randomized trials.” A more precise question is **whether the AI component adds value beyond additional practice, better materials, timely feedback, or increased instructional support**. For example, [[one-click-away-khanmigo-two-year-school-experiment-2026|One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment]] reports modest achievement gains across 18 middle schools, alongside limited substantive engagement with the tutor. The authors note that the gains resembled those associated with structured practice without AI. This establishes evidence about an implemented instructional package, but does not cleanly isolate the incremental contribution of its AI component. The problem underneath is conceptual as well as empirical. "ChatGPT" names a tool, not a method, and [[weidlich-chatgpt-effect-search-cause-2025|Weidlich et al. (2025)]] audit 19 ChatGPT-in-education comparisons to show what that costs: only 4 (21%) specify all three of a replicable treatment, an operationalized control, and a valid learning measure (74% well-defined treatment, 42% well-defined control, 53% a learning outcome). Their larger point is that when new AI is introduced alongside new activities, feedback, or interface design, the medium and the method are confounded and no effect can be attributed to the AI. Research therefore needs more comparisons with strong, realistic non-AI alternatives, holding curriculum, practice opportunities, and support as constant as possible. Independent replications should test whether benefits survive changes in institution, instructor, subject, and model. As [[research-methods-aied|Efficacy Research Methods]] emphasizes, different methods answer different questions: [[qualitative-research|qualitative]] and [[design-based-research|design-based]] studies help explain implementation, while appropriately designed experiments strengthen causal claims. ## 2. Following durable learning and independent capability over time Improved work while using AI is not necessarily evidence of learning that persists after assistance ends. The [[learning-gains|Learning Gains]] and [[cognitive-offloading|Cognitive Offloading]] syntheses repeatedly distinguish assisted performance from retained knowledge, independent reasoning, and transfer to unfamiliar tasks. [[making-ai-tutoring-productive-mastery-math-2026|Making AI Tutoring Productive]] illustrates the measurement problem. In a randomized experiment involving more than 6,000 middle-school students, a three-correct-in-a-row mastery rule increased platform-defined success without, by itself, producing detectable learning gains one week later. The strongest delayed-test evidence emerged when AI was embedded in the mastery workflow and was concentrated on practiced material. **The remaining gap concerns trajectories of capability, not merely an additional post-test.** Studies should examine retention over months, transfer across tasks, performance after assistance is withdrawn, and learners’ accuracy in judging what they know. They should also distinguish failure to acquire a skill from deterioration of an already-established skill. Research on offloading should test when delegation supports those trajectories and when it displaces the practice necessary to develop them. ## 3. Explaining which instructional components work—and why “AI-supported learning” often combines several changes: new feedback, additional reflection, peer discussion, different task sequences, and altered assessment incentives. A successful package does not establish which components are necessary or which mechanism produced the benefit. The multisite experiment [[genai-feedback-design-multisite-experiment|Human-centered GenAI feedback design in higher education]] provides a useful advance. Among 1,176 first-year undergraduates, reflective and hybrid feedback designs outperformed direct [[ai-feedback-quality|AI feedback]] on delayed AI-free transfer. The hybrid condition combined self-evaluation, [[peer-assessment|peer feedback]], and AI critique. This supports investigating how feedback is organized and used, rather than treating access as the intervention. Further studies should isolate the contribution and timing of initial independent attempts, self-explanation, peer input, corrective feedback, hints, and fading assistance. They should test how those components interact with [[prior-knowledge|prior knowledge]] and task difficulty. The accompanying theory gap is equally important: **naming a [[learning-theories|learning theory]] is not the same as testing it**. Research should connect a theoretical prediction to a specific system behavior and a measurable learning process. Statistical mediation can inform that explanation, but does not by itself establish a causal mechanism. See [[theory-development-aied|Theory Development in AI in Education]] and [[scaffolding]]. ## 4. Validating measures, automated judgments, and simulated learners Constructs such as “engagement,” “[[critical-thinking|critical thinking]],” “AI literacy,” and “[[personalized-learning|personalization]]” are measured inconsistently. Self-reports can describe perceptions and experiences, but cannot substitute for demonstrated competence. Technical accuracy, expert-rated output quality, and student learning also represent different evaluation targets. These distinctions are central to [[educational-measurement|Educational Measurement]] and [[ai-ed-evaluation|AI Ed Evaluation]]. A particularly important gap concerns [[ai-technologies|AI systems]] used to evaluate other AI systems. In [[llm-student-simulation-misconception-faithfulness|Simulating Students or Sycophantic Problem Solving?]], [[simulating-students|simulated students]] frequently abandoned assigned [[misconceptions]] after corrective feedback regardless of whether it addressed the misconception. Their responses could therefore make ineffective instruction appear successful. Targeted training improved the study’s faithfulness measure, but improvement on that measure is not equivalent to validation against human learning. Research needs to establish which automated scores and simulated behaviors predict outcomes with real learners, including learners and settings not used during development. Human judgments also require scrutiny: agreement among raters is not automatically evidence that the right construct is being assessed. **The gap is validation of the evaluation chain—from model behavior, to [[pedagogy|pedagogical]] judgment, to learner response, to educational outcome.** ## 5. Establishing assessment validity when AI can produce and evaluate the evidence AI creates two connected assessment problems: it can help produce the work being assessed, and it can influence how that work is scored. The knowledge base’s [[ai-agents-complete-lms-assessment-validity-2026|study of AI agents completing assessed LMS tasks]] documents agents navigating a live undergraduate course and completing assessed activities. These demonstrations establish a capability that challenges assumptions about student-produced evidence; they do not establish the prevalence of such use or invalidate every asynchronous assessment. The research question is **which assessment designs still support defensible conclusions about the learner**. [[eportfolio|Portfolios]], reflections, staged submissions, and activity logs should themselves be validated rather than assumed to establish authorship or understanding. Studies should examine combinations of evidence against independently observed competence, while accounting for accessibility, workload, privacy, and student anxiety. See [[assessment-validity|Assessment Validity]]. For [[automated-assessment|automated scoring]], [[llms-do-not-grade-essays-like-humans-2026|LLMs Do Not Grade Essays Like Humans]] reports systematic disagreements between out-of-the-box models and human raters. Its findings are configuration-specific, but demonstrate why internal consistency is insufficient. A useful conceptual distinction comes from [[human-capability-test-learning-outcomes-ai-2026|A Human Capability Test for Learning Outcomes in the AI Era]]: assess what learners must do independently, what they may accomplish with AI, and what they must verify and defend. That is a proposed framework requiring empirical validation, not an established assessment solution. [[karr-ai-detection-humanization-2026|Karr et al. (2026)]] quantify why detection is a dead end. On 642 published English abstracts, two commercial AI detectors at τ = 0.50 flagged guideline-compliant light AI editing at 38–80%, flagged unmodified 2023–25 originals at 9–15% (non-[[stem-education|STEM]] far above STEM, p < 0.001), and after humanization caught fewer than 4% of AI-labeled rewrites (false-negative rate > 96%). A score that penalizes honest assistance while missing evasion cannot be the basis for a defensible conclusion about the learner; the gap it exposes is designs and process evidence that do not depend on such a score. ## 6. Showing that AI literacy transfers into responsible behavior AI-literacy intervention research is substantial enough to support synthesis. [[liu-ai-literacy-interventions-meta-analysis-2026|AI Literacy Interventions in Education: A Meta-Analysis of Effects and Moderators]] includes 59 studies and 7,211 participants. It reports a positive average effect, but substantial variation across studies and a wide prediction interval spanning zero. Knowledge-focused outcomes showed stronger effects than skills, attitudes, or ethics. The sharper gap is therefore not simply developing more competency frameworks. It is determining **whether literacy instruction changes how people act when using AI**. Can learners recognize unsupported claims, verify sources, reject misleading suggestions, identify inappropriate agreement, and choose when not to delegate? Do those behaviors persist under time pressure and transfer to unfamiliar systems and disciplines? Research should combine performance-based assessments with observations of actual decisions and delayed follow-up. It should distinguish conceptual knowledge, operational proficiency, and critical judgment rather than treating them as interchangeable. The [[ai-literacy|AI Literacy]] and [[trust-calibration|Trust Calibration]] syntheses provide useful starting points for these distinctions. ## 7. Understanding generalizability and equitable outcomes—not just equitable access Findings from one course, institution, language, or learner population often provide limited grounds for decisions elsewhere. The [[limitations-in-aied-research|Limitations in AIEd Research]] synthesis identifies this as a recurring problem. More evidence is needed about how instructional effects vary across developmental stages, disciplines, prior knowledge, disability, language, and resource conditions. Equity research must also distinguish access, skills, and outcomes. The [[digital-divide|Digital Divide]] synthesis makes clear that providing devices or tool access does not establish equal capacity to benefit. For example, [[school-ai-education-readiness-gaps-agency-2026|Does School-Based AI Education Narrow Readiness Gaps?]] followed 752 Hong Kong junior-secondary students. Psychological readiness gaps narrowed, while differences on an objective AI-literacy test persisted. All groups improved, but overall improvement did not eliminate inequality. Because prior-learning profiles were not randomly assigned, the study does not establish their causal effects. The knowledge base's own coverage shows where the group-level evidence is thin as well as uneven. [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]] counts the literature by who it studies: second-language and multilingual learners and students with disabilities are the deepest strands, gender and neurodiversity next, while first-generation students and international students have a single study each in this corpus, gifted and high-achieving students are effectively unstudied as a group, and refugee, immigrant and displaced learners appear in none. An evidence base with that shape cannot answer equity questions by aggregation; it has to be sampled deliberately. **The research priority is identifying which designs reduce differences in demonstrated capability, participation, and agency.** Studies should examine subgroup outcomes and burdens, not merely average gains. Accessibility research should distinguish removing barriers to participation from replacing a capability the learner is intended to develop. See [[equity-in-ai-education|Equity in AI Education]] and [[accessibility]]. Capacity also varies within a single national system in ways governance categories do not capture. [[adeniranye-ai-integration-nigerian-higher-education-2026|Adeniranye et al. (2026)]] scored AI integration across 45 Nigerian universities and found only moderate overall adoption (M = 4.79, range 1.83–7.83 on a 10-point scale), with institution type failing to predict integration once age and geography were controlled (age β = 0.43; South-West location β = 0.31). Internal capabilities intercorrelated at r = 0.79–0.80 and international collaborations with industry partnerships at r = 0.74, so well-connected institutions accumulate compounding advantages. The gap is to test which capacity-building designs change outcomes at newer, less-connected institutions. ## 8. Determining how control should be shared between learners, teachers, and agents As AI systems plan, initiate actions, maintain memory, and coordinate tools, the educational question becomes more specific than whether human–AI collaboration is beneficial: **who should control which parts of the learning process, and when should that control change?** [[agentic-ai-education-scoping-review|Agentic AI in Education: A Scoping Review]] maps 474 studies and identifies limited longitudinal validation, concentrations in [[higher-ed|higher education]] and STEM, and weak educational-theory integration. Only 29% of the reviewed studies explicitly drew on educational theory—a finding about that corpus, not all AIED research. Research should compare configurations in which learners or agents initiate help, set goals, select strategies, monitor progress, and make final decisions. It should test whether support can be gradually withdrawn as competence develops and whether learners retain the ability to challenge the system. The [[agentic-ai|Agentic AI]] and [[human-ai-collaboration|Human–AI Collaboration]] syntheses also raise questions about [[teacher-role|teacher]] intervention and accountability in multi-agent environments. Greater autonomy should be evaluated as a pedagogical design choice, not assumed to represent educational progress. ## 9. Connecting pedagogical and relational safety to real educational consequences Educational safety extends beyond factual accuracy, offensive content, or prohibited requests. A tutor can provide a correct answer while undermining the learner’s opportunity to reason, reinforcing an underlying misconception, or encouraging inappropriate dependence. See [[pedagogical-safety|Pedagogical Safety]]. [[hazra-safetutors-pedagogical-safety-2026|SafeTutors: Pedagogical Safety in AI Tutoring]] identifies failures such as excessive answer disclosure and abandonment of scaffolding, with substantially more failures under multi-turn testing. These are [[benchmark]] findings under specified testing conditions—not estimates of the prevalence or severity of harm in classrooms. The unresolved issue is how such failures affect real learners over sustained use. Which produce temporary confusion, persistent misconceptions, reduced motivation, or weakened independent capability? Which safeguards reduce those risks without excessive refusal or frustration? Longer-term research should also examine trust, willingness to seek human help, [[agency|learner agency]], and relationships with peers and teachers. These questions are especially important for children and require developmentally appropriate studies that connect system behavior to educational and relational outcomes. ## 10. Explaining how implementation, teacher development, and costs shape outcomes Technical capability does not establish that a tool will be used productively or that professional development will improve student learning. The missing link often runs from **teacher preparation, through changed classroom practice, to student outcomes**. In [[pedagogy-first-technology-second-teacher-knowledge-2026|Pedagogy First, Technology Second]], a multilevel study of 46 teachers and 2,832 students found that pedagogical AI knowledge was associated with students’ perceptions and intentions, but neither measured teacher-knowledge component was directly associated with student AI-knowledge gains. These associations do not establish that a particular training intervention would cause better learning. Research should investigate which combinations of coaching, [[curriculum-design|curriculum alignment]], review routines, scheduling, and institutional support produce sustained improvements. It should observe enacted teaching, not only teacher confidence or intention to adopt. See [[teacher-ai-competency|Teacher AI Competency]] and [[educational-development|Educational Development]]. Comparative cost-effectiveness is another priority. Evaluations should include verification, correction, training, supervision, maintenance, and implementation time—not just subscription or model-use costs—and compare AI-supported provision with realistic alternatives. The relevant question is what educational benefit the complete arrangement delivers for the resources it requires. ## 11. Evaluating governance, privacy, and meaningful participation [[ethics|Ethical]] principles and governance frameworks are necessary, but their existence does not establish that they change practice or protect learners. [[agarwal-ethical-values-norms-aied-2026|Identifying the Ethical Values and Norms for Artificial Intelligence in Education]] reviews 25 articles and finds end users largely passive in the reviewed ethics literature, with student voices essentially absent. It also identifies tensions among values and power asymmetries between stakeholders. This describes the reviewed literature; it should not be generalized into a claim that students never participate in AIED design. The research gap concerns **which governance arrangements make a measurable difference**. Does student and teacher participation change procurement, tool design, assessment rules, or remedies after mistakes? Do human-review procedures catch consequential errors? Are alternatives to AI use genuinely available? Privacy research should similarly examine the educational value of additional data collection rather than assume that more detailed learner monitoring is justified. Studies can compare data-minimizing designs with more intrusive alternatives, assessing both learning and learner autonomy. See [[governance|AI Governance]] and [[privacy]]. ## 12. Building reproducible studies and trustworthy cumulative evidence AIED faces an unusually difficult reproducibility problem. Studies may omit prompts, model versions, settings, code, or instructional details; proprietary systems can also change during or after an intervention. These issues are documented in [[limitations-in-aied-research|Limitations in AIEd Research]]. Reproducibility requires describing the instructional arrangement as well as the model: learning tasks, content sources, permitted actions, interface, teacher support, assessment conditions, and changes during deployment. Researchers should distinguish reproducing one configuration from testing whether its pedagogical principle transfers to another. The reporting discussion in [[research-methods-aied|Efficacy Research Methods]] addresses this need for transparent descriptions. Evidence synthesis requires comparable care. The [[meta-analysis-systematic-review|Meta-Analysis and Systematic Review]] page highlights weak primary studies, publication bias, heterogeneous interventions, and incompatible outcomes as limitations on pooled conclusions. A single average “AI effect” can conceal the distinctions educators most need. Reviews should separate assisted performance from independent learning, distinguish intervention types and comparison conditions, and make coding and analytic decisions auditable. Null findings, failed implementations, and boundary conditions are essential contributions to this cumulative evidence base. The synthesis literature is itself a gap. [[oneill-presumed-effective-meta-analysis-2026|O'Neill's (2026) forensic audit]] of 14 high-impact [[meta-analysis-systematic-review|meta-analyses]] found that none provided a valid basis for its claims: no coherent construct (a tool treated as a single intervention and multidimensional outcomes pooled), invalid publication-bias assessment in all 14, reported I² between 77.2% and 94.4% in every analysis that reported it (12 of the 13 above 80%), and 61% of randomly vetted primary studies carrying validity concerns. Those 14 analyses had accumulated more than 2,000 citations in roughly 16 months, and a retracted meta-analysis was still cited as authoritative in 60% of post-retraction citing papers without acknowledging the retraction. Uncritical uptake is not confined to retracted work: in a sample of 14 papers citing another audited meta-analysis, whose abstract advertised a large effect of "AI education" that in fact measured teaching students about AI, only 2 cited it appropriately while 8 read it as evidence that integrating AI improves learning and 4 were incorrect in other ways. The audit's recommendations target that chain directly, asking journals to require full data transparency for meta-analyses (search protocols, coded study characteristics, extracted statistics, and analytic code), asking editors not to treat a publication record as proof of reviewing competence, and asking that retractions be made visible wherever an article is discovered, exported, or cited. Reproducibility therefore covers the integrity of the synthesis chain, not only of individual studies. Instructors who want the classroom-level version of these concerns — which measures and comparison designs to use — can follow the method guidance in [[evaluating-ai-interventions-methods]]. ## Overall takeaway The central research need is not simply more studies showing that students like AI, teachers save time, or AI-supported work receives higher scores. It is stronger evidence answering: > Which educational design, for which learners, in which context, through which mechanism, produces which durable human capabilities—and with what distribution of benefits, costs, and harms? Answering that question requires complementary methods: well-specified experiments, longitudinal follow-up, validated assessments, qualitative and design-based investigation, equity-focused sampling, and transparent synthesis. The aim is an evidence base that explains not only whether an intervention worked, but why it worked, where it may fail, and what educators can responsibly carry into another setting. --- ## [Should We Use AI Detectors?](https://edtechdev.github.io/aied/faqs/should-we-use-ai-detectors/) # Should We Use AI Detectors? **Not as evidence in a misconduct case, and not as an institution's first line of defense. A detector score cannot be validated against ground truth, cannot be cross-examined, and does not meet the balance-of-probabilities standard that [[academic-integrity|academic integrity]] findings require.** Worse, its errors are patterned rather than random: the strongest controlled study in the knowledge base found that detectors flag honest, guideline-compliant AI *editing* far more readily than unmodified student prose, while deliberate evasion passes almost untouched — and the students most likely to be flagged are non-native English writers ([[karr-ai-detection-humanization-2026]], [[teichmann-detecting-undetectable-misconduct-2026]]) — one of several places where an AI system's error rate is not distributed evenly across learners, as [[differential-effects-across-learner-groups|Differential Effects Across Learner Groups]] documents. [[bassett-ai-detectors-education-2026|Bassett et al. (2026)]] go further and argue detection should not be used in education at all, because the technology cannot tell "work created *with* AI" from "work created *by* AI." This page is for the people who have to decide: instructors, academic integrity officers, and administrators. The design-side playbook lives in [[reduce-ai-cheating|How Can I Reduce AI Cheating in My Course?]]. ## 1. The accuracy numbers do not support a finding Three independent lines of evidence converge on the same conclusion. **The false-positive rate on legitimate work is high, and it punishes transparency.** A controlled study of 642 published English abstracts across four domains and two time periods found that at a 0.50 threshold, two commercial detectors flagged guideline-compliant *light* AI editing at **38–80%**, flagged unmodified 2023–25 originals at **9–15%** (non-[[stem-education|STEM]] far above STEM, p<0.001), and — once text had been run through a humanizing service — caught **fewer than 4%** of AI-labeled rewrites, a false-negative rate above 96%. The authors describe this as an integrity catch-22: students who disclose and edit lightly are the ones the tool catches, while students who deliberately evade it are not. Their recommendation is explicit — detector scores should never serve as standalone misconduct evidence.([[karr-ai-detection-humanization-2026]]) **The best independent benchmark does not reach usable accuracy.** In the most comprehensive early [[benchmark]], none of fourteen tools reached 80% accuracy, and simple paraphrasing, minor editing, or humanizing services roughly halve even that performance. At realistic base rates, false positives outnumber true positives.([[teichmann-detecting-undetectable-misconduct-2026]]) **And it misses wholesale AI use entirely when it counts.** In a covert field study, researchers injected wholly AI-generated submissions into a live online [[summative-assessment|examination]] system across five psychology modules: **94% went undetected**, and the AI work on average outscored the real students. Contract cheating already demonstrated the same structural problem — it leaves no reliable trace to find.([[teichmann-detecting-undetectable-misconduct-2026]]) Accuracy also varies by task in ways a policy cannot anticipate. When researchers tested whether [[generative-ai|generative AI]] can reliably detect its own output, detection was dependable for programming and longer reflective writing but poor for short answers, where the model often judged its own text as *more* human-like than authentic student work — and minor prompt variations sharply reduced accuracy.([[llm-detecting-llm-generated-content-education]]) Any threshold you set will hold for some assignments and fail for others. **Detection has never been clearly better than a careful human, and the margin is not the point.** [[leaton-gray-ai-digital-cheating-ethical-pedagogies-2025|Leaton Gray, Edsall and Parapadakis (2025)]] report machine detection of AI or paraphrased text at roughly **80%** against **78.4%** for human reviewers, and cite evidence that AI-generated text has passed as human-authored in **up to 80% of cases** — which they read as a margin far too narrow to ground a misconduct finding, since the machine's advantage disappears into the same error band the human brings. Their review also undercuts the assumption that detection deters the capable: Krou et al.'s meta-analysis finds self-efficacy correlates negatively with cheating while actual ability does not correlate inversely with it at all, so students who could do the work may cheat when they judge the assessment unfair. Detection is therefore neither a reliable instrument nor an obvious deterrent. ## 2. The errors are patterned, and they land on the wrong students This is the part that should decide the question for anyone responsible for equity. The features detectors treat as signals of AI — long-token density, academic word frequency, uniform style — are also features of competent second-language writing, so detectors misclassify non-native English speakers systematically rather than randomly. The harm of misclassification therefore lands on students who are already disadvantaged.([[teichmann-detecting-undetectable-misconduct-2026]]) The same skew appears by discipline: non-STEM abstracts were flagged far above STEM ones in the abstract study, which is a property of the writing conventions of those fields, not of their authors' conduct.([[karr-ai-detection-humanization-2026]]) In effect, a detector is a style test, and the style it punishes correlates with language background, discipline, and register — not with whether a student used AI. That is an [[equity-in-ai-education|equity]] problem before it is a technical one, and it is also a [[trust]] problem. Because detectors are unreliable and formal processes demand detection-grade proof that is functionally unavailable, faculty end up with what one practitioner account calls "suspicion without recourse," while [[student-rationalization-ai-writing|students rationalize]] their own use and case files stall.([[best-response-student-ai-dialog-2026]]) ## 3. Why a score cannot carry an integrity case If you sit in a hearing, this is the section that matters. - **The evidential standard is not met.** Academic misconduct findings require evidence meeting the balance of probabilities. Detector scores — alone *or* in combination with linguistic markers, style comparisons, [[llm]] judgments, or a student's silence — do not satisfy that standard.([[bassett-ai-detectors-education-2026]]) - **Silence and speech are both being misused.** Students retain the right to silence; refusing to respond does not tip the scales against them. The legitimate question is not "did you use AI?" but whether the student can demonstrate the [[learning-gains|learning]] the [[assessment]] claims to measure, which an oral response can answer.([[bassett-ai-detectors-education-2026]]) - **The probability is unverifiable.** Unlike a spam filter or a medical test, a detector's output cannot be independently checked: in real submissions there is no ground truth about how the text was produced, so validation becomes circular, and false-positive/false-negative metrics apply only in controlled tests.([[bassett-ai-detectors-education-2026]]) - **The reference data is outdated by construction.** Detectors are trained and tested on pre-generative-AI human writing — Turnitin's own validation used 700,000 pre-2019 papers — which assumes that corpus reflects current student prose. That assumption is unverified, and it shifts with every model release.([[bassett-ai-detectors-education-2026]]) - **The tool cannot be interrogated.** Detectors publish no thresholds or training data and do not permit independent replication, so a flagged student cannot answer or cross-examine the accusation. You cannot defend a finding you cannot explain.([[teichmann-detecting-undetectable-misconduct-2026]]) - **The binary is the wrong question.** Students' work is frequently created *with*, not *by*, AI, across a hybrid continuum, and policies that say "in assessment" rarely define when an assessment begins — leaving enforcement to subjective judgment rather than principled criteria.([[bassett-ai-detectors-education-2026]]) - **You take on data risk.** Detectors store student work on third-party servers, sometimes overseas under weaker [[privacy]] protections, creating breach, retention, and commercial-exploitation exposure your institution owns.([[bassett-ai-detectors-education-2026]]) - **More monitoring can backfire.** Deterrence runs through a detector's *discrimination* between hidden use and legitimate work, not its raw catch rate. When extra sensitivity produces more new false positives than new true positives, stronger monitoring makes concealment relatively *more* attractive — you spend credibility on students who did nothing wrong, faster than you identify the ones who did.([[mohamed-temimi-assessment-imperfect-information-disclosure-2026]]) ## 4. What to do instead **Ask for verification, not provenance.** Grand Canyon University moved from "Did this student use AI?" to "Can this student demonstrate understanding of what they submitted?", using short conversations, early drafts, and recorded explanations, implemented institution-wide in fall 2025. The argument is practical: faculty are already qualified to judge understanding, and the shift restores their authority instead of leaving them waiting on proof that will never arrive.([[best-response-student-ai-dialog-2026]]) **Make at least one high-stakes task unaided.** Asynchronous oral assessments — just-in-time prompts with brief, time-limited recorded responses graded against embedded rubrics — performed comparably to in-person multiple-choice exams in one study and significantly better in another, with students reporting more active preparation and higher perceived professional relevance. For an [[administrator]], the relevant property is that this is scalable without proctoring.([[asynchronous-oral-assessment-2026]]) **Redesign tasks so AI shortcuts are less attractive.** A case study of undergraduate economics students found only about a third reported any AI use, with disclosure rarer still — and non-disclosure read as rational caution under ambiguous policy rather than dishonesty. Those same students favored real-world, data-based tasks as the fix.([[kirsanov-beyond-detection-ai-online-assessments-2026]]) Expecting, declaring, and scrutinizing AI use beats policing it, because [[authentic-assessment|authenticity]] has to be designed rather than enforced.([[beyond-detection-authentic-assessment-ai-2025]]) **Make honest disclosure the safe option.** Vague or punitive policies drive concealment; specific declaration frameworks tied to cognitive stages, paired with an assurance that truthful disclosure is not penalized, get more information out of students than surveillance does.([[gonsalves-student-non-compliance-ai-declarations-2025]])([[chang-should-i-tell-my-teacher-ai-disclosure-2026]]) ## 5. The objections you will hear - **"Our vendor advertises a 1% false-positive rate."** Vanderbilt's licensed detector claimed exactly that, and when the university sought to validate the figure it could not — 1% of 75,000 annual submissions implied roughly 750 mislabelled students — so it disabled the tool. Independent benchmarks put no tool above 80% accuracy.([[teichmann-detecting-undetectable-misconduct-2026]]) - **"We only use it as one piece of evidence."** The standard is not "some evidence" but evidence meeting the balance of probabilities, and detector scores combined with markers, style comparisons, LLM judgments, or silence still do not satisfy it. If you would not say the reasoning out loud in a hearing, it is not evidence.([[bassett-ai-detectors-education-2026]]) - **"If we stop detecting, cheating wins."** Detection is not what stops the determined: 94% of injected AI submissions passed a live online exam unnoticed, humanizing defeats detection more than 96% of the time, and contract cheating already left no reliable trace.([[teichmann-detecting-undetectable-misconduct-2026]]) What detectors do catch, with measured reliability, is honest light editing and non-native prose.([[karr-ai-detection-humanization-2026]]) - **"We only use it privately, to start a conversation."** That is the least harmful use, and it still costs something: the students most likely to be flagged are the least likely to be misusing AI, so suspicion lands disproportionately on the wrong people.([[teichmann-detecting-undetectable-misconduct-2026]]) Grand Canyon's answer was to stop starting from suspicion at all.([[best-response-student-ai-dialog-2026]]) ## 6. If you write or revise policy, put these in it - **Detector output is never standalone evidence, never triggers an automatic consequence, and never grounds a finding.** State this in the policy itself, not in a memo.([[karr-ai-detection-humanization-2026]]) - **Students must be able to see the accusation and respond to it**, with the right to silence preserved and an oral verification route available for any allegation that rests on style alone.([[bassett-ai-detectors-education-2026]]) - **Procurement terms should cover data retention, use for [[pedagogical-llm-training|model training]], storage location, breach notification, and appeal rights** — the risk sits with your institution, not the vendor.([[bassett-ai-detectors-education-2026]]) - **Fund the alternative.** The documented faculty complaint is "suspicion without recourse," so the budget line that matters is verification capacity and guidance, not a detection license.([[best-response-student-ai-dialog-2026]]) - **Ask for error rates in writing, then try to validate them locally.** If the advertised rate cannot be reproduced on your own submissions — as Vanderbilt found — that is your answer.([[teichmann-detecting-undetectable-misconduct-2026]]) ## 7. The legal risk when a student is wrongly accused A detector score that becomes an accusation is where this stops being a teaching question. Institutions are not exposed because someone was accused, but because of how the accusation was built and handled, and the exposure usually surfaces first as an internal appeal or a regulator complaint rather than a lawsuit filed in court. - **The procedure is the first thing tested.** Principles of natural justice require that a student be informed of the allegation and given an opportunity to respond *before* any determination is made, obligations codified in regulatory standards as well as in sound academic integrity policy; the response opportunity is typically an investigative meeting or panel interview, and whatever the student says becomes part of the evidentiary record. A finding reached without that step is vulnerable regardless of whether the underlying suspicion was reasonable.([[munoz-misconduct-allegation-evidence-2026]]) - **The evidence standard is the second.** Misconduct findings require evidence meeting the balance of probabilities, and detector output does not get there on its own: the scores cannot be validated against ground truth in real submissions, the tools cannot be interrogated about how a given verdict was reached, and a student cannot cross-examine a number. If your case cannot be stated without the detector score, the case is weak on its face.([[bassett-ai-detectors-education-2026]]) - **Patterned error turns a technical defect into a fairness and equality problem.** Detector error is not random: the strongest evidence in the knowledge base shows flagging concentrated on non-native English writing and on one discipline's prose conventions over another's, with measured accuracy of 0.69 and 0.61 for two widely used commercial tools and both of them failing on hybrid human-AI text. A finding built on an instrument that misclassifies by language background is a finding that invites a discrimination argument.([[hadra-ai-detector-accuracy-efl-2026]])([[van-vlasselaer-ai-detector-reliability-2026]]) - **Blanket "AI use" bans can remove an accommodation.** Rules that do not separate transcription and OCR from generative drafting may criminalise the assistive tools students with conditions affecting motor control, handwriting legibility or typing accuracy rely on, several of which have been discontinued with AI transcription filling the gap. Over-inclusive policy is a legal exposure, not just an imprecise one.([[wright-transcription-not-generation-2026]]) - **The data is your liability.** Detectors store student work on third-party servers, sometimes overseas under weaker [[privacy]] protections, which puts breach, retention and onward commercial use inside your institution's risk register rather than the vendor's marketing.([[bassett-ai-detectors-education-2026]]) - **Vendor claims will not protect you.** A licensed detector advertising a 1% false-positive rate failed validation when the university tried to reproduce it, implying roughly 750 mislabelled students among 75,000 annual submissions; that institution disabled the tool. The claim in the contract does not transfer the risk in the hearing.([[teichmann-detecting-undetectable-misconduct-2026]]) What lowers the risk is procedural rather than technical: never treat detector output as standalone evidence or an automatic trigger; document the evidence standard your process applies; give notice and a genuine opportunity to respond on the record; offer an oral verification route when the case rests on style alone; put retention, training-use and breach terms in procurement; write policy scope so that [[assistive-technology|assistive]] and transcription tools are explicitly addressed; and keep the audit trail that shows all of it happened. The knowledge base's account of the legal exposure itself is in [[legal-issues-and-risks]], and it is honest about its limits: it documents procedures, evidence categories and instrument reliability, not litigated outcomes. ## How this page differs from the neighboring FAQs - [[reduce-ai-cheating|How Can I Reduce AI Cheating in My Course?]] is the instructor's design playbook: guardrailed tools, assessment redesign, verification, declarations, [[ai-literacy|AI literacy]]. This page answers the narrower prior question of whether detector output can be used at all. - [[redesign-assessment-ai-era]] covers assessment redesign in depth; this page only points to the redesign moves that specifically replace detection. - [[course-ai-policy]] covers writing and communicating a course policy; section 6 here covers the detection-specific clauses an [[educational-policy-ai|institutional policy]] needs. - [[legal-issues-and-risks]] is the concept page behind section 7, covering wrongful accusation, surveillance, accessibility and data-protection exposure together. ## The bottom line Do not build a finding on a detector score, and do not buy one expecting it to secure your assessments. The measured behavior of these tools is the opposite of their marketing: they catch honest, disclosed, lightly edited work and competent second-language writing, while deliberate evasion passes 96% of the time and wholly generated submissions passed a live exam system 94% of the time.([[karr-ai-detection-humanization-2026]])([[teichmann-detecting-undetectable-misconduct-2026]]) The integrity question you can actually answer is whether a student can demonstrate understanding of what they submitted — and that is a teaching capacity worth funding. --- ## [How Should I Use AI to Study and Learn Effectively?](https://edtechdev.github.io/aied/faqs/study-with-ai/) # How Should I Use AI to Study and Learn Effectively? It is 11 p.m., you have a problem set due in the morning and a paper you have not started, and the chat window is already open. Every one of these tools will answer almost anything you ask, instantly, in clean prose that sounds like someone who knows what they are talking about. That is exactly why using them well is a skill rather than a convenience. Here is the bottom line, and it is the one sentence worth carrying off this page: **the same model, used at a different point in your study process, produces opposite results.** A chat that ends your thinking before it starts makes the work feel faster and leaves less behind. The same chat, used to test you, to explain something you have already attempted, or to show you where your reasoning broke down, helps you learn. The second thing to hold on to: work that came out fine is not evidence that you learned anything. Work you produced with a model in the room shows what you and the model can do together; exams, placements and the next course in the sequence test what you can do alone — what the research calls [[learning-gains|learning]]. Those are different things, and the gap between them is the whole game. ## The short version: three rules you can use tonight **1. Attempt first, prompt second.** Write your own answer — an outline, a wrong guess, a messy attempt — before you open the chat. Then make the tool withhold the solution until you have one. This is the single move with the strongest evidence behind it, and it is below. **2. Ask for teaching, not production.** Hints, explanations, worked steps, quizzes and corrections: yes. Deciding what your essay argues, finding the source, writing the sentence, solving the problem: yours. If you could not do the step with the chat closed, you are not studying, you are transcribing. **3. Finish by closing the tool.** Answer from memory, out loud, with nothing open. If you cannot, you have not learned it yet — and you found that out tonight instead of in the exam hall. Everything below is detail on those three. ## When to use it All of these leave your thinking intact: - **Explaining something you have already read and failed to follow.** Ask for the concept as though you will be tested on it, then explain it back. - **Quizzing, one question at a time**, with instructions to hold the answer until you have tried. - **Showing a worked example of a problem type** so you can see its shape, then doing the next one yourself. - **Getting unstuck on a point you can name.** Not "help me with this assignment" but "I keep getting the wrong sign when I integrate by parts; here is my attempt; where does it break?" - **Making practice cards from what you are already reading**, and scheduling them. - **Explaining your own errors after feedback.** Ask what was wrong with your version, not just what the right answer is. The pattern: the tool is doing something *to* your learning rather than *instead* of it. The four-stage process the STEM students in [[viberg-efficiency-effectiveness-srl-llm-help-seeking-2026|Viberg, Feldman-Maggor and Wong (2026)]] described is worth copying — deciding whether help was needed, choosing whom to ask, deciding what kind of help to request, and judging what they got. Those 20 interview participants put the model first because it was the lowest-barrier option ("first ChatGPT, then classmates, and lastly teachers"), and they deliberately asked for hints, step-by-step guidance and concept explanations rather than direct solutions, treating the chatbot as "a hint, an assisting tool, but not the standalone solution". ## When not to use it, or when to use it differently The gap between what students intend and what they do is the most consistent finding in this research. [[yan-cognitive-outsourcing-genai-assessments-2026|Yan and colleagues (2026)]] interviewed 38 undergraduates in Japan and China while they walked through their own chat histories for unsupervised essay assessments. One group handed the task over entirely — teacher assigns, student passes it to GenAI, GenAI generates, student checks formatting, student submits — and one participant described the interaction as having "completely replaced my brain". That was the small group. Thirty-one of the 38 said their intention was to use AI as a learning assistant, and they still reported [[cognitive-offloading|overreliance]], mental complacency and fast forgetting; one said "the speed at which you forget it is also very fast". Yan and colleagues call this the efficiency paradox: convenience bought at the cost of the cognitive work that builds understanding. Only 8 of the 38 worked as what the authors call cognitive partners, and their defining feature was not that they used less AI. It was that their total effort did not fall — it moved. One shifted effort from searching to quality checking; another wrote short reflection notes after every AI session to counter the fading of instantly retrieved information. What that looks like as instructions to yourself: - **Do not paste the assignment and take the output.** In the same study, **76.32% of students relied on an ask–get answer–stop pattern**, typically pasting the assessment title without saying what they actually needed and then resubmitting the same prompt when the answer disappointed them. Sustained, iterative dialogue appeared in only 23.68%, almost all of them cognitive partners. - **Do not detach it from your own reading and writing.** 78.94% used AI either before starting or after drafting, and that detachment is the signature of the group that learned least. - **Do not stop at the first reply.** The [[help-seeking]] research shows the same shape. [[student-ai-inquiry-types-cs2-2026|Amoozadeh and Alipour (2026)]] classified 830 prompts from 72 students across two programming tasks and found the same few moves over and over: assertions (reports of confusion rather than questions), verification prompts asking whether something was right, and instrumental or procedural prompts asking for the next step. Comparison, prediction and feature specification — the moves that make you think — stayed rare. Wherever your own prompts cluster, you can predict your learning from it. - **Let it do logistics, not analysis.** Summarizing, translating, reformatting, generating practice cards at the point of reading and planning a study schedule are reasonable uses. [[ai-guided-learning-audiovideo-2026|Kawamura (2026)]] built systems that adapt spoken playback to the difficulty of each segment (averaging about 1.30x) and generate multimodal video summaries that cut viewing time by 53% with no statistically significant difference in quiz scores. Moving through material faster is fine. Skipping the work that changes you is not. ## How to study with it so it sticks **Attempt before you prompt.** This has the best evidence on the page. [[adaptive-pretesting-retention|Akgun and Toker (2026)]] gave 89 undergraduates the same adaptive pretesting session and the same instruction, then seven weeks of different practice. Adaptive spaced retrieval — the AI probing misconceptions, demanding elaboration of thin answers, advancing only on genuine conceptual engagement — produced the highest posttest scores (M = 78.19) and the highest practice effort (M = 0.85), against free chat with the model (M = 67.28 and 0.49; d = 0.92 on scores). Pretesting helps even when your first answers are wrong, because the attempt activates what you know and exposes what you do not. The advantage came from the agent's refusal to give direct solutions, a policy you can impose on any chatbot in a sentence: *explain this as though I will be tested on it, quiz me one question at a time, do not give me the answer until I have attempted it, tell me what I got wrong and why.* **Retrieve on a schedule instead of re-reading.** [[memdora-ai-spaced-repetition|Zhang (2026)]] describes Memdora, an AI spaced-repetition system built on the finding that roughly **70% of newly learned material is forgotten within 24 hours** without review. It generates cards from whatever you are reading, at the point of reading, and schedules them with FSRS-6. Be precise about how much that is worth: the paper reports better retention than traditional flashcard tools, but its contribution is a design and interaction taxonomy rather than a controlled retention trial. Take the spacing and retrieval as evidenced and the specific interaction design as promising. **Process corrections rather than swapping answers in.** [[verification-quality-reliance-calibration-genai-2026|Wei and Shang (2026)]] report an error-correction study in which effort during correction mattered for learning, while simple answer substitution was unlikely to deliver the same benefit. When something gets fixed for you, ask what was wrong, why your version failed and what the underlying rule is — then redo the step yourself. **Take your own notes.** Yan's cognitive partners counteracted fast forgetting by writing short reflection notes after each session. That is the cheapest version of the same habit, and it costs two minutes. ## How to check output before you trust it "Critical use" is a vague phrase, so it helps to know what research can separate. Wei and Shang separate seven targets: epistemic evaluation, whether you start checking, how well you checked, whether the check succeeded, what you did with the output, how you performed on the task, and what you learned independently. The point to hold on to is that **checking is not the same as successful checking** — a good process can end inconclusive, and a weak one can land on the right answer by accident. [[trust-calibration|Calibrated reliance]] means your decision to accept or reject an output matched that output's actual quality, which you cannot judge from how confident the answer sounded. Three findings make this concrete. Dávila et al. (2025) gave learners advice that was correct about half the time and found that how much they weighted it varied with their [[prior-knowledge|prior knowledge]] and gender — trust was not tracking accuracy. Zheng et al. (2025) identified a "failed application" pattern in which correctly following guidance that was itself correct still ended in a wrong answer. And [[self-efficacy]] is not a safe guide: Rheu and Cho (2025) found that understanding how language models work was associated with more self-reported fact-checking, while some forms of confidence and feature knowledge were associated with less of it. So verify anything you will be assessed on against a source you can name. The STEM students in the Viberg study checked outputs against their coursework and instructors, one explaining that "I only go to the TA if we can't tell whether ChatGPT is making things up". When you genuinely cannot tell, mark the point unresolved and ask a person rather than adopting it to get the page finished. And decide your own ground rules before you need them: [[ethical-conditions-llm-exam-preparation-2026|Pérez-Portabella and colleagues (2026)]] surveyed 151 undergraduates and found that ethical judgments were necessary conditions for intending to use a large language model for exam preparation, with consequentialist and deontological reasoning predicting that intention, and intention strongly predicting actual use. Deciding in advance what you will and will not do is not a formality; it is what your behavior follows. ## How to tell whether you are actually learning Do not judge by how the session felt or how polished the output looks. Independent learning means retention, [[transfer-of-learning|transfer]], unaided performance and finding your own errors once the AI support is withdrawn. Three tests follow, and none of them needs a researcher: - **Close the tool and answer.** If you cannot produce the explanation or the solution with the chat shut, you have not learned it yet. - **Test after a delay, not immediately.** The seven-week posttest is what separated the conditions in the statistics study; a score taken while the conversation is still on screen tells you very little. - **Explain it out loud, from memory, with no notes.** If you need the tool's phrasing, the understanding is not yours yet. If the app you use reports outcomes per item, read them — Memdora's classroom layer tracks learning at the individual card level, which tells you far more than a streak count or total minutes studied. ## "I'm paying for this — why not use it?" Because a subscription buys you something better than answers, and the answers are the cheap part. What you are paying for is a patient tutor that will explain the same idea five ways, quiz you until you can do it cold, and tell you what went wrong with your own attempt — a tutor most students could not have at 11 p.m. Getting your money's worth means asking for the expensive things, not the cheap one. The cheap one is paste-and-take, and it is the use that measurably costs you: the group in the Yan study that handed the task over entirely was the small group, while the students who reported forgetting fastest were the much larger group that believed they were using AI well. ## "Everyone else does it this way" Most of them are doing it in the way the studies show works badly: 76.32% ask, get an answer and stop; 78.94% never involve the tool in the part of the work that would have changed them. Being in that majority is easy and unremarkable. The 8 of 38 who worked as cognitive partners are not a lane reserved for the gifted — their distinguishing feature was a rule you can adopt tonight: effort moves rather than disappears. One moved effort from searching to quality checking. One wrote reflection notes after each session. ## "It's faster" It is, and that is the trap the research names the efficiency paradox: speed goes up while the thing that made the task worth setting goes down. The retrieval study is the cleanest version of the trade — free chat did less practice work (M = 0.49 effort against M = 0.85) and scored lower seven weeks later (M = 67.28 against M = 78.19, d = 0.92). Both conditions had the same AI and the same starting session; the difference was whether the tool pushed back. You will be faster tonight and slower in the exam, and you get to choose which one you want. ## The checklist - Write your own answer before you open a chat, even a bad one, and make the AI withhold the solution until you have attempted it. - Ask for hints, explanations, worked steps and quizzes rather than completed work. - Name the concept you are stuck on and paste your own attempt; do not paste the assignment title and accept the first reply. - Stay in the conversation: push back, ask why, ask what you got wrong — instead of resubmitting the same prompt. - Retrieve on a schedule, using cards drawn from your own reading, rather than re-reading. - Verify anything assessed against a named source, and treat fluency in the output as no evidence at all. - When a correction is handed to you, redo the step yourself instead of pasting the fix. - Keep your reading and drafting your own: use AI before brainstorming and after drafting, not in place of either. - Finish by closing the tool — unaided answer, delayed test, spoken explanation. If you also teach, one finding is worth carrying across. All 38 students in the Yan study reported that their instructors forbade copying and gave almost no concrete guidance, and that vacuum pushed even well-intentioned students toward outsourcing. Showing what sustained dialogue with a model looks like, setting tasks that require an attempt before the prompt, asking for process evidence such as notes on how AI was used, and marking unaided performance separately from assisted work all do something the prohibition alone does not. For the wider evidence base, see [[does-ai-help-students-learn]] and [[how-ai-impacts-students]]. --- ## [What Are the Top 10 Findings from AI in Education Research That Instructors Should Know About?](https://edtechdev.github.io/aied/faqs/top-10-findings-ai-education-instructors/) # What Are the Top 10 Findings from AI in Education Research That Instructors Should Know About? You are not being asked to become an AI researcher. You are being asked to make ordinary teaching decisions — what to allow on an assignment, what a grade is supposed to prove, what your students should be able to do without the tool — and you would like them to rest on more than opinion and vendor claims. The bottom line, across the research this knowledge base has collected: **how AI is built into the learning activity decides whether it strengthens thinking or replaces it.** It is not simply "use AI" or "ban AI," and the studies behind these findings are recent, often short-term, and tied to specific courses. Treat them as the strongest available guidance rather than settled law. ## The short version - Decide what thinking the task exists to build, then give AI a job that does not do that thinking for the student. - Judge success by what students can do **later, without the tool** — not by how good the work looks now. - Use AI as a tutor, coach, or critic far more than as an answer machine. - Treat a polished submission as weak evidence of learning, and collect some evidence of the process too. - Put an explicit AI rule on each major assignment and say **why** it is that rule. - Teach students to check, question, and push back on AI output. It does not develop from exposure alone. ## 1. Work that looks better with AI is not proof of better learning Students can produce stronger work and finish faster with [[generative-ai|generative AI]] while learning less on their own. [[cognitive-offloading|Cognitive offloading]] is the risk: it is hardest on learning when the AI performs the reasoning the student was supposed to practice. The distinction that matters is between *performance while assisted* and *learning demonstrated later without assistance*. The strongest anchor is a [[kumar-genai-computing-education-systematic-review-2026|systematic review of 72 peer-reviewed computing-education studies]]: generative AI reliably raised short-term completion and cut time on task in 36 studies — the best-replicated result in the whole corpus — yet those efficiency gains "do not transfer to independent performance" in 21 studies. A [[yan-cognitive-outsourcing-genai-assessments-2026|study of 38 undergraduates using think-aloud interviews]] found the same split in how they actually worked: 76.32% stayed in a single-turn ask–answer–stop pattern, and only 21.06% alternated AI use with independent reading and drafting. A [[critical-thinking-paradox-genai-learning-2026|2026 framework paper]] names the pattern a critical-thinking paradox: grades and products can rise while the mental work that produces durable learning falls. **In the classroom:** build in at least one task per unit where students retrieve, explain, solve, or defend ideas with no AI present, and grade that. ## 2. AI earns its place as a tutor, not an answer machine Decades of [[intelligent-tutoring|intelligent-tutoring research]] point the same direction: diagnose what the student understands, ask questions, give graduated hints, ask them to explain, give feedback — instead of handing over the solution. Current [[pedagogical-agent]] work draws the same line between *teaching behavior* and *answer production*. The [[thermomix-genai-education-analogy-2026|Thermomix kitchen-machine analogy]] makes it concrete: one appliance can either do the cooking for you, so you lose the skill, or act as a testing partner for ideas you still have to make sense of. Those modes line up with the [[icap-framework|ICAP framework]], which predicts different learning from passive, active, constructive, and interactive use. Newer work shows the same principle inside a narrow tool. A [[structrag-diagram-reasoning-ai-tutoring|tutor that reads engineering diagrams structurally]] reached 93.0% edge-level F1 against 89.3% whole-diagram accuracy, which means it can name the specific missing or misread connections even when the whole diagram is wrong — feedback a student can act on, rather than a pass or fail. **In the classroom:** tell students to ask for "one hint," "ask me questions," or "critique my reasoning," and make "solve this" the unusual case. ## 3. Do not smooth out the productive struggle [[productive-failure|Productive failure]] is the finding that learners often retain more when they attempt a problem before being shown how. Making learning frictionless can remove exactly the work that creates the learning. In the [[yan-cognitive-outsourcing-genai-assessments-2026|38-student interview study]], the largest group (n = 31) described mastery goals but worked in fragmented single turns, then reported overreliance, mental complacency, and fast forgetting — one participant put it as "the speed at which you forget it is also very fast." The [[thermomix-genai-education-analogy-2026|Thermomix analogy]] compresses the risk into five words: with a Thermomix you lose the ability to cook. **In the classroom:** use an **attempt → AI assistance → revision → reflection** sequence rather than opening the tool at the first second of every task. ## 4. Feedback only counts once a student does something with it [[ai-feedback-quality|AI feedback]] can be timely, specific, scalable, and acceptable to students, including in [[higher-ed|higher education]]. Whether it *teaches* depends on accuracy, [[pedagogy|pedagogical]] fit, and students' [[feedback-literacy|feedback literacy]] — their ability to judge feedback and act on it. A [[mcinnes-salvaging-constructive-alignment-genai-2026|critical analysis of 14 institutional guidance documents]] warns that generic prompts produce outcomes, activities, and assessments in isolation: because the tool cannot know how interconnected a topic is, its feedback stays general instead of diagnostically precise. Two 2026 studies show how much the inputs move the output. In a [[teacher-ai-literacy-prompt-feedback-quality-2026|study of AI feedback on learning goals]], the model alone explained 26.9% of the variation in feedback quality, and adding the prompt raised it to 42.8% — an extra 15.9%. The one prompt feature that mattered was subject-specific terminology; swapping it for everyday paraphrases made feedback significantly worse. Model choice mattered too: Claude 3 and Gemini Advanced produced significantly lower-rated feedback than ChatGPT-4. Meanwhile a [[llm-automated-grading-programming-comparison-2026|comparison of 18 language models grading 6,081 programming submissions]] found average grades from 0.290 to 0.608 depending on the model, with exact-agreement rates as low as 0.20 for some and 0.74 at best. **In the classroom:** have students weigh AI feedback against your rubric, decide what to accept or reject, and explain what they changed. ## 5. Instructional design matters more than which model you use A comparison of a theory-informed chatbot that scaffolded student explanations against ordinary ChatGPT and business-as-usual teaching found no significant immediate differences — but four weeks later, the scaffolded group retained more conceptual knowledge. It is one study, not a universal effect, and it is the clearest illustration that **design can outweigh model capability**. The [[kumar-genai-computing-education-systematic-review-2026|72-study review]] lands on the same design layer, calling critical engagement with AI output "the common mechanism linking every effective intervention in the corpus" and recommending [[scaffolding|graduated access]] — introducing generative AI only after foundational competence is shown. **In the classroom:** design AI activities around self-explanation, retrieval, comparison, argumentation, teaching, or critique rather than content generation. ## 6. Stop trying to catch AI; start producing evidence of learning AI detectors have well-documented reliability and fairness problems, and the deeper issue is [[assessment-validity|validity]]: a polished take-home product no longer shows that the person who submitted it has the competence. Detection-framed policy also chills legitimate use. In a [[zou-is-this-a-trap-student-teachers-genai-2026|mixed-methods study of 85 student teachers]], 62.4% declined to use generative AI even where it was permitted, 41.5% of those who declined cited fear of being accused of [[academic-integrity|plagiarism]], and 9 of 11 interviewees read the permissive policy itself as a trap. New evidence raises the stakes on automated judging. In a [[llm-grading-self-preference-bias-2026|study of 1,426 psychology dissertations spanning ten academic years]], all four AI graders scored student-written work lowest and AI-written work highest — 10 of 16 comparisons were large enough that differences of that size are rare in education research, up to the largest gap observed. The bias was strongest for fully AI-generated text, which means a grader can reward text for being machine-like even when its instructions say to judge content. **In the classroom:** assess process alongside product — drafts, reasoning, critiques, oral defenses, demonstrations, reflections. ## 7. One AI rule for every assignment will not hold A useful assessment framework distinguishes three cases: **restrict AI** when independent competence is what you are measuring, **scaffold AI** when bounded help does not compromise that competence, and **require AI** when skilled human–AI collaboration is itself the thing students must learn. The [[zou-is-this-a-trap-student-teachers-genai-2026|student-teacher study]] shows why the conditions have to be explicit and consistent: only 37.6% of students used permitted generative AI at all, their choices tracked program culture and [[assessment|assessment design]] more than any one course's permission, and their own disclosure declarations under-reported actual use in every course. **In the classroom:** state the AI condition for each major assessment, and explain why that assignment has that rule. ## 8. AI literacy is much more than writing good prompts Higher-education frameworks now treat [[ai-literacy|AI literacy]] as conceptual understanding, operational skill, [[critical-thinking|critical evaluation]], [[ethics|ethical]] judgment, and awareness of limits — not just [[prompt-engineering|prompt engineering]]. Students need to learn when to distrust the tool, verify claims, spot bias, recognize uncertainty, and stay responsible for conclusions. The [[kumar-genai-computing-education-systematic-review-2026|computing-education review]] is concrete about this: prompt engineering, output verification, and analyzing AI errors are teachable skills that do not develop through exposure, and verification is the first of its three design requirements. Two 2026 studies of educators show how far this is from automatic. Among [[science-educators-ai-literacy-postqualification-2026|science teachers who had already completed AI training]], average AI literacy was 16.7 out of 30, below the reference sample's 18.79, and it was unrelated to age, gender, years of service, or how much they had used AI. A [[ai-tpack-mathematics-teacher-education-2026|survey of 412 prospective mathematics teachers]] found readiness at an early stage: teaching beliefs scored highest (mean 5.24 on a 7-point scale) while technical AI knowledge scored lowest (4.23). **In the classroom:** give students deliberately imperfect AI output and grade their ability to verify, critique, improve, and contextualize it. ## 9. AI can widen gaps even when everyone has access The [[digital-divide|digital divide]] now runs along at least three lines: access to tools, skill in using them, and who actually gets a useful result. Prompting skill alone creates a "prompt privilege" where more experienced users get better output from the same system. The [[kumar-genai-computing-education-systematic-review-2026|72-study review]] separates two mechanisms: a **skill gap**, where students with stronger [[prior-knowledge|prior knowledge]] convert help into durable gains while under-prepared students substitute it for practice, and a **resource gap**, where reliable internet and paid access sustain better tool use across institutions. Only six studies in the corpus directly examined equity — which the authors treat as the problem. A [[co-learning-ai-agent-hidden-rules-2026|four-experiment study of learners discovering hidden rules with an AI agent's help]] found the aid cut moves needed by 33–52%, but the benefit concentrated in the weakest performers: stronger learners were largely unaffected. Help can narrow a gap, in other words, but only for the students who engage with it rather than the ones who already did not need it. **In the classroom:** do not let prior AI experience become a hidden prerequisite. Provide [[equity-in-ai-education|equitable]] access, worked examples, direct instruction, alternatives, and accommodations. ## 10. Human judgment is still the part that does not automate AI integration raises linked questions about bias, privacy, transparency, learner [[agency|autonomy]], accountability, and [[pedagogical-safety|pedagogical safety]]. Faculty therefore need [[teacher-ai-competency|pedagogical AI competence]] rather than technical familiarity alone; reviews of educator preparation describe it as pedagogical reasoning plus critical and ethical judgment. The [[mcinnes-salvaging-constructive-alignment-genai-2026|guidance-document analysis]] proposes a concrete design answer: a bounded, institutionally configured [[rag|retrieval-augmented]] agent that guides thinking without supplying answers, flags misalignment, and escalates to a person at the edges — authority that is "derivative and bounded" rather than autonomous. There is now direct evidence for keeping people in the loop instead of only in front of it. A [[instructional-agents-multi-agent-course-gen|course-material generation system]] scored better when people stayed involved: the mode with the most human input improved reviewer scores by 0.5–0.9 points over the fully autonomous mode. Its autonomous reviewers also behaved differently from human ones — the AI reviewers clustered tightly around 2.9–3.1 while human evaluators spread out and discriminated more — so the authors kept human judgment as the primary quality signal. **In the classroom:** keep consequential instructional and assessment decisions under meaningful [[human-in-the-loop-ai|human oversight]], especially where accuracy, fairness, privacy, or student progression is at stake. ## The pattern underneath all 10 **AI that replaces thinking → riskier for learning.** **AI that elicits thinking → potentially valuable for learning.** So instead of asking *"Should students use ChatGPT?"*, ask three better questions: **What thinking do I need students to practice? → What role should AI play without doing that thinking? → What evidence will show me the student learned it?** That leads to activities such as **attempt-before-AI, AI as a [[socratic-method|Socratic]] tutor, critique-the-AI, compare human and AI solutions, AI feedback plus student judgment, process [[eportfolio|portfolios]], and short oral defenses** — and away from the false choice between open use and blanket prohibition. ## What you can do this week Pick one assignment you are uneasy about and make two changes: state the AI rule explicitly with a one-line reason, and add a short in-class or recorded element that shows the student's reasoning with no AI present. That single pair usually settles the question of whether the assignment is measuring what you meant it to measure, and it costs you almost no class time. ## Objections you are likely to hear - **"My students say AI helps them."** It usually does help with the work in front of them; the finding is about what remains after. Ask what they can still do unaided and you get a different answer. - **"Detection tools are all we have."** They are unreliable, they misjudge legitimate work, and fear of them suppresses permitted use — 41.5% of the non-adopters in one study cited that fear. Process evidence is stronger and fairer. - **"I teach 200 students; I cannot read drafts."** You do not have to read everything. Short oral checks, in-class writing, and reflection notes on the AI interaction are cheaper than full draft review and far more diagnostic. - **"I do not teach AI; this is not my subject."** The findings here are about your subject: when AI does the practice your course exists to provide is exactly when it interferes. - **"Banning it is simpler."** Simpler, and it usually fails — adoption in one study tracked program culture rather than any single course's policy, and disclosure under-reported actual use in every course. For the stakeholder-by-stakeholder version of these misunderstandings, see [[addressing-common-misconceptions-ai-education|How Can We Address Common Misconceptions About AI in Education?]]; for the design requirements implied by findings 9 and 10, see [[equity-ethics-pedagogical-safety-research|How Should AI in Education Research Incorporate Equity, Accessibility, Privacy, Ethics, and Pedagogical Safety?]] and [[redesign-assessment-ai-era|How Should Assessment Be Redesigned for the AI Era?]]. One important caveat: the generative-AI evidence base is developing rapidly. Much of it consists of short interventions, [[self-report-measures|self-report]] studies, single disciplines, or emerging 2025–2026 work, and findings from mature [[intelligent-tutoring]] research are generally stronger than claims about unrestricted general-purpose [[conversational-ai|chatbots]]. Be especially skeptical of results that show only **student satisfaction, task speed, output quality, or immediate assisted performance** without measuring delayed or unassisted learning. --- ## [How Do I Teach Students to Verify AI Output?](https://edtechdev.github.io/aied/faqs/verify-ai-output/) # How Do I Teach Students to Verify AI Output? You have warned them. The chatbot is not a search engine, check the sources, use it responsibly — and the same fabricated citations and confident wrong answers keep arriving in the submissions. The warnings fail because they ask for a disposition. Careful students have it, hurried ones do not, and there is nothing in the sentence a student can act on at 11 p.m. with an output on screen that reads perfectly well. The research gathered here says that framing is too coarse to teach or grade. Checking is a set of analytically separate actions, and the broad label "responsible use" hides the difference between a student who attempted a check and one who settled whether the output was correct. The bottom line: verification becomes teachable the moment you stop treating it as a trait and start treating it as a sequence — a standard fixed before the check, a check that produces a record, and a decision the student must defend on domain grounds. That is a design problem you can solve this week — and a grading problem you can solve with artifacts the assignment already generates. This page addresses the mechanics of verification itself. For what students should understand about how these systems work, and how that understanding is defined, sequenced and assessed, see [[incorporating-ai-literacy]] and [[ai-literacy-evidence]]. A student can know a great deal about [[llm|large language models]] and verify nothing. ## What verification actually requires of a student [[verification-quality-reliance-calibration-genai-2026|Wei and Shang (2026)]] separate seven targets that "critical AI use" usually collapses into one: epistemic evaluation, verification initiation, process quality, verification success, reliance decisions, immediate task performance, and independent learning. Your syllabus rubric says "checked sources," which sits at step two and says nothing about whether the check was competent (process quality) or whether it changed anything (success, then the reliance decision). Initiation is not success: a strong process can end inconclusive, a weak one can land on the right answer, and a student who ran a search and stayed confused scores the same as one who resolved the question. The second distinction is between [[trust]] and reliance. The same review defines [[trust-calibration|reliance calibration]] as a judgment about whether a reliance decision was appropriate given the actual quality of the AI output — an output-contingent classification, not a stage on a timeline and not a score on a trust scale. One audited study found that false ChatGPT information shifted participants' reported trust, but never recorded whether they later accepted or rejected a specific recommendation. Another logged accept and reject decisions on ChatGPT [[feedback]] for 78 translation students, but without independent expert evaluation an acceptance cannot be called appropriate, nor a rejection justified. Without ground truth about output quality, reliance cannot be classified at all — which tells you what every verification task needs: a defensible answer key. That is why [[trust-calibration-chatbots-design-problem-2026|Jaidka and Cai (2026)]] treat miscalibration as a design problem rather than a student deficit. Their typology crosses ability to verify with motivation to verify, and the profiles need different remedies: the high-ability, low-motivation user is prone to complacent overtrust, while the low-ability, low-motivation user is the most exposed — so two students failing the same check may need opposite remedies. ## How to teach it Wei and Shang's synthesis drew on 493 deduplicated records from Web of Science Core searches and mapped 14 priority empirical studies onto verification, reliance and outcome columns. Its constructive output is the sequence to build tasks around. Adjudicate output quality first, capture whether verification was initiated, code process quality, score whether it succeeded, record the accept–revise–reject decision, and evaluate that decision against the adjudicated quality. One warning belongs here too: interventions that reduce inappropriate acceptance must also be checked for the unintended rejection of correct assistance, so leave room for a student to accept good output for a stated reason. **Predict before the output, then compare.** [[ai-writes-code-student-writes-model-2026|Gousopoulos (2026)]] builds a measurement program for construction tasks around *predict before you run*: the learner states what should happen, then the artifact's behavior is judged against that prediction on domain grounds rather than by whether it runs. Its audit of 24 studies found the same tool producing opposite outcomes under different task structures: a conventional ChatGPT setup ended significantly lower in achievement, [[self-efficacy]] and flow, while a condition adding verification requirements and error-reflection modules showed stronger [[critical-thinking|higher-order thinking]]. Because the prediction precedes the output, the reasoning is on record — and the record is what you grade. **Require source checks against the record, with a named target.** [[citation-errors-hallucinations-computing-education-2026|Denny et al. (2026)]] traced 113,588 references from 5,225 computing education papers published since 2021 and manually verified 828 suspicious records, finding 30 containing verifiably fabricated bibliographic information across 14 papers, all from 2025 and 2026. Thirteen were entirely fabricated; the other 17 combined a real title with fabricated or incorrect authorship, venue or year. At the SIGCSE Technical Symposium the count rose from 3 in the 2025 proceedings to 17 in 2026, or 2.3% of 2026 proceedings papers. Author fields fail more often than any other part of a generated reference, so "open the source and read the author list" targets the field most likely to be wrong. Do not ask students to "make sure sources are real"; ask them to confirm authors, venue and year against the record itself. The paper asks that every cited work be verified and that any checker stay [[human-in-the-loop-ai|human-in-the-loop]] — the practice a citation-verification requirement rehearses. **Teach the discrimination, not only the caution.** Gousopoulos formalizes verification as a signal-detection problem with two independent parameters: sensitivity, the ability to discriminate correct from flawed output, and criterion, where the learner sets the threshold for rejection. Over-reliance splits accordingly — warnings, checklists and hallucination prompts shift the criterion, changing when a student rejects, while domain instruction, worked comparisons and seeded-error practice raise sensitivity, because the learner can then tell. The prerequisite follows: verification instruction is educative only where sensitivity can exceed zero, which requires enough domain knowledge to distinguish correct from incorrect output. In the same audit, novices asked to explain [[llm]]-generated code succeeded on roughly a third of tasks, which is why judging AI feedback or code against one's own reasoning is a real check only where that reasoning has substance. If your students cannot yet do the domain thinking unaided, more warnings buy you nothing; teach the content first and attach the check to it. **Forewarn, immediately before the task.** [[chatgpt-inoculation-training-verification-2026|Vu, Cummings and Park (2026)]] showed a generic inoculation (forewarning) message immediately before two tasks to 100 US-based students, 40 domestic and 60 international EFL. Inoculated students were significantly more likely to verify the academic-source-summary task (M = 0.34 versus 0.18), while self-reported verification intentions did not move. Two lessons: a short warning at the point of use changes enacted behavior, and the effect was task-dependent — appearing for the source-summary task but not uniformly for a mathematics quiz on exponentiation and large-number multiplication. **Make correction cost something.** In an error-correction paradigm the review audited, effort during correction mattered for learning; simple answer substitution is unlikely to deliver the same benefit. Requiring a student to reproduce a step, rewrite a passage, or state the domain reason an output is wrong converts a check into work. [[tripartite-feedback-framework-ai-assessment-2026|Venetsanos (2026)]] sets the bar for what may be checked mechanically: documented criteria, comparison against established knowledge without interpretive judgment, and a single correct answer or pre-specified acceptable alternatives. Anything interpretive fails that bar, so separate the mechanical layer — dates, formulas, citations, calculations — from the judgment layer, where the student's own reading is the instrument of the check. ## How to assess it without policing students Gousopoulos draws the consequence directly: if the AI can produce the artifact, the artifact cannot be the assessment, so evaluation relocates to the specification, the validation reasoning and the interpretation. Its model authorship construct has four facets — specification, conceptual model, verification, interpretation — at four ordered levels from delegated to authored, where the authored level requires a verifiable specification preceding the first prompt, rejection of output on domain grounds with a stated reason, and interpretation beyond the artifact's own report of itself. Because the rubric is scored from materials a construction task already generates, the same instrument serves as [[formative-assessment|formative assessment]] and as a research measure. That is the answer to the busywork worry: you are not adding an assignment, you are scoring a byproduct. Three decisions keep this from becoming surveillance. First, grade the check against adjudicated quality, not effort — otherwise you build an incentive to perform checking theater. Second, state provenance, as Venetsanos requires: students must understand the provenance, nature and limitations of the feedback they receive — which parts were machine-verified, which evaluatively judged, and that human judgment has primacy — because students cannot weigh feedback they cannot situate. Third, keep enforcement off detection tools. [[bassett-ai-detectors-education-2026|Bassett et al. (2026)]] argue that AI detection should not be used in education at all: its estimates are probabilistic and cannot be independently verified because real-world text origin is unknown; its scores do not meet the balance-of-probabilities standard integrity investigations require; and the human-versus-AI dichotomy is meaningless for work created with, rather than by, AI. Their conclusion is that detection "does not safeguard academic integrity; it undermines it" — surveillance regimes foster suspicion and erode student [[trust]]. Keep the two senses of detection apart: [[ai-detection|detecting AI-generated text]] has no defensible evidentiary role here, while teaching students to detect errors in AI output is the point of the exercise. Separation is architectural, and worth telling students about: [[vetting-dual-llm-safety-education|Li, Zhang and Botelho (2026)]] check output with a second model rather than embedding the check in the generator — and a student who sees why can see why "the tool said it was right" is not a verification argument. ## Why students skip the check **Calibration.** Jaidka and Cai's typology puts the high-ability, low-motivation student at risk of complacent overtrust while the low-ability, low-motivation student is the most exposed. The audited literature shows the same split in miniature: a design that supplied advice correct about half the time, where the weight students gave it varied with [[prior-knowledge|prior knowledge]]. Calibration failures are not fixed by exhortation: the student's own confidence signal is what is miscalibrated. **Fluency.** Output that reads well is treated as right. Novices asked to explain [[llm]]-generated code succeeded on roughly a third of tasks, and a field study of student–ChatGPT quiz conversations found that following correct guidance still produced a wrong answer. When sensitivity is near zero, skipping the check costs the student nothing they can perceive — the failure is invisible until it is graded. **Social proof and intention.** Students report that they intend to verify, and the report is worthless: in Vu, Cummings and Park's study, self-reported verification intentions did not move while enacted behavior did (M = 0.34 versus 0.18). Do not grade the intention. Norm-based appeals are not a reliable fallback either — message-based norm nudges showed no significant effect in a direct tournament Jaidka and Cai report. Time pressure is a documented reason students skip checks, and a syllabus that leaves verification unscheduled says it is optional. ## Three objections **"They can just do it for the grade."** Partly true — so grade materials the AI cannot produce on a student's behalf. A specification written before the first prompt, a prediction recorded before running the artifact, a source list checked against the record, a stated domain reason for rejection: each of these requires the student to hold a position the generator cannot supply. At the authored level, Gousopoulos requires rejection of output on domain grounds *with a stated reason* — a fabricated reason is as visible as a fabricated citation. The residual risk is bounded by the sensitivity threshold: where a student has no domain knowledge to check against, verification is theater, and no rubric rescues it. That is an argument for teaching the domain first, not for abandoning the check. **"There is no time in the syllabus."** Forewarning is one message immediately before the task; the seeded-error task is twenty items, eight with a domain-level error, and it tells you whether students can discriminate at all; and the rubric rides on materials the task already generates, so it costs marking time on evidence you would otherwise lack rather than a new assignment. What costs time is the alternative: unverified submissions, fabricated references, and the integrity conversations that follow. The constraint is still genuine — time pressure is a documented reason students skip checks, so verification that is not scheduled into the task will not happen. **"The tool is usually right."** Then the check is cheap and the exceptions are the point. The record on references is not reassuring: 30 verifiably fabricated items across 14 papers, all from 2025 and 2026, 17 of them pairing a real title with fabricated or incorrect authorship, venue or year, 13 entirely fabricated, and a jump from 3 in the 2025 SIGCSE proceedings to 17 in 2026, or 2.3% of 2026 proceedings papers. Accuracy on the parts you happen to notice says nothing about the parts you do not — author fields fail more often than any other part of a generated reference. And an accepted answer can be wrong even when the guidance was right: the field study of student–ChatGPT quiz conversations recorded exactly that. The reason to teach verification is not that the tool is usually wrong, but that the student cannot tell which case they are in — precisely the skill your course exists to build. ## What remains unknown - Whether any intervention improves verification success and the reliance decision that follows, judged against adjudicated output quality: no study among Wei and Shang's 14 priority cases measured both, and few followed immediate performance with delayed retention or [[transfer-of-learning|transfer]]. - Whether the components relate in the order the map implies. The seven targets are an analytic ordering, not a validated causal model. - Whether design propositions work. Jaidka and Cai's eight propositions are untested predictions, and message-based norm nudges showed no significant effect in a direct tournament they report. - Whether checking built into feedback develops self-verification habits or dependency on external validation, which Venetsanos raises and leaves open. - Whether verification instruction works below a domain-knowledge threshold. Gousopoulos predicts it cannot, and notes that some domains furnish an external criterion — physics, chemistry, ecology, epidemiology — while history, literature and [[ethics]] largely do not. ## What to do this week - **Fix the standard before the check.** Decide what adjudicated correctness means for the task so a check has something to resolve against. - **Ask for the prediction first.** Have students commit in writing to what should happen, then compare the output on domain grounds, not fluency. - **Require citation verification with a named target.** Author lists are the least reliable part of a generated reference; have students open the record and confirm authors, venue and year. - **Diagnose sensitivity and criterion separately.** A student who accepts flawed output because they cannot tell needs domain practice; one who can tell and accepts anyway needs the threshold moved. - **Grade the checking, not only the artifact.** Score specifications, prediction records, validation logs, source checks and stated reasons for rejection. - **State provenance.** Tell students which feedback was machine-verified and which was judged by a person. - **Keep enforcement off detectors, and schedule verification.** Time pressure is a documented reason students skip checks. - **Run the seeded-error task.** Twenty items, eight with a domain-level error, tells you whether students can discriminate at all. - **Say the warning at the point of use.** Forewarn per task type, not once per syllabus. For the surrounding work, see [[incorporating-ai-literacy|how to incorporate AI literacy into a course]] and [[ai-literacy-evidence|what the AI literacy evidence shows]] — the two pages that build the understanding this one puts to work. --- ## [What Are Best Practices for Writing Instruction in the Context of AI?](https://edtechdev.github.io/aied/faqs/writing-instruction-ai-best-practices/) This FAQ is written for instructors who have to decide, course by course and assignment by assignment, what AI should be allowed to do in student writing. It draws on the [[ai-education|AI in Education]] knowledge base, especially its syntheses of [[writing-education|AI in writing education]], [[cognitive-offloading|Cognitive Offloading]], [[ai-feedback-quality|AI Feedback Quality]], [[feedback-literacy|Feedback Literacy]], and [[academic-integrity|Academic Integrity]]. The research points to one consistent principle: > **Use AI to increase feedback, reflection, critique, and revision — not to remove the intellectual work that writing is meant to develop.** That principle matters because [[writing-education|writing]] is not the production of polished text but a cognitive, rhetorical, and social process. AI can improve the product while weakening the processes of reasoning, authorship, source evaluation, and revision that your assignment exists to teach. Everything below is a way of applying that distinction. ## Eight research-informed practices for instructors **1. Decide first what intellectual work students must retain, then set the AI boundary.** Before deciding whether AI is "allowed," name the construct the assignment develops or assesses. If the goal is argumentation, students must formulate and defend claims. If it is disciplinary interpretation, they must make interpretive judgments. If it is scientific reasoning, they must connect evidence to conclusions. AI use that replaces that work undermines the [[assessment-validity|validity]] of the assignment even when the resulting prose is excellent — the product no longer warrants the inference you want to draw from it. This is the central logic of [[writing-education|writing education]] and [[academic-integrity|academic integrity]], and it is why task-level rules beat course-level bans. [[nash-preservice-teachers-classroom-ai-policies-2026|Nash and Burriss]] show how hard that boundary is to draw well. Across 27 [[teacher-education|preservice]] English teachers' classroom [[educational-policy-ai|AI policies]], 22 permitted AI for ideation and brainstorming, 22 disallowed or left unclear the use of AI to compose sentences, paragraphs or papers, and 22 said nothing at all about *reading*. The contradiction the authors surface is instructive: the same teachers called ideation and revision "thinking" while locating real thinking only in final written text. If thinking is distributed across planning, drafting, revising and evaluating — as the composition research the paper draws on argues — then a policy that guards only the final product is guarding the wrong step, and reading deserves the same explicit treatment as writing. **2. Prefer "student thinks, AI responds" over "AI writes, student edits."** This is the most actionable principle in the evidence base. [[layer-sensitive-cognitive-offloading-writing-2026|Chen's layer-sensitive study]] of 168 undergraduates across six intact classes distinguished surface, structural, idea, and reasoning offloading: open AI collaboration produced the strongest AI-supported writing but the weakest later independent performance, and delegating *reasoning* was most negatively associated with independent [[critical-thinking|higher-order thinking]]. A bounded condition that restricted delegation and required students to explain how they accepted, modified, or rejected AI suggestions did better on independent writing, argument depth, and revision. The study is quasi-experimental with only six intact classes, so treat the causal claim cautiously — but the pattern aligns with the wider [[cognitive-offloading|offloading]] literature. A practical rule follows: require students to establish their own interpretation, hypothesis, argument, or analysis *before* asking AI to critique or develop it. AI can be asked, "Here are three objections to my argument." It should rarely be asked, "Write my argument for me." **3. Be especially cautious about AI-generated first drafts.** In a study of 253 writers, any AI assistance reduced perceived ownership, but *drafting* assistance reduced it most while *planning* assistance reduced it least; more AI-contributed text and ideas went with better essay quality but lower ownership ([[ai-writing-support-stage-ownership-2026|the planning-to-revision ownership study]]). The trade-off is real: you can get a better essay and a weaker sense of authorship in the same submission. If preserving authorship matters in your course, separate AI that prompts planning from AI that generates paragraphs — the first is defensible almost anywhere, the second rarely is when composing is the target skill. **4. Use AI as one feedback partner, and keep [[feedback|human feedback]] in the system.** The strongest classroom evidence here is [[pairr-ai-peer-review-2025|PAIRR]], involving 654 students across ten writing courses and three writing-intensive [[stem-education|STEM]] courses: students drafted, completed [[peer-assessment|peer assessment]], obtained rubric-based AI feedback, compared and evaluated both sources, made revision plans, revised, and reflected. Fifty-eight percent preferred combined peer + AI feedback against only 6 percent who preferred AI feedback alone. Students found AI useful for broad rubric-oriented revision advice while peers supplied contextual knowledge and authentic audience response. Be precise about what this shows: PAIRR measured students' experience and evaluation of feedback rather than long-term [[learning-gains|learning gains]], so it is strong evidence for a *feedback design* and weaker evidence for durable improvement in writing. **5. Teach students to evaluate feedback, not just to prompt for it.** Access to good feedback is not enough. Students need [[feedback-literacy|feedback literacy]]: the capacity to seek feedback, judge its quality, manage their reaction to it, and turn it into revision. Students with higher feedback literacy benefit far more from AI feedback, while those who treat it as authoritative gain less ([[ai-feedback-quality|AI Feedback Quality]]). A useful assignment component is a short feedback decision table — *AI suggestion → accept / reject / modify → why → the rubric criterion or evidence supporting the decision*. That turns AI output into material for [[evaluative-judgment|judgment]] rather than instructions to obey, and it gives you something gradeable that is not the prose itself. **6. Keep some writing and reasoning independently observable.** AI-assisted performance is not independent capability. Retain occasional withdrawal conditions: brief no-AI writing, in-class interpretation, oral explanation, a conference, spontaneous revision, or a follow-up problem that requires transferring the same reasoning to a new case. This matters most when the final paper carries substantial grade weight. It does **not** mean converting every assignment into a proctored exam — a few strategically placed independent samples give both you and the student a baseline to interpret the assisted work against ([[cognitive-offloading|Cognitive Offloading]]; [[layer-sensitive-cognitive-offloading-writing-2026|Chen's study]]). This is the writing-specific case of the general assessment-redesign argument in [[redesign-assessment-ai-era]]. **7. Use disclosure as reflection, not as a trap.** Disclosure is not neutral: students conceal AI use when policies are ambiguous, punitive, or stigmatizing, and honesty can even attract suspicion. Disclosure works pedagogically when it makes decision-making visible — what AI was used for, what was supplied to it, which suggestions mattered, and what the student ultimately accepted or rejected ([[ai-use-disclosure|AI Use and Disclosure Statements]]; [[student-rationalization-ai-writing|student rationalization research]]). Task-specific guidance is better than a blanket "AI permitted" or "AI prohibited" rule, because the right boundary differs between brainstorming, argument development, sentence editing, source work, and final composition. State how disclosure will affect grading, or students will assume the worst. The assessment-design modeling of [[mohamed-temimi-assessment-imperfect-information-disclosure-2026|Mohamed and Temimi]] supplies the mechanism: disclosure becomes the attractive option only when the cost of honesty stays low, and because a detector's false positives fall on honest students too, heavier monitoring can make concealment relatively *more* attractive. Their advice is to design for the student most tempted to conceal and to read a declared use as context rather than a confession. **8. Do not treat generic [[llm]] judgment as a substitute for your judgment, especially in [[summative-assessment|summative assessment]].** Two apparently conflicting results are compatible: carefully calibrated systems with detailed rubrics and examples can score particular tasks well, while out-of-the-box LLM grading diverges substantially from human judgment. [[llms-do-not-grade-essays-like-humans-2026|Mathew et al.]] found weak human–LLM agreement that varies systematically with essay quality — LLMs over-reward short, superficially readable essays and under-reward longer, stronger essays with minor surface errors, clustering toward the middle of the scale. AI is therefore much easier to justify for low-stakes [[formative-assessment|formative]] feedback, comment drafting, or triage than as an autonomous final grader ([[automated-essay-scoring|Automated Essay Scoring]]). Students draw a version of this line themselves. In an undergraduate technical communication course where [[generative-ai|ChatGPT]] scored handwritten writing and students were told so, all 13 participants found the feedback clear and useful for surface-level revision, yet most separated *feedback utility* from *evaluative authority* — accepting the critique while insisting that the instructor decide the grade ([[student-perspectives-ai-writing-grading-2026]]). That two-judgment pattern argues for designing [[human-in-the-loop-ai|human oversight]] in as the point at which AI output becomes a grade, rather than treating it as an optional courtesy. ## A default workflow you can adopt or adapt A robust AI-era writing sequence keeps the reasoning with the student and puts AI in a consulting role: | Stage | Student responsibility | Appropriate AI role | What you can collect | |---|---|---|---| | **1. Encounter evidence** | Read, observe, annotate, calculate, run the experiment | Usually none, or clarification only | Notes, annotations, observations | | **2. Form an initial position** | Generate interpretation, question, hypothesis, claim | May ask questions or challenge assumptions | Claim or hypothesis memo | | **3. Plan** | Decide evidence, sequence, audience, genre | May critique an outline or suggest alternatives; student decides | Outline + rationale | | **4. Draft** | Produce substantive prose and reasoning | Restricted according to the learning goal | Draft and version history | | **5. Human response** | Give and receive peer or instructor feedback | None necessary | Peer comments | | **6. AI feedback** | Ask for criterion-referenced critique | Critic, reader, counterargument generator, clarity checker | AI feedback transcript | | **7. Evaluate feedback** | Compare peer, instructor, and AI suggestions; accept or reject with reasons | AI output becomes the object of evaluation | Feedback decision memo | | **8. Revise** | Make and justify substantive changes | Test clarity, offer alternatives | Revision | | **9. Verify** | Check every factual or source-dependent claim | AI cannot verify its own output | Sources and evidence | | **10. Reflect and disclose** | Explain AI's role and what changed in their thinking | None | Brief process reflection | | **11. Independent check when needed** | Explain or apply the reasoning without AI | None | Oral defense, quick write, new case | This is essentially [[pairr-ai-peer-review-2025|PAIRR]] extended: draft → human feedback → AI feedback → critical comparison → revision → reflection, with an independent baseline and clearer boundaries around reasoning. ## Discipline-specific guidance ### Composition and writing courses Composition instructors have the strongest reason to protect the writing process itself, because drafting, rhetorical decision-making, revision, audience awareness, and the development of voice are not just ways of displaying learning — they *are* the learning objectives. Teach AI use progressively rather than as a binary. Early in the term, collect several independent samples so both you and the student know what the writer can currently do. Then introduce AI mainly as audience, critic, and revision partner: ask it to identify where a reader loses the thread, generate objections to a thesis, compare a draft against the rubric, flag unsupported claims, or explain why a paragraph feels incoherent. Students judge the suggestions themselves. Keep peer review even when AI feedback is available: peers supply contextual and audience knowledge that AI does not, and evaluating the difference between the two is itself writerly training. Be conservative about AI generating long stretches of a first draft when learning to compose is the objective, as the ownership and offloading evidence above both indicate. At the same time, avoid blanket bans on grammar, phrasing, or language assistance — they create unnecessary barriers for [[multilingual-learning|multilingual writers]]. The better move is to separate *language support* from *intellectual authorship*: students may get help with expression while remaining responsible for ideas, evidence, rhetorical decisions, and meaning. Be aware that AI feedback is not language-neutral: [[marked-pedagogies-linguistic-bias-writing-feedback|Tan et al.]] found identical writing drew more praise and less substantive critique when student demographic or educational attributes were included in the prompt, so [[bias-mitigation|bias]] auditing belongs in your feedback design. ### Humanities and social science courses Interpretation is frequently the target capability here, so make **source encounter precede AI encounter**. Students should annotate the primary text, historical source, artwork, archival item, interview, or theoretical passage and formulate an initial interpretation *before* consulting AI. Afterward, AI becomes educationally useful as a foil: "Offer an alternative reading"; "What evidence would challenge my interpretation?"; "Which assumptions does my argument make?"; "Generate an interpretation from a contrasting theoretical perspective." The student decides which reading the actual text warrants. Assignments can also make AI critique itself part of disciplinary learning: give students an AI interpretation of a poem, event, argument, or social phenomenon and ask what it notices, what it misses, which evidence supports its claims, whose perspective is absent, and where it collapses ambiguity into a smooth answer. That preserves the interpretive judgment the [[humanities-education|humanities]] teach instead of treating an LLM as an oracle. Assess interpretive decisions and evidential justification rather than polish — a short conference question such as "Why did you read this passage this way rather than the alternative you rejected?" is far stronger evidence of understanding than guessing whether a sentence "looks AI-generated" ([[ai-detection|AI detection]]). ### STEM courses and lab reports **A qualification first:** the knowledge base holds much stronger evidence about AI-supported academic writing generally than about AI in laboratory-report writing specifically; hands-on and laboratory [[pedagogy]] remain under-covered, including within the [[engineering-education|engineering education]] literature. The guidance below is a reasoned application of writing, offloading, validity, and engineering evidence rather than a conclusion from a large literature on lab reports. Draw the AI boundary around **scientific reasoning rather than prose as a whole**. Students remain responsible for observations, raw data, calculations, uncertainty, figures, analysis choices, results, and claim–evidence reasoning. Once those exist, AI can reasonably help with organization, readability, transitions, and disciplinary conventions — the same surface-versus-reasoning distinction as above ([[layer-sensitive-cognitive-offloading-writing-2026|Chen's study]]). A strong AI-era lab-report package therefore includes the report **plus** selected raw data, calculations or notebook output, a figure with a student-written interpretation, and a brief statement of how the central conclusion follows from the evidence. For high-stakes reports, ask one or two individualized follow-up questions about a graph, an anomalous result, a [[research-methods-aied|methodological]] choice, or a limitation. Do not let writing fluency stand in for conceptual understanding: the knowledge base reports that automated [[physics-education|physics]] scoring can underestimate conceptual understanding when linguistic expression is weaker, which falls hardest on [[multilingual-learning|multilingual students]] ([[automated-essay-scoring|Automated Essay Scoring]]). If scientific reasoning and scientific writing are both objectives, score them separately. Engineering courses add professional accountability. A review of [[ethical-use-ai-engineering-education-review-2026|99 empirical engineering-education studies]] identified transparency, [[human-in-the-loop-ai|human oversight]], student independence, privacy, authorship, fairness, and beneficence as recurring [[ethics|ethical]] concerns, and argued that AI ethics belongs in professional formation rather than rule compliance alone. In a lab or design report, ask students not only to disclose AI assistance but to certify that they have verified calculations, sources, assumptions, and safety-relevant claims. ## What to stop doing - **Requiring polished prose while ignoring the process.** It makes the product progressively harder to interpret as evidence of learning. - **Leaning on AI detectors.** [[ai-detection|Detection]] addresses detection, not learning or [[assessment-validity|assessment validity]], and its evidentiary record is poor: in a covert field study 94% of AI-generated exam submissions went undetected and outscored real students, while a controlled study found detectors flagged compliant light AI editing at 38–80% yet missed more than 96% of humanized rewrites ([[teichmann-detecting-undetectable-misconduct-2026]]; [[karr-ai-detection-humanization-2026]]). - **Giving unlimited AI access without teaching feedback evaluation.** That assumes the [[feedback-literacy|feedback literacy]] many students have not yet developed. - **Substituting AI feedback for peer and instructor interaction.** It removes the contextual and relational information students consistently report valuing. - **Writing "AI allowed" and leaving it there.** Students reasonably read that as covering everything from spellchecking to generating the central argument. Task-level expectations are clearer and make disclosure less of a guessing game ([[ai-use-disclosure|AI Use and Disclosure Statements]]). ## A quick rule for deciding what AI may do For each task, ask three questions: | Question | If the answer is yes | |---|---| | Is this cognitive activity itself a learning objective? | Keep substantial responsibility with the student. | | Will students need to perform this capability independently later? | Add independent practice and some no-AI assessment. | | Can AI help students evaluate, practice, or revise the capability without performing it for them? | Usually the strongest case for AI integration. | Applied: if the goal is argumentation, AI may challenge an argument but should not routinely supply one. If the goal is historical interpretation, AI may offer a competing reading the student critiques. If the goal is lab-report communication, AI may improve prose after the student has done the analysis and reasoning. This is the [[coach-not-crutch-ai-writing|coach-over-crutch]] boundary in practice. ## How strong is the evidence? Promising, but not yet strong enough to justify a single universal AI-writing policy. Some of the best evidence comes from substantial authentic classroom studies such as [[pairr-ai-peer-review-2025|PAIRR]] with its 654 students. Other findings rest on quasi-experiments, small experimental samples, conceptual analyses, and emerging 2025–2026 work: [[layer-sensitive-cognitive-offloading-writing-2026|Chen's bounded-writing study]] involved 168 students but only six intact classes; the [[ai-writing-support-stage-ownership-2026|ownership study]] used a short experimental writing task; and a study reporting immediate quality gains after ChatGPT practice involved only 21 first-year international students and did not establish long-term [[transfer-of-learning|transfer]] ([[chatgpt-academic-writing-quality-ownership-2026]]). Treat the consensus as **design principles rather than settled prescriptions**. The most reliable cross-study finding is to preserve [[agency|student agency]], [[evaluative-judgment|evaluative judgment]], independent competence, disciplinary reasoning, and human feedback, while using AI to expand the availability of critique, practice, revision support, and linguistic assistance. **The short version for a syllabus or faculty workshop:** students do the intellectual work first; AI mainly questions, critiques, explains, and supports revision; students evaluate rather than merely implement its output; and assessment includes enough process or independent evidence to show what the student can actually do. ---