Concept
Pedagogical Patterns
Pedagogical Patterns — the ordered sequences of activity that research in this knowledge base has tested, with attention to where generative AI enters the sequence and where human judgment has to remain. Where Pedagogies and Teaching Strategies catalogs approaches (active learning, problem-based learning, collaborative learning) and Learning Design describes how a course is designed, this page catalogs what students and teachers actually do, in what order, and what happened when it was tried. PAIRR — draft, peer review, AI review, reflection, revision — is the best-documented example, and the pattern behind it recurs across disciplines: effort first, AI second, human judgment at the stakes.
Questions to Consider
- A pattern is a sequence, not a tool. Take one assignment you teach and write the order of moves a student makes. Where in that order would AI help, and where would it do the work the student is supposed to be doing?
- Several patterns here deliberately give AI a weak role — hints instead of answers, questions instead of corrections. Why would a deliberately less helpful tutor produce better learning, and what does that imply about the AI tools your institution is buying?
- The best-evidenced pattern on this page (learning by teaching an AI) asks students to explain, and it improves explanation and question quality but not objective recall. If you adopted it, what would you change about how you assess?
- When AI feedback was higher quality than teacher feedback, students did not revise more. What does that suggest about the difference between producing feedback and getting students to use it?
- Contexts change the answer: some patterns were tested online and asynchronous, others face-to-face with a lab. Which of these could you run in your own setting without new tools, and which would need infrastructure you do not have?
- Nearly every pattern here keeps a human at the point of judgment — grading, verifying, or interpreting. Is that a design choice, an evidence-based necessity, or a limitation of what has been tested so far?
Introduction
Pedagogy answers how should we teach; this page answers a narrower, more operational question: in what order should the moves happen, and where does AI belong in that order? The distinction matters because the same tool produces opposite outcomes depending on its position in a sequence. A generative-AI assistant placed before a student attempts a problem reliably depresses later unassisted performance; placed after an attempt, with hints rather than answers, the same class of system removes that harm.
Every pattern below is reported with an evidence status, because the knowledge base's coverage is uneven and the difference matters to anyone deciding what to adopt:
- Tested — at least one article reports a controlled or comparative test.
- Mixed — tested, but without a control, with conflicting results, or with the tested variable entangled with something else.
- Design proposals — the idea appears only as a proposal or framework, with no test reported. These are collected separately at the end of the page, in Design proposals (not yet tested), and are not evidence.
The patterns are grouped by the function they serve in a lesson: placing effort before help, pairing AI feedback with human feedback, verifying understanding rather than output, making the learner the teacher, structuring collaboration, and confronting a specific misconception. Contexts (online, face-to-face, blended) and disciplines are reported with each, and summarized at the end.
Patterns that place effort before help
These patterns share a structural claim: the learner must commit to an attempt before the AI contributes. It is the most consistently supported design rule in the knowledge base.
Retrieve or attempt before the AI answers
Evidence: tested. The sequence is: attempt from memory, receive instruction or an example, practice in spaced sessions, consult the AI only after committing to an attempt, receive response-contingent feedback that probes the misconception, and advance only once the response shows adequate engagement.
An adaptive spaced-retrieval condition produced the highest posttest scores (M = 78.19) and significantly outperformed learner-directed AI study (M = 67.28, d = 0.92, p = .003) across 89 students in a blended statistics course, while fixed spacing was statistically indistinguishable from adaptive (Akgun & Toker, 2026). The contrast case is decisive: in a randomized trial of 120 students, the group that studied with unrestricted ChatGPT retained less on a surprise test 45 days later — 57.5% correct against 68.5% for traditional learners, t(83) = −3.19, p = .002, d = 0.68 — and had also studied roughly 45% less, with the disadvantage surviving a study-time covariate (Barcaui, 2025). A preregistered randomized field experiment in an online MBA found that gains tracked completed weeks rather than minutes of exposure (+2.00 points per additional completed tutoring week, p = .018), reading as spaced practice mattering more than total time (Yang et al., 2026).
Productive failure: attempt before instruction
Evidence: mixed, and thinner than its reputation. Students attempt a problem targeting a concept they have not been taught, the tutor withholds the solution and elicits multiple attempts, help comes only when strictly necessary, and consolidation follows with comparison and direct instruction.
The one field study that tests the full sequence with a steered tutor used 17 high school students in Singapore: the steered condition achieved a higher productive-failure score, significant for problem consistency (p = .046), and students produced on average 2.6 representations per session (p = .05), but no learning outcome was measured (Puech et al., 2025). The strongest support is indirect and comes from a randomized within-subjects experiment with 26 students, where immediate-answer Scaffolding performed significantly worse than peer, TA and tutor roles on model abstraction (β = −0.692, p = .015; β = −1.039, p < .001; β = −0.769, p = .005) even though students preferred the directive tutor — preference ran against competence (Zhu et al., 2026).
Error analysis and erroneous examples
Evidence: mixed. Students diagnose an error in an artifact — an AI-generated diagram violating referential integrity, a query that drops a join, an Large Language Models (LLMs) code snippet — receive clues that force them to infer the fix rather than being handed the correction, repair it, and reflect on which parts of the output were untrustworthy.
A pre-post study of 13 students in an online database course rose from 4.25 to 6.83 out of 7 (t(12) ≈ 5.10, p < .001, d = 1.49) using weekly critique-and-refinement cycles built on deliberate AI failure cases, but with no control group the gain cannot be separated from the curriculum or instructor (Hosseini, 2026). A systematic review of 72 computing-education studies reports that error analysis is a distinct competence: students performed significantly worse at correcting LLM-generated code than at traditional programming exam tasks (Kumar et al., 2026). No article in the knowledge base reports a controlled test of erroneous-example instruction as such.
Worked examples with self-explanation
Evidence: mixed — the two halves diverge. A worked example is presented with missing justifications to complete, or with errors to find; the student completes or repairs it, explains their reasoning, and then attempts the next problem.
Adaptively assigning guided and buggy examples beat random problem-type assignment in a classroom study of 113 students (posttest M = 72.3 and 72.5 against 65.7 control, A = .58, p = .005 and p = .002), and the Knowledge Tracing variant narrowed the achievement gap by 77.1% for low Prior Knowledge students (β = 9.4, p = .001) (Dey Tithi et al., 2026). But adding the self-explanation step to elaborated AI feedback lost on every measure in a preregistered experiment with 302 participants: it doubled feedback time (4.1 vs 2.1 min, p < .001), cut problems completed by 40% (2.0 vs 3.4, p < .001), produced no gain per episode (OR = 1.03, p = .486), and lowered end-of-session mastery (65% vs 79%, d = .41, p < .001) (Asher et al., 2025).
Patterns that pair AI feedback with human feedback
The knowledge base's most replicated finding about AI feedback is that it works better combined with human feedback than alone — and that its quality is not what determines whether students use it.
PAIRR: peer and AI review with reflection
Evidence: mixed (widely implemented, not controlled). Students read and reflect on how AI and feedback work, draft, give and receive peer review, prompt the AI for criteria-driven feedback on the same draft, critically compare both, write a revision plan, revise, and reflect on which feedback changed what.
The largest study of college students' use of AI feedback to date followed 654 students across ten writing courses and three writing-intensive STEM courses: 58% preferred combined ChatGPT and peer feedback, 36% peer alone and only 6% AI alone; 75% found the two similar and mutually reinforcing; AI feedback was called "overly general" by 31% while peer feedback was more specific for 28%; and only 5.3% showed overconfidence in AI feedback (Sperber et al., 2025). The model was also run in an upper-division business writing course with 34 students, where about a quarter of coded reflections expressed skepticism about or noted inaccuracies in AI feedback (MacArthur et al., 2025). Both report perception data; neither has a control condition, which is why the pattern is mixed rather than tested.
Combined peer and AI feedback
Evidence: tested. The comparative evidence comes from outside the PAIRR program. A quasi-experiment with 122 Chinese EFL students found AI-plus-peer integrated feedback raised behavioral, affective and cognitive engagement against peer-feedback-only (all p < .001; partial η² = 0.28, 0.28, 0.32) and improved writing on all four IELTS dimensions (F(1,119) = 42.68, p < .001, partial η² = 0.26), largest on task achievement (d = 1.41) — with no delayed post-test, so durability is unmeasured (Liu, 2026). In a randomized study of 45 student teachers in 12 groups, GenAI-supported peer feedback beat plain peer feedback on argumentation, and the prompt-scaffolded variant performed best on advanced elements such as rebuttal data and addressing the opposing view (Chang et al., 2026).
AI critique then revision
Evidence: tested — with an important null. Draft, prompt the AI for rubric-driven feedback, critically assess that feedback against the rubric and the sources, write a revision plan, revise, and reflect.
A controlled 2 × 2 factorial experiment with 120 English majors found writing-quality gain was highest for the group trained in both filtering and appraisal (M = 7.92), against 6.10, 4.56 and 3.10 for the other conditions, with deep revision rising from 28% to 48% and the advantage persisting on a new topic and after AI support was removed (Dai, 2026). The null is the instructive part: in a randomized three-group experiment with 70 students, chain-of-thought-prompted AI feedback was significantly higher quality than both zero-shot AI feedback (p = .01) and teacher feedback (p = .008), yet this quality advantage did not translate into greater revision gains — teacher feedback produced comparable improvement (Farrokhnia et al., 2026).
Human-in-the-loop review of AI output
Evidence: mixed — the review step is rarely the tested variable. AI generates draft output, automated verifier agents check it for realism, readability or hallucination, failed checks loop back for refinement, and a teacher reviews, edits, and accepts or discards before anything reaches students.
A four-agent loop with 8 teachers produced 212 problems of which 166 were accepted as-is, and realism checks worked as intended (10 realism problems flagged, 20 quantity or unit edits, no mathematics error found in a final problem) — but interest fit was the weak point, with students rejecting the topic in 160 of 422 responses (Walkington et al., 2026). An educator-in-the-loop feedback tool rated by 30 teachers never fell below 4.1/5 across nine items and cut median time per assignment from 10–30 minutes to under 5, but the authors acknowledge no student evaluation, so no learning claim is supported (Zhao et al., 2025). A red-team experiment makes the stakes concrete: 2 of 5 prompt injections changed a grade undetected, at 100% (9/9) and 94% (17/18), leaving the teacher as the only real check on AI grading output (Humble, 2026).
Patterns that verify understanding rather than output
Because AI can produce a competent artifact, these patterns move assessment to evidence the artifact cannot supply on its own.
Oral and viva verification
Evidence: tested as a format, but results are about scores and affect rather than learning. A coding assignment is submitted with AI permitted, followed within 48 hours by a mandatory 15-minute oral code review in which the student explains the program and runs integration tests live, graded 70% on the review and 30% on the rubric.
A three-semester quasi-experiment with 96 students found no statistically significant change in exam performance despite the new policies (~2% improvement on one exam), while pasted-to-total characters rose from 61.0% to 68.1% (p < 0.0001); 90% of students said the reviews motivated them to understand their code better and 65% that they helped avoid over-reliance (Fowles et al., 2026). Asynchronous recorded oral responses produced significantly higher scores than in-person multiple-choice (midterm Md = 92.5 vs 70, p < .001; final Md = 94.2 vs 86.4, p = .002) with only moderate cross-format correlations (τ = .44 and .25) — and the authors caution these are format score differences, not evidence of learning gains, with cheating behavior unmeasured (Pentland et al., 2026). Direction is not uniformly positive: students were calmer in a chat-based viva (M = 6.50 vs 5.86, p = .028) but rated the face-to-face viva significantly better for understanding their own work (p = .004) (Yusuf et al., 2026).
Staged checkpoints and process evidence
Evidence: mixed — no controlled test of the mechanism itself. Coursework runs as staged modules, each ending in a checkpoint that verifies both the output and the approach — a correct result reached by hardcoding is rejected — with a pre-advancement check that returns the learner to skipped steps.
A case study of 5 graduate students in a self-paced quantum-information course logged 75 interactions and confirmed the dual output-and-approach checkpoint functioned as intended, with no control group (Elhaimeur & Chrisochoides, 2026). A 27-participant pilot using stop-block checkpoints reported significant Self-Efficacy gains across all ten assessed skill areas (p < 0.001) in a within-subjects pre-post design where gains cannot be separated from practice effects (Naboulsi, 2026). A three-year quasi-experiment with 248 biomedical engineering students found higher A-rates after adding four-module problem-based learning with milestones and rubrics (66.4% vs 39.1%, Δ = +27.3 points, p = 0.042), persisting after excluding the pandemic-affected year, but the comparison is historical and non-randomized (Nnamdi et al., 2026).
Patterns that make the learner the teacher
Learning by teaching an AI tutee
Evidence: tested, and the best-evidenced pattern on this page. The student studies the content, then explains it to an AI prompted to hold a novice stance that never reveals the target explanation; the AI asks for explanations, examples and verification reasoning, sequenced from lower- to higher-order, and persists until the explanation is satisfactory.
A quasi-experiment with 68 preservice teachers found explaining to a GAI novice learner scored higher on defining the flipped classroom (M = 4.18 vs 3.29, p < 0.001, r = 0.474) and its activities (M = 4.91 vs 3.06, p < 0.001, r = 0.642), generated more and higher-quality questions (both p < 0.001) — but showed no group difference on objective questions (M = 23.18 vs 21.57, p = 0.416) (Wang et al., 2026). A randomized lab experiment with 41 students found higher knowledge-test scores (adjusted 11.86 vs 10.53, F = 35.54, η² = 0.74) and clearer, more readable code, but no difference in code correctness (Chen et al., 2024). An 11-week deployment across 546 students found each additional deep-learning act associated with a 2.7% decrease in expected quiz attempts (IRR 0.973, p < .001), with the comparison confounded by time-on-task and late-semester circumvention rising to 30–35% external content reuse (Wang et al., 2026). The consistent shape: gains in explanation and generative work, not in objective recall.
Patterns that structure collaboration
Scripted roles with one shared AI
Evidence: tested, but in controlled or uncontrolled settings rather than ordinary classrooms. Two learners share one AI and are assigned explicit roles with rotation rules; the AI is configured to take a role as needed and its output goes to the whole group. In a pair-programming variant the shared AI models the dyad's joint attention and effort, forecasts a breakdown up to 30 seconds ahead, and escalates scaffolds in tiers from doing nothing to a directive hint.
A within-subjects experiment with 26 dyads found the feedback condition achieved higher debugging success (t[49.96] = −13.51, p < .0001) and finished faster (t[44.70] = 4.39, p < .0001), though it required dual eye-tracking and pupillometry hardware and did not test transfer to unsupervised pair work (Golrang et al., 2026). A quasi-experiment with 58 graduate students in 16 groups found role design raised mind-map content scores from 3.65 to 4.59 on a 1–5 SOLO scale (z = 3.771, p < 0.001) while node and branch counts stayed flat — but with no control group, practice effects cannot be ruled out (Cheng et al., 2026).
AI-assisted discussion
Evidence: mixed — the one direct implementation is a case description. Students analyze a scenario and answer guided questions independently, prompt ChatGPT with a standardized prompt on the same questions, evaluate the AI responses for accuracy against their own, refine their answer, and close with a whole-class discussion.
The economics activity deliberately exploits an AI error — ChatGPT calls the song's depicted behavior "perfectly elastic" demand when the correct answer is inelastic — turning validation of AI output into the discussion (Beck & Brodersen, 2025). It reports instructor impressions, not a measured outcome. A tested collaborative-discussion sequence with 67 teacher-education students found the experimental group outperformed a lecture control (M = 51.45 vs 43.89, p = 0.001, g = 0.839) with co-regulation rising (p = 0.043, g = 0.512) — but AI was used to design the technique, not to mediate the discussion (Tutal, 2026).
Patterns that confront a specific misconception
Refutation and conceptual change
Evidence: tested — with directly conflicting results. Elicit the learner's specific belief, present a refutation text or a personalized AI dialogue that confronts it, engage with the counter-evidence and the correct explanation, restate the correct conception, then retest after a delay.
A preregistered experiment with 375 adults found personalized misconception AI dialogue produced significantly larger immediate belief reductions than both textbook-style refutation and neutral AI dialogue, persisting at 10 days but converging with textbook refutation by 2 months (Corbett & Tangen, 2026). A Solomon four-group quasi-experiment with 413 tenth-graders found the reverse: expert-written and AI-generated conceptual-change texts were both significantly more effective than interactive ChatGPT dialogue, which showed no significant advantage over control, and gains were almost exclusively limited to high-achieving students (Akdogan, 2025). The second paper flags the conflict explicitly and attributes it to prompt design and domain. Refutation text itself beat control in both.
Contexts and disciplines
The pattern determines what the context requires, and several patterns were tested in only one setting:
- Online and asynchronous. Retrieval and spacing, the Socratic variants, worked examples with self-explanation, prompt-scaffolded use, asynchronous oral assessment, and the learning-by-teaching deployments. Asynchronous settings make the sequencing load-bearing, because the system cannot see whether the student attempted first.
- Face-to-face and blended. Productive failure, error analysis, the flipped variants, oral code review, scripted collaboration, and the conceptual-change studies. Class time is often reallocated rather than replaced — in the oral-review pattern, lectures moved to video so class time could hold the interviews.
- Disciplines represented in the tested evidence. Writing and language learning (PAIRR, combined peer and AI feedback, AI critique then revision), mathematics (retrieval and spacing, productive failure, worked examples, error analysis), computer science (Socratic assistants, error analysis, oral code review, learning by teaching, scripted pair work), medicine (Socratic scaffolding in clinical interviews), teacher education (scripted argumentation, learning by teaching), business (flipped MBA tutoring), and physics, science and vocational settings for the conceptual-change and oral-assessment studies.
What the evidence does not yet support
Stated plainly, because these are the findings most likely to be quietly dropped:
- Better AI feedback does not produce more revision. Higher-quality chain-of-thought feedback beat teacher feedback on quality and produced no revision advantage (Farrokhnia et al., 2026).
- Socratic questioning is not automatically better. A randomized trial of 132 students found the Socratic assistant with full context rated significantly worse for supporting task completion than all other configurations (mean rank 48.63, μ = 3.53, against 4.27, 4.16 and 4.12; χ²(3) = 12.14, p = .007), with the most external LLM use (23%) and the fewest full-comprehension responses (48% vs 67%) (Eastwood et al., 2026) — while a medical RCT found a multi-agent system containing a Socratic tutor beat its control on examination and communication scores (Yang et al., 2026).
- Students prefer the less effective role. Directive tutoring was preferred while immediate-answer scaffolding depressed model abstraction (Zhu et al., 2026).
- Verification formats change scores without demonstrating learning. Oral formats raised scores and reduced anxiety while one study found face-to-face better for understanding, and no study measured cheating.
- No controlled test isolates human review of AI output, and the checkpoint mechanism has never been tested as the manipulated variable.
- Productive failure and error analysis rest on small, uncontrolled studies (n = 17 and n = 13) that measure strategy fidelity or self-report rather than learning outcomes.
Design proposals (not yet tested)
The patterns below come from a faculty guide supplied by the knowledge base maintainer (AI-Ready Course Design, September 2026). That guide states explicitly that its examples are design proposals, not tested interventions, and no article in this knowledge base tests them. They are recorded here as design ideas worth trying and evaluating, and must not be read as evidence.
- Argument + revision trail (composition, humanities). Replace an essay-only submission with an initial thesis, two annotated source passages, a revised essay and a 150-word decision note; permit AI critique after the first draft. Assess the claim–evidence connection and one accepted or rejected suggestion justified against the sources.
- Data + reasoned claim (science, laboratory courses). Replace a polished lab report with raw observations, a graph, an uncertainty note and an explanation linking results to a claim; AI may critique a supplied interpretation, and students verify that critique against their data.
- Attempt + error analysis (precalculus, calculus). Replace answer-only homework with an initial attempt, analysis of a flawed worked solution, and a corrected explanation; hints are permitted only after the attempt. This is the design form of the error-analysis pattern above, and it inherits that pattern's weak evidence base.
- Position + challenge + reconsideration (psychology, sociology). Replace "post once, reply twice" with a case-based claim using a course concept; a peer supplies a counterexample and the author revises or defends with evidence.
- Project + linked check (business, health professions). Pair an AI-permitted recommendation for a fictional organization or patient case with a short explanation of two key decisions and a response to a changed constraint; publish the grading relationship between the two components.
- Plan + try + adapt (college success). Replace a generic time-management reflection with a one-week study plan, a brief record of trying it, and a revision tied to what happened; AI may suggest scheduling options after the student identifies constraints.
The guide's own cautions apply: recorded video, reflections and logs can themselves be AI-assisted, so in fully asynchronous courses a recording should not be treated as verification of independent mastery. Two of its framings — assessment twins and asynchronous oral assessment as a cheating deterrent — remain frameworks awaiting validation; the asynchronous oral-assessment studies cited above measured format scores and did not measure cheating at all.
Connected Concepts
- Pedagogies and Teaching Strategies — the umbrella of teaching approaches this page operationalizes into sequences
- Learning Design — where patterns are chosen, sequenced and embedded in a course
- Scaffolding — the support-and-fading principle that governs where AI help belongs
- Feedback — the system this page's feedback patterns instantiate
- Formative Assessment — the assessment purpose most of these patterns serve
- Peer Assessment — the human half of the PAIRR and combined-feedback patterns
- AI Feedback Quality — why feedback quality alone does not determine revision
- Evaluative Judgment — the appraisal students must exercise on AI output
- Feedback Literacy — the capability the critical-appraisal steps are building
- Human-in-the-Loop — the oversight structure of the review patterns
- Oral Assessment — the format underlying the verification patterns
- Process-Oriented Assessment — the logic behind staged checkpoints
- Productive Failure — the concept behind attempt-before-instruction
- Retrieval, Spacing and Interleaving — the evidence base for spaced retrieval patterns
- Desirable Difficulties — why effortful sequences outperform fluent ones
- Misconceptions about AI — what the conceptual-change patterns target
- Refutation Text — the text form of misconception confrontation
- Learning by Teaching — the pedagogy behind the AI-tutee pattern
- Socratic Method — the questioning pattern and its conflicting evidence
- Collaborative Learning — the context for scripted shared-AI work
- Cognitive Offloading — the risk every effort-first pattern is designed to avoid
- Metacognition — what reflection steps in these sequences are meant to trigger
- Prompt Engineering — the scaffolding layer in structured-use patterns
- Transfer of Learning — the outcome most patterns are ultimately judged on
- Assessment Validity — the reason process evidence is proposed at all
- Academic Integrity — the driver behind oral and process verification
- AI Literacy — the capability developed by critiquing AI output
- Online Teaching and Learning — the context that makes sequencing load-bearing
- Higher Education — the level where most of this evidence was generated
- K-12 — the level of the productive-failure, conceptual-change and oral-assessment studies
Connected Articles
- Peer and AI Review + Reflection (PAIRR): A Human-Centered Approach to Formative Assessment — Peer and AI Review + Reflection (PAIRR): the flagship sequence, N = 654 (Sperber et al. 2025)
- GIFT-AI: Teaching the Game and Leveling the Field: Peer and AI Review + Reflection in a Business Writing Course — PAIRR applied in a business writing course (MacArthur et al. 2025)
- AI-peer integrated feedback in second language writing classes: exploring students' engagement and writing performance — Integrated AI-plus-peer feedback raised engagement and all four IELTS dimensions (Liu 2026)
- Leveraging generative AI to facilitate peer feedback in collaborative argumentation learning — Prompt-scaffolded GenAI peer feedback in collaborative argumentation (Chang et al. 2026)
- Dai (2026) — Training students to filter and appraise AI feedback: the FRAC and APCA conditions
- Generative AI offers more, but students revise less: comparing the effects of teacher and AI feedback on student essay revisions — Higher-quality AI feedback produced no greater revision gains (Farrokhnia et al. 2026)
- Guardrails or Roadblocks? Effects of Pedagogical Style and Context Awareness in AI Teaching Assistants for Programming — Socratic plus full context was rated worse than every other assistant configuration (Eastwood et al. 2026)
- Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training — Need-triggered Socratic scaffolding in clinical interview training, N = 100 (Yang et al. 2026)
- Analyzing Undergraduate Problem-Solving in Physics Through Interaction With an AI Chatbot — Socratic chatbot in introductory mechanics: question specificity rose from 10–15% to 100% (Hashmi et al. 2025)
- The Effects of Agent Type and Feedback Style on Self-Directed Learning: A Mixed-Methods Study — Socratic versus directive agent feedback, with the order not counterbalanced (Han et al. 2026)
- Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention Study — Adaptive spaced retrieval beat learner-directed AI study (Akgun & Toker 2026)
- ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention — Unrestricted ChatGPT during study lowered 45-day retention (Barcaui 2025)
- When AI Tutors Speak: Evidence from a Randomized Field Experiment — Structured tutoring gains tracked completed weeks, not minutes (Yang et al. 2026)
- Evidence and Theory for why the Best Example-Problem Ratio To Optimize Learning Gain Depends on Knowledge Content — Examples versus practice cross over by knowledge type (Rachatasumrit et al. 2025)
- Adaptive Scaffolding for Cognitive Engagement in an Intelligent Tutoring System — Adaptive guided and buggy examples in an intelligent logic tutor (Dey Tithi et al. 2026)
- Benefit or Bottleneck? Assessing the Impact of Structured Reflection on Learning from AI-Driven Explanatory Feedback — Adding self-explanation to AI feedback lost on every measure (Asher et al. 2025)
- Generative AI without guardrails can harm learning: Evidence from high school mathematics — Guarded tutoring removed the exam harm that unguarded GPT caused, ~1,000 students (Bastani et al. 2025)
- Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics — Guided versus unrestricted LLM use in statistics (Amanlou et al. 2026)
- Preferred Scaffolding Does Not Lead to Better Learning Performance: Empirical Evidence from AI-Supported Mathematical Modelling — Immediate-answer scaffolding depressed model abstraction while being preferred (Zhu et al. 2026)
- Towards the Pedagogical Steering of Large Language Models for Tutoring: A Case Study with Modeling Productive Failure — Steering an LLM tutor to withhold solutions for productive failure (Puech et al. 2025)
- The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking — Weekly critique-and-refinement cycles built on deliberate AI failure cases (Hosseini 2026)
- Generative AI in computing education: A systematic review and a framework for responsible integration — Error analysis as a distinct competence, across 72 computing-education studies (Kumar et al. 2026)
- Clue before correction: ChatGPT-enhanced strategy for promoting autonomous and reflective language learning — Guided clues instead of direct error correction in L2 revision (Lukešová & Jennings 2026)
- The Contribution of Generative Artificial Intelligence as a Novice Learner to Students in the Learning by Teaching Model — Explaining to an AI novice learner, N = 68 (Wang et al. 2026)
- Learning-by-Teaching with ChatGPT: The Effect of a Teachable ChatGPT Agent on Programming Education — Randomized test of teaching a ChatGPT agent to program (Chen et al. 2024)
- Turning 500+ Students into Teachers: A Semester-Long Study of an AI Teachable Agent in an Undergraduate Algorithms Course — Learning by teaching deployed to 546 students over 11 weeks (Wang et al. 2026)
- Learning by Teaching: Engaging Students as Instructors of Large Language Models in Computer Science Education — Students designing questions an LLM cannot answer (Yang et al. 2025)
- Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom — Mandatory oral code review interviews as a response to GenAI in CS1 (Fowles et al. 2026)
- Asynchronous Oral Assessments: Enhancing Integrity, Engagement, and Communication in the AI Era — Asynchronous recorded oral assessment versus in-person multiple-choice (Pentland et al. 2026)
- Exploring student anxiety and experience in performance-based assessments using AIvaluate: an LLM-augmented emotionally — Chat-based viva reduced anxiety while face-to-face was rated better for understanding (Yusuf et al. 2026)
- Designing AI-Supported Oral Assessment in TVET — AI surfacing rubric evidence for teacher judgment in vocational workshops (Adams 2026)
- From Prototype to Classroom: An Intelligent Tutoring System for Quantum Education — Output-and-approach checkpoints in a self-paced graduate course (Elhaimeur & Chrisochoides 2026)
- Agentic Education with AI Coding Assistants — Stop-block checkpoints and a pre-advancement check, N = 27 (Naboulsi 2026)
- Advancing Problem-Based Learning in Biomedical Engineering in the Era of Generative AI — Four-module problem-based learning with milestones and rubrics (Nnamdi et al. 2026)
- Mathematics Teachers’ Interactions with a Multi-Agent System for Personalized Problem Generation — A four-agent review loop with teachers, and where interest fit failed (Walkington et al. 2026)
- LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop — Educator-in-the-loop feedback with verifier scores (Zhao et al. 2025)
- Ethical implications of prompt injection in AI-mediated grading: An adversarial red-team evaluation — Prompt injections that changed grades undetected, leaving the teacher as the only check (Humble 2026)
- ProPACT: A Proactive AI-Driven Adaptive Collaborative Tutor for Pair Programming — A shared AI forecasting collaboration breakdown in pair programming (Golrang et al. 2026)
- Enhancing Human-Generative Artificial Intelligence Online Collaboration Outcomes: The Pivotal Function of Symbiotic Role Design — Scripted learner and AI roles in group knowledge construction (Cheng et al. 2026)
- ParaTutor: LLM Mediated Parent Child Tutoring through Role Separated Scaffolding Interface in Real Time — Role-separated AI support in parent–child tutoring (Luo et al. 2026)
- AI tutors vs. tenacious myths: Evidence from personalised dialogue interventions in education — Personalized AI dialogue beat textbook refutation immediately, converging by two months (Corbett & Tangen 2026)
- Comparing the effectiveness of expert-written text, AI-generated text, and interactive AI dialogues on students' — Conceptual-change texts beat interactive AI dialogue, the reverse result (Akdogan 2025)
- AI-Supported Inquiry-Based Learning in Photosynthesis and Respiration: Implications for Sustainable Science Teacher Education — AI-supported guided inquiry and conceptual understanding (Aydin 2026)
- From Traditional Classroom to AI-Enhanced Flipped Classroom: A Three-Year Pedagogical Evolution for International Students in Pharmacology — Three-cohort comparison of traditional, flipped and AI-enhanced flipped (Liu et al. 2026)
- Integrating Flipped Learning and Generative AI for Practice-Based Design Education: Evidence from a Knit Yarn Design Course — Flipped studio course with parameter-guided prompt scaffolding (Qu et al. 2026)
- The Impact of Generative Artificial Intelligence on Learning Outcomes in Higher Education: A Meta-Analysis — Teaching model moderated outcomes: flipped g = 1.96 against traditional g = 0.46 (Jing et al. 2026)
- Student Adoption of an AI Tutor for Statistical Programming: A Longitudinal Study — Weekly homework-tutor use predicted transfer-task scores in a flipped course (Préau et al. 2026)
- Fostering Generative AI Literacy in Economics: A Hands-on Approach — A five-step AI-critique discussion built on an AI error (Beck & Brodersen 2025)
- Artificial intelligence assisted design of a novel cooperative learning technique for higher education — A tested cooperative-learning sequence with assigned roles and a gallery walk (Tutal 2026)
- Developing an AI-Assisted Seminar-Based Learning Platform With Embedded Librarian Support: Enhancing Information Literacy of Engineering Research Teams — Seminar and peer-discussion module with an AI retrieval recommender (Huang 2026)