Research Article
Exploring Mathematics Teacher Educators' Lesson Study Experiences in Supporting Pre-Service Teachers' Dialogic Engagement With AI
Synthesis: Robinson, Erbilgin, Johnson, Hashem and Gningue (2026) report a qualitative study at one teacher education institution in the United Arab Emirates in which five teacher educators used lesson study to design AI-supported mathematics lessons for pre-service teachers (PSTs). Two research lessons were taught, one in a secondary and one in a primary mathematics course, reaching 36 PSTs (12 secondary, 24 primary), with ChatGPT (version GPT-4o) used for iterative task generation and refinement. Thematic analysis of observation notes, PST work, interviews and the educators' reflections yielded three educator themes (developing pedagogical knowledge for AI integration, ethical and critical positioning of AI, shifts in pedagogical confidence and design) and three PST themes (use of AI for educational purposes, AI as a dialogic partner in building conceptual tasks, increased prompt engineering skill). Of 26 analyzable prompts drawn from 15 worksheets, 16 incorporated conceptual task features. The authors argue that prompt engineering operated as pedagogical rather than technical thinking.
Key Findings
- Two cycles, 36 pre-service teachers. Five teacher educators ran two lesson study cycles around a 75–90-min research lesson in secondary and primary mathematics courses, one cycle in each course.
- Educators built pedagogical knowledge. They reported learning prompt engineering, the design of structured AI-supported activities, and the use of AI as a dialogic resource rather than an output generator.
- Beliefs were reinforced or repositioned. Some educators said the process reaffirmed commitments to critical and ethical AI use, while one moved from a tool-counting view toward a pedagogy-first perspective.
- Practice became more explicit. Educators modeled prompt refinement, anticipated PST misconceptions collectively, and planned prompts requiring students to critique AI output; the revised second lesson ran as intended.
- PSTs negotiated with AI rather than accepting it. They compared and refined successive outputs and judged for themselves when a task was conceptual enough; of 26 analyzable prompts, 16 carried conceptual task features.
- Self-reported learning rose in the revised lesson. In research lesson 1, 5 of 7 PSTs reported learning to distinguish conceptual from procedural tasks; in research lesson 2, 10 of 14 reported learning that distinction plus prompt engineering.
How the study was designed
The study paired lesson study with Generative AI in Math Education. The five educators, four mathematics educators plus one colleague with AI-focused expertise, agreed on two shared challenges: PSTs struggled to distinguish procedural from conceptual tasks, and PSTs accepted AI-generated outputs at face value. Planning adopted two scaffolds, a prompt engineering framework modified from Giray (2023) and Lin (2024), and a modified Mathematical Task Analysis framework (Stein et al. 2009) for judging cognitive demand. In the main activity, PSTs generated a task with ChatGPT, evaluated its cognitive demand, revised their prompt to raise conceptual depth, and repeated the cycle. The first lesson ran in a secondary course, the revised version in a primary course. Data included observation forms, PST artifacts, interviews and the educators' reflections, analyzed with Braun and Clarke's (2006) six-phase approach and collaborative coding.
What the teacher educators reported
Educators described gains in Teacher AI Competency. One framed prompt engineering as "not a technical skill; it's a cognitive one," describing how watching students iterate shifted her understanding of AI as a source of better questions rather than answers. Another concluded that "robust pedagogical/content/PCK knowledge is needed to engage with AI tools critically," which the authors read as treating AI integration as a design task grounded in Technological Pedagogical Content Knowledge (TPACK). On beliefs, one held that AI "should not replace thinking, it should provoke it," while another said lesson study let her enact existing beliefs. Practice changes centered on clarity and structure: modeling the first iteration, Scaffolding through paired work, visible displays of both frameworks, and questions requiring students to critique AI output. The team also reported logistical obstacles (one class section, virtual observation, scheduling), pedagogical disagreements, and inexperience with teaching via AI, which they worked through collectively.
How pre-service teachers engaged with AI
PST interactions with AI became dialogic. Instead of requesting an answer, they negotiated the procedural or conceptual character of a task across rounds of prompting, using both frameworks as mediating artifacts, evidence of Human AI Collaboration and of growing Student Engagement in task design. One PST said the main activity "requires you to look at the depth of conceptual tasks, so you know how to generate one in the future with the help of AI." Prompt artifacts support the claim: later prompts added explicit criteria, for instance a final prompt on addition within 20 that asked for non-algorithmic thinking, a real-world context and multiple solution paths. Challenges persisted. One PST noted that asking AI to make a task more conceptual sometimes produced a more complicated task, later reporting that a more detailed prompt worked better; several also wanted to learn more ChatGPT features.
What this means for practice
- Faculty developers. Treat prompt engineering as pedagogical rather than tool training, since these educators described it as a cognitive skill in which clarity and specificity shape output.
- Instructors. Pair AI with a disciplinary task analysis framework so prompt revision is anchored in conceptual depth, as the team did when judging cognitive demand.
- Instructors. Plan critique-and-iterate prompts that ask students to evaluate AI output against explicit criteria and revise it, building Critical Thinking into the activity rather than saving reflection for the end.
- Institutions. Fund sustained Collaborative Learning through joint planning and shared reflection on student thinking, since learning came from inquiry into real lessons rather than standalone workshops.
Limitations
- Different PST groups took each research lesson and completed mathematically different tasks, so the lesson revision coincided with changes in group composition and content; cross-cycle comparisons should be made cautiously.
- The study cannot establish that lesson study caused the reported changes, or that it outperforms other professional learning; wider practice change rests on the educators' reflections, and sustainability was not examined.
- The educators held overlapping roles as participants, designers, data collectors and researchers, which may have influenced reflections and interpretations despite collaborative analysis.
- Participation varied across data sources, and ChatGPT was the sole generative AI platform, so prompt refinement patterns may not transfer to models with different interactional features.
Citation
Robinson, Jennifer M.; Erbilgin, Evrim; Johnson, Jason D.; Hashem, Reem; Gningue, Serigne M.. (2026). Exploring Mathematics Teacher Educators' Lesson Study Experiences in Supporting Pre-Service Teachers' Dialogic Engagement With AI. Journal of Computer Assisted Learning, 42, e70321.