Research Article
AI in Action Learning Tour
Synthesis: Instruction Partners spent the 2025–26 school year studying 20 student-facing AI-powered instructional products already in use across at least 1,000 school systems, observing 32 classrooms in 16 school systems across seven states and interviewing 50+ teachers, 25+ school and system leaders and 100+ students. This is a report, not a peer-reviewed study: its judgments come from fieldwork rather than measured effect sizes. Three product categories emerged: three general-purpose chatbots, five education-focused multipurpose platforms and 13 targeted instructional tools. The report's sharpest claim is that the chatbots pose the clearest threat to student thinking, because they make it easy to bypass the reasoning and productive struggle learning requires while appearing to complete the work. Its procurement number is sobering: seven products have independent reviews that examine student achievement across student groups in the US. Teachers stayed central, implementation mattered as much as design, and the recommended actions are the authors' proposals rather than findings.
Key Findings
- Products in use, not in demonstration. The team observed 32 classrooms in 16 school systems across seven states, with 50+ teacher, 25+ school and system leader and 100+ student interviews.
- Three categories of product. Three of the 20 are general-purpose chatbots, five are multipurpose platforms and 13 are targeted instructional tools; the chatbots pose the clearest threat to student thinking.
- Independent evidence is thin. Of the 16 profiled products, seven have independent reviews checking achievement across student groups in the US; two studied abroad, four have studies underway and three rely on internal data.
- Teacher-role design tracked with use. Nine of the 16 profiled products specify teacher actions during use and were used closer to their intended approach; all 13 targeted tools have a teacher dashboard.
- The same product can land well or badly. Activities diverged by classroom according to how the teacher framed the purpose, whether they watched the dashboard, and how fast they acted on it.
- Users valued feedback they could act on. Fast, specific, individualized feedback that pushed students back into thinking drew the most praise; teachers objected to loopholes that faked mastery and feedback that misread correct work.
Three product categories, one clear source of risk
The general-purpose chatbots, ChatGPT, Claude and Gemini, are generative AI tools with different access rules: ChatGPT limits access below the age of 13 and requires parental consent for 13–17 year-olds, Anthropic restricts Claude to account holders who are 18+, and Google allows some K–12 students to use Gemini. Five of the 20 are education-focused multipurpose platforms, so what students see depends on what a teacher or system assembles. Thirteen are targeted instructional tools: five for multiple subjects, four for literacy and four for math. The report found those targeted tools produced the strongest and most consistent experiences, and that products with a clearer point of view about the teacher's role were used more consistently.
Implementation decides whether a product helps
The report is emphatic that implementation matters as much as design. The same product produced different experiences depending on how the assignment was introduced, what the teacher did while students worked, and how the teacher used what the tool reported. The strongest implementations shared conditions: a clear point in the scope and sequence, scheduled time, predictable routines, logistics settled in advance, and a clear vision for teacher actions during use. Where those were missing, ten-minute activities dragged out all class, parts meant to be informal and social were graded as assignments, and teachers inserted moves that kept students from attempting the work independently. Teachers needed sustained coaching, not one rollout session, and where leaders knew a product well, teachers took the desired actions more often. That is change management work, and the loopholes teachers complained about are misuse that harms learning.
What the evidence base does and does not support
Of the 16 products with complete profiles, seven have independent reviews that examine student achievement across student groups in the US, two have conducted studies outside the US, four have external studies underway and three look only at internal data. The authors argue that independently evaluated causal studies, covering priority groups, should be the expectation for all student-facing products, which puts interpreting the evidence in adoption decisions. In literacy, benefits included more writing practice with faster feedback and precise diagnosis of letter sound combinations, while risks included practice disconnected from any text and easy text releveling that lowers cognitive demand. In math, products that captured reasoning rather than only final answers let teachers see thinking, but misdiagnosed misconceptions sent students back to relearn material they already knew.
What this means for practice
- School and system leaders. Set Guardrails through policy: state where AI has a role in instruction and where it does not, and lock down chatbot access during assignments and assessments.
- Before adopting anything, inventory the program. Start from gaps in your instructional program rather than shopping, pilot at one grade, and decide upfront what conditions a product must meet to scale.
- Teachers. Frame tool use as learning rather than completion, monitor the dashboard during the lesson, and keep the AI an assistant to a goal you set and assess.
- Product developers and funders. Treat implementation as a design problem by specifying what teachers should do during use, and fund independent studies before treating impact claims as established.
Limitations
- This is a report, not peer-reviewed research: it is qualitative and reports no control conditions, effect sizes or standardized outcome measures.
- The sample was not selected scientifically: products came from what partner systems were already using plus a market scan built to vary grade levels, subjects and designs.
- Several counts describe the 16 products with complete profiles rather than all 20, and the authors plan to revise the analysis as the tour continues.
- The report states that it did not analyze data privacy, technical integration, pricing, how products promote AI Literacy, or environmental impact.
Citation
Instruction Partners. (2026). AI in Action Learning Tour. Report, not peer reviewed. Written chiefly by Emily Freitag and Doe Kim.