On this page

Randomized controlled trial (RCT) — a research design in which participants are randomly assigned to a treatment or control condition to estimate the causal effect of an intervention on an outcome. In AI in education, RCTs are the gold standard for establishing whether an AI tool or pedagogical approach causes learning gains, engagement changes, or other outcomes, rather than merely correlating with them.

Questions to Consider

  • If a school tells you 'students who used the AI tool scored higher,' why might that still fail to prove the tool caused the gain — even if the difference is large?
  • Randomization balances known and unknown confounders across groups. Before you read, what does random assignment accomplish that simply comparing two intact classrooms cannot, no matter how well-matched they look?
  • The page calls the RCT the gold standard but lists real costs: artificial settings, fast-changing AI that dates trials, underpowered small samples, and ethical constraints on withholding helpful tools. Which of these trade-offs do you think is most often ignored in education research headlines?
  • An RCT with 1,174 participants found GenAI closed about three-quarters of an education-based productivity gap. But a well-run RCT can still be conducted on a narrow task in a contrived setting. What should you check about the outcome measure before trusting the causal claim?
  • Consider the ethics problem directly: if you had genuine reason to believe an AI tutor helps students learn, is it defensible to randomly deny it to half a classroom for a semester? How would you design an ethically sound study that still isolates the cause?

Introduction

Randomization is what distinguishes an RCT from other designs: by randomly assigning learners to conditions, an RCT balances known and unknown confounders across groups, so any observed difference in outcomes can be attributed to the intervention with high internal validity.

How RCTs appear in the research

  • Micro-RCTs as a response to fast-moving technology: Harrison et al. (2026) argue that conventional large-scale trials cannot keep pace with tutoring platforms that change materially during a study, and use teacher-led micro-randomized controlled trials across English secondary schools (644 of 929 students completing post-testing, g = 0.33) to keep causal estimation repeatable. The trade-offs are stated in their own design: 30.7% attrition, curriculum-aligned rather than independently standardized outcomes, and only four weeks of follow-up.
  • Causal efficacy claims: RCTs in AIED test whether an AI tutor, tool, or pedagogical treatment improves outcomes. A randomized experiment on generative AI with 1,174 participants found GenAI substantially narrows education-based productivity gaps, closing roughly three-quarters of the initial performance difference — a clear causal estimate of AI's effect.
  • Comparison to the gold standard: The research methods page situates RCTs as the strongest design for internal validity while noting their trade-offs — cost, artificial conditions, fast-changing AI, small underpowered samples, and ethical limits on withholding potentially helpful tools.

Strengths and limitations

  • Strengths: strongest causal inference; clean outcome measurement; supports effect-size estimation; balances confounders through randomization.
  • Limitations: costly and slow; artificial settings can reduce ecological validity; AI tools change faster than trials can run; small samples often underpower detection of meaningful effects; ethical constraints on withholding potentially beneficial AI from a control group.

For the fuller treatment of experimental design in AI in education — including when an RCT is appropriate versus quasi-experimental, survey, or computational designs — see Research Methods in AIED.

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.