On this page

Research methods in AIED — the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study AI in education: whether and how AI tools support (or harm) learning, and under what conditions. The knowledge base's corpus spans experimental, survey, qualitative, design-based, computational-benchmark, and review methods. Each has distinct strengths and limitations, and choosing among them involves trade-offs among internal validity (confidence in causal claims), external validity (generalizability), ecological validity (real-world authenticity), and the feasibility of studying fast-moving AI tools.

Questions to Consider

  • The page's central tension: the strongest designs for causal inference (randomized experiments) are the hardest to run in real classrooms, while the most authentic settings offer weaker causal control. If you had to decide whether an AI tutor helps learning, which of these two failures would you rather live with — and why?
  • Before you read, can you name the difference between internal, external, and ecological validity? The page argues every design trades these off. How might a study that's rigorously causal still tell you almost nothing useful about a real classroom?
  • A benchmark shows an AI scores high on accuracy, but the page insists high benchmark accuracy does not entail educational effectiveness. Why might a system that 'passes the test' still fail to help students learn — and what kind of evidence is missing?
  • Design-based research iterates on a real intervention but can't attribute gains to a specific mechanism, while an RCT isolates causes but runs in artificial conditions. Given the fast pace of AI change, how long do you think a rigorous RCT remains relevant before the tool it tested is obsolete?
  • Delphi expert consensus establishes agreement among experts, not empirical effect. When is it legitimate to build a competency framework from what experts believe, versus from data about what works — and how would you tell the difference in practice?
  • The page advocates triangulation — combining benchmark evaluation, experiments, measurement, and qualitative work to judge both whether a tool works and how. Before you read, where in a claim like 'this AI improves learning' would each method be needed to make you confident?

Introduction

The central tension in AIED research is that the strongest designs for causal inference — randomized experiments — are often the hardest to run with authentic AI tools in real classrooms, while the most authentic settings (field deployments, case studies, log-data analyses) offer weaker causal control. No single method resolves this; the field advances by triangulating across methods, and by being explicit about what kind of claim each design can support. Every method also carries cross-cutting limitations — generalizability, measurement validity, the fast pace of AI change, reproducibility, and weak theory use — that readers must weigh; see Limitations in AIEd Research.

The page's subject is method rather than findings. Learning Sciences is the substantive field these methods serve: where this page covers how a study should be designed, measured and reported, that page covers what the field has established about how people learn and how learning environments should be designed, and it treats design-based and mixed methods as the learning sciences' signature approaches rather than two options among many.

Reporting rigor and the TEP-AIED model

The reporting quality of AI-in-education studies is itself a research concern. The TEP-AIED model (Hwang, Xie, Wah & Gasevic, 2026) offers a structured framework for presenting AI-in-education research with rigor, organizing the essential components of a study report — theory/technology/educational problem framing, design, data, analysis, and results — so that readers and reviewers can assess whether claims are supported and whether the work is reproducible. It responds to the field's chronic weaknesses in reporting (vague tool descriptions, unstated model versions, omitted evaluation details) that the limitations page documents. Reporting frameworks like TEP-AIED sit alongside established reporting checklists (e.g., CONSORT-style guidance for trials, PRISMA-style guidance for reviews) as part of the field's broader move toward methodological transparency and reproducibility.

RAISE (Allison, 2026) approaches the same problem from the opposite direction — as a checklist rather than a narrative structure. It sets out 30 items across ten thematic domains (educational justification and theoretical grounding, AI system specification, AI role and interaction, Accessibility and cultural fit, setting and participants, human involvement, study design and evaluation, ethics and trustworthiness, transparency and reproducibility, and limitations and implications), with an editable version and a companion Ethics and Risk Matrix covering learner agency, equity of access, data AI Governance and algorithmic transparency. The two frameworks address each other directly: TEP-AIED characterizes RAISE as comprehensive but faults its breadth, arguing that "its breadth and granularity may make it complex and less accessible for routine empirical applications," while RAISE's own framing is that it mandates no method or model and only requires that choices be made visible. Read together they mark the trade-off in this literature — the fuller audit checklist versus the leaner three-dimension narrative — and they converge on the same non-negotiables: name and version the AI system, disclose prompts and interaction design, define treatment and comparison conditions, report ethical review and risk mitigation, and state whether outcomes measure performance, retention, or transfer. For a study to be adjudicated on those terms, the reporting instrument has to be adopted at design time rather than assembled at manuscript stage, which is the point both frameworks insist on. Corpus-level historical analysis is itself a methodological choice with transparency obligations: Rismanchian & Doroudi locate each paper in their AI×Ed framework based on author judgment of abstracts and full texts, explicitly acknowledge that this is not a systematic or scalable data-driven categorization, and make their full dataset publicly available as supplementary material for replication.

Experimental and quasi-experimental designs

An efficacy study tests whether an intervention produces its intended learning effect, typically using experimental or quasi-experimental designs that compare outcomes with and without the intervention. Experiments randomly assign learners to conditions (e.g., AI tutor vs. human tutor, or AI-scaffolded vs. unassisted) to estimate causal effects on outcomes like learning gains, engagement, or motivation. Randomized controlled trials are the gold standard for internal validity. A randomized field study of human support plus AI tutoring and an RCT on generative AI in teaching use assignment to isolate causal effects. Quasi-experimental designs (pre/post, between-subjects, or matched groups without randomization) are more feasible in intact classrooms but weaker on causal claims.

Survey and structural-equation-modeling studies

Cross-sectional surveys measure self-reported attitudes, perceptions, motivation, Self-Efficacy, and technology acceptance, often modeled with regression or structural equation modeling (SEM/PLS-SEM) to test hypothesized relationships and mediators. These dominate the knowledge base's corpus, particularly for acceptance, motivation, and psychological-mechanism questions.

  • Strengths: large samples; broad, low-cost coverage; can test complex mediational models of psychological mechanisms; feasible for studying attitudes that are hard to observe.
  • Limitations: cross-sectional data cannot establish causation; common-method/self-report bias; convenience sampling limits generalizability; mediators inferred from covariance, not manipulation.

The instrument itself deserves separate scrutiny. What a questionnaire, interview, or diary can and cannot establish — and the documented gap between what people report and what they do — is gathered on Self-Report Measures.

Qualitative methods

Interviews, focus groups, and thematic analysis produce rich, contextual accounts of how students and teachers experience AI tools, the meanings they attach to them, and the tensions and harms that standardized measures miss. Research on AI tutor safety and how AI changes teaching workflows rely heavily on qualitative evidence. See the dedicated Qualitative Research concept page for the full treatment of qualitative approaches — thematic analysis, grounded theory, phenomenology/phenomenography, discourse analysis, observations and ethnography, case studies, and interviews/focus groups — each with knowledge base exemplars.

Mixed-methods designs

Mixed-methods studies combine quantitative and qualitative strands — often sequentially (e.g., QUAL→QUAN→qual) — so that qualitative data explains or contextualizes quantitative findings. A mixed-method study of GenAI and sustainable learning pairs three-wave surveys with educator interviews; the competence-paradox study uses instructor focus groups, a student survey, and follow-up interviews.

Design-based research (DBR)

DBR iteratively designs, implements, and refines an educational intervention in authentic contexts, cycling between theory, design, and real-world practice. It is prominent in the knowledge base for developing AI learning environments and pedagogical models. See the dedicated Design-Based Research concept page for the full DBR cycle, exemplars, and its strengths/limitations. A canonical AIEd example is the AI-Assisted Collaborative Learning model study (Putra et al.), which ran a four-phase DBR cycle — needs analysis, model design, eight-week classroom implementation, and model refinement — iterating on a four-stage learning cycle (problem identification → AI-assisted collaborative inquiry → collaborative Problem Solving → reflection and presentation). Other exemplars develop AI Literacy teacher training (Development and evaluation of artificial intelligence literacy training for teacher education students) and GenAI Scaffolding for critical thinking (Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher).

DBR trades the causal control of experiments for ecological authenticity and iterative refinement: it is the right tool for "how do we design this AI learning environment to work in practice?" questions, and its evidence is strongest as proof-of-concept and design guidance rather than causal efficacy. Reading DBR learning gains requires the same caution as other designs — without an unassisted, controlled outcome measure, gains can reflect the same AI-inflated-performance confound documented under learning gains.

Systematic reviews and meta-analyses

Reviews synthesize the evidence base rather than running a new experiment. Systematic and scoping reviews apply a transparent protocol to search, screen, appraise, and synthesize a body of studies; meta-analyses additionally pool effect sizes across studies to produce a weighted summary estimate and test moderators. A comprehensive ITS review and a systematic review of GenAI in higher education exemplify the approach.

See the dedicated Meta-Analysis and Systematic Review concept page for a fuller treatment of systematic review and meta-analysis in AI in education — including their relationship to primary designs, PRISMA reporting, and their strengths and limitations.

Computational and benchmark evaluation

Computational evaluation assesses AI systems directly — against benchmarks, ground-truth labels, or human judgments — rather than studying human learners. This includes benchmarks, Large Language Models (LLMs)-as-judge approaches. This is the closest method to AI Ed Evaluation (see the distinction below).

Other designs: longitudinal, case, and simulation studies

Beyond the major families, the knowledge base uses longitudinal designs that track learners over time (a longitudinal LMS study), case and in-the-wild studies of authentic usage (large-scale analysis of real student interactions), and simulation studies in which LLMs stand in for students or patients (LLMs as simulated learners, Simulation). These trade breadth or control for realism and for access to phenomena that are otherwise hard to observe.

Expert-consensus methods: the Delphi technique

The Delphi method is a structured technique for establishing expert consensus on a question where the answer is not yet known empirically — most often used in the knowledge base to develop frameworks, competency lists, and definitions that practitioners and researchers can agree on. In a Delphi study, a panel of experts responds to successive rounds of questionnaires; after each round, an anonymized summary of the group's responses is fed back, and experts revise their answers until the group converges on agreement (typically defined by a pre-set threshold, e.g., 75%). It is a way to build construct validity and professional consensus through iterative, anonymized consultation rather than a single survey or vote.

  • Strengths: produces consensus from a diverse expert panel without in-person group pressures (anonymity reduces dominance effects); well-suited to defining constructs, competencies, and frameworks when no validated measure exists; iterative rounds let experts refine and converge; feasible where full experiments or large samples are impractical.
  • Limitations: consensus reflects expert judgment, not empirical evidence — it establishes agreement, not effect; results depend on panel composition and the (subjective) consensus threshold; can be slow across multiple rounds; a single panel's judgment may not generalize.
  • Exemplars: the SAIL framework study (three rounds, 17 experts, refining AI-literacy competency levels), the HCAP framework study (three rounds, 30 teachers, defining 25 AI-teacher competencies), the AI Literacy Heptagon (which used expert input/consensus alongside a PRISMA-guided review), and.

Delphi is often combined with other methods — for example, expert consensus can be used to validate a framework (as in SAIL and HCAP) that is then tested or implemented via design-based research or survey studies. It sits alongside qualitative and expert-judgment approaches and contributes to the validity of framework-based instruments.

Research vs. evaluation: connections and distinctions

Research and evaluation are closely related but distinct. Research asks generalizable questions about how AI affects learning — "does scaffolding improve learning outcomes?" — and aims to build theory and evidence that transfers beyond the specific study. Evaluation (see AI Ed Evaluation) assesses whether a specific AI tool or system works — is accurate, reliable, pedagogically sound, and fit for purpose — against benchmarks, rubrics, or stakeholder-defined criteria. Research emphasizes internal validity and generalization; evaluation emphasizes system quality and local decision-making.

The boundaries blur: benchmark studies are evaluation that can feed research, and evaluation instruments (rubrics, ground-truth sets, validity frameworks) depend on the Educational Measurement and Assessment Validity concerns that research clarifies. Conversely, research findings on what supports learning should inform how AI tools are evaluated. The knowledge base treats them as complementary: computational and benchmark evaluation (Benchmark, AI Ed Evaluation) tells us whether an AI system is technically sound, while efficacy and survey research (RCT) tells us whether it helps people learn.

Choosing among methods

Method choice follows the research question. Causal-effect questions favor experiments (RCT); mechanism and perception questions favor surveys and qualitative work; system-quality questions favor computational evaluation (Benchmark, AI Ed Evaluation); synthesis questions favor reviews and meta-analyses; design questions favor DBR; and questions about what experts agree a construct, competency, or framework should contain favor expert-consensus methods like the Delphi technique. Given the field's heterogeneity and the speed of AI change, the knowledge base's corpus reflects a deliberate move toward triangulation — combining computational evaluation with efficacy, qualitative, and expert-consensus evidence to judge both whether a tool works and whether it helps learning.

Equally important is reading any single study with awareness of the cross-cutting limitations that affect AIED research as a whole — methodological constraints, the fast pace of AI change versus slow publication, reproducibility and FAIR-practice gaps, reliance on proprietary tools, and weak or uncritical theory use. See Limitations in AIEd Research.

Contrasting the major research traditions

The three major research traditions — quantitative, qualitative, and experimental — differ fundamentally in what they can claim, what they sacrifice, and when each is appropriate. Understanding these contrasts is essential for both designing and reading AI-in-education research.

What each tradition establishes

Dimension Quantitative / survey Qualitative Experimental
Core question How much? How related? What does it mean? How is it experienced? Does X cause Y?
Primary data Numbers, scales, self-report Words, observations, artifacts Outcome measures across assigned conditions
Inference target Patterns, correlations, mediation Meaning, mechanisms, categories Causal effects
Internal validity Weak (correlational) Weak (no control) Strong (random assignment)
External validity Strong (large samples) Limited (small, context-bound) Moderate (controlled conditions)
Ecological validity Moderate High Lower (artificial conditions)
  • Quantitative research measures and models relationships among variables — surveys, SEM/PLS-SEM, measurement, longitudinal tracking. It provides breadth, precision, and generalizability but cannot establish causation from cross-sectional data and inherits measurement limitations (including self-report bias).
  • Qualitative research interprets meaning and experience — interviews, focus groups, thematic analysis, grounded theory, phenomenography, discourse analysis, observation/ethnography, case studies. It provides depth, mechanism, and theory-building (see Theory Development in AI in Education) but limited generalizability and weak causal support.
  • Experimental and quasi-experimental designs (see RCT) estimate causal effects via random assignment or matched comparison — the gold standard for internal validity, at the cost of cost, speed, and ecological validity.

Quantitative work depends on Educational Measurement — reliable, valid instruments for the constructs being studied. Qualitative work reveals the mechanisms and meanings those instruments may miss. Experimental work estimates whether an intervention causes the outcomes the instruments measure. The three are complementary layers: instruments quantify constructs, experiments establish causality, and qualitative work explains the how and why behind the numbers.

Mixed-methods designs intentionally combine quantitative and qualitative strands so their strengths offset each other's weaknesses — quantitative breadth plus qualitative depth, with triangulation increasing confidence.

Usability and HCI research

A distinct methodological strand — usability and HCI research — evaluates how users interact with an AI system: its usability, usefulness, learnability, and user experience, using think-aloud protocols, structured user studies, interviews, and observation. It is the closest to AI Ed Evaluation and answers a prerequisite question: even a pedagogically sound tool fails if it is unusable. Usability research shares data-collection methods with qualitative research but aims at evaluating an artifact rather than interpreting meaning.

Benefits and limitations across traditions

  • Quantitative/survey: benefits — large samples, broad coverage, tests complex mediators, efficient. Limitations — no causation, self-report bias, convenience sampling, instruments may measure the wrong construct.
  • Qualitative: benefits — deep insight, surfaces unexpected phenomena and harms, essential for theory-building, centers under-represented voices. Limitations — limited generalizability, researcher dependence, small samples, weak causal support, hard to synthesize.
  • Experimental: benefits — strongest causal inference, clean outcome measurement, effect-size estimation. Limitations — costly/slow, artificial conditions, fast-changing AI dates results, underpowered small samples, ethical constraints.
  • Mixed-methods: benefits — triangulation, breadth + depth, explains unexpected results. Limitations — complex, resource-intensive, integration can be shallow, inherits each strand's weaknesses.
  • Usability/HCI: benefits — identifies adoption barriers, actionable design guidance, fast and cheap. Limitations — does not establish learning effects, small samples, self-report satisfaction can mislead.

In practice, AI-in-education research rarely falls cleanly into one tradition. The strongest evidence triangulates: a computational or usability evaluation establishes that a system works, an experiment establishes that it causes learning, quantitative instruments measure the constructs, and qualitative work reveals the mechanisms and meanings — together answering both whether a tool helps learning and how and why.

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.