π·οΈ Concept
Research Methods in AIED
Research methods in AIED β the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study AI in education: whether and how AI tools support (or harm) learning, and under what conditions. The wiki's corpus spans experimental, survey, qualitative, design-based, computational-benchmark, and review methods. Each has distinct strengths and limitations, and choosing among them involves trade-offs among internal validity (confidence in causal claims), external validity (generalizability), ecological validity (real-world authenticity), and the feasibility of studying fast-moving AI tools.
The central tension in AIED research is that the strongest designs for causal inference β randomized experiments β are often the hardest to run with authentic AI tools in real classrooms, while the most authentic settings (field deployments, case studies, log-data analyses) offer weaker causal control. No single method resolves this; the field advances by triangulating across methods, and by being explicit about what kind of claim each design can support.
Experimental and quasi-experimental designs
Experiments randomly assign learners to conditions (e.g., AI tutor vs. human tutor, or AI-scaffolded vs. unassisted) to estimate causal effects on outcomes like learning gains, engagement, or motivation. Randomized controlled trials are the gold standard for internal validity. A randomized field study of human support plus AI tutoring and an RCT on generative AI in teaching use assignment to isolate causal effects. Quasi-experimental designs (pre/post, between-subjects, or matched groups without randomization) are more feasible in intact classrooms but weaker on causal claims.
Survey and structural-equation-modeling studies
Cross-sectional surveys measure self-reported attitudes, perceptions, motivation, self-efficacy, and technology acceptance, often modeled with regression or structural equation modeling (SEM/PLS-SEM) to test hypothesized relationships and mediators. These dominate the wiki's corpus, particularly for acceptance, motivation, and psychological-mechanism questions.
Qualitative methods
Interviews, focus groups, and thematic analysis produce rich, contextual accounts of how students and teachers experience AI tools, the meanings they attach to them, and the tensions and harms that standardized measures miss. Research on AI tutor safety and how AI changes teaching workflows rely heavily on qualitative evidence.
Mixed-methods designs
Mixed-methods studies combine quantitative and qualitative strands β often sequentially (e.g., QUALβQUANβqual) β so that qualitative data explains or contextualizes quantitative findings. A mixed-method study of GenAI and sustainable learning pairs three-wave surveys with educator interviews; the competence-paradox study uses instructor focus groups, a student survey, and follow-up interviews.
Design-based research (DBR)
DBR iteratively designs, implements, and refines an educational intervention in authentic contexts, cycling between theory, design, and real-world practice. It is prominent in the wiki for developing AI learning environments and pedagogical models.
Systematic reviews and meta-analyses
Reviews synthesize the evidence base. Systematic and scoping reviews map and appraise the literature; meta-analyses pool effect sizes across studies. A comprehensive ITS review and a systematic review of GenAI in higher education exemplify the approach.
Computational and benchmark evaluation
Computational evaluation assesses AI systems directly β against benchmarks, ground-truth labels, or human judgments β rather than studying human learners. This includes benchmarks, grading accuracy, teaching-ability evaluation, and LLM-as-judge approaches. This is the closest method to AI Ed Evaluation (see the distinction below).
Other designs: longitudinal, case, and simulation studies
Beyond the major families, the wiki uses longitudinal designs that track learners over time (a longitudinal LMS study), case and in-the-wild studies of authentic usage (large-scale analysis of real student interactions), and simulation studies in which LLMs stand in for students or patients (LLMs as simulated learners, Simulation). These trade breadth or control for realism and for access to phenomena that are otherwise hard to observe.
Expert-consensus methods: the Delphi technique
The Delphi method is a structured technique for establishing expert consensus on a question where the answer is not yet known empirically β most often used in the wiki to develop frameworks, competency lists, and definitions that practitioners and researchers can agree on. In a Delphi study, a panel of experts responds to successive rounds of questionnaires; after each round, an anonymized summary of the group's responses is fed back, and experts revise their answers until the group converges on agreement (typically defined by a pre-set threshold, e.g., 75%). It is a way to build construct validity and professional consensus through iterative, anonymized consultation rather than a single survey or vote.
Delphi is often combined with other methods β for example, expert consensus can be used to validate a framework (as in SAIL and HCAP) that is then tested or implemented via design-based research or survey studies. It sits alongside qualitative and expert-judgment approaches and contributes to the validity of framework-based instruments.
Research vs. evaluation: connections and distinctions
Research and evaluation are closely related but distinct. Research asks generalizable questions about how AI affects learning β "does scaffolding improve learning outcomes?" β and aims to build theory and evidence that transfers beyond the specific study. Evaluation (see AI Ed Evaluation) assesses whether a specific AI tool or system works β is accurate, reliable, pedagogically sound, and fit for purpose β against benchmarks, rubrics, or stakeholder-defined criteria. Research emphasizes internal validity and generalization; evaluation emphasizes system quality and local decision-making.
The boundaries blur: benchmark studies are evaluation that can feed research, and evaluation instruments (rubrics, ground-truth sets, validity frameworks) depend on the Educational Measurement and Assessment Validity concerns that research clarifies. Conversely, research findings on what supports learning should inform how AI tools are evaluated. The wiki treats them as complementary: computational and benchmark evaluation (Benchmark, AI Ed Evaluation) tells us whether an AI system is technically sound, while efficacy and survey research (Efficacy Study, RCT) tells us whether it helps people learn.
Choosing among methods
Method choice follows the research question. Causal-effect questions favor experiments (RCT); mechanism and perception questions favor surveys and qualitative work; system-quality questions favor computational evaluation (Benchmark, AI Ed Evaluation); synthesis questions favor reviews and meta-analyses; design questions favor DBR; and questions about what experts agree a construct, competency, or framework should contain favor expert-consensus methods like the Delphi technique. Given the field's heterogeneity and the speed of AI change, the wiki's corpus reflects a deliberate move toward triangulation β combining computational evaluation with efficacy, qualitative, and expert-consensus evidence to judge both whether a tool works and whether it helps learning.