🧠 AI Ed Wiki

Research methods in AIED β€” the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study AI in education: whether and how AI tools support (or harm) learning, and under what conditions. The wiki's corpus spans experimental, survey, qualitative, design-based, computational-benchmark, and review methods. Each has distinct strengths and limitations, and choosing among them involves trade-offs among internal validity (confidence in causal claims), external validity (generalizability), ecological validity (real-world authenticity), and the feasibility of studying fast-moving AI tools.

The central tension in AIED research is that the strongest designs for causal inference β€” randomized experiments β€” are often the hardest to run with authentic AI tools in real classrooms, while the most authentic settings (field deployments, case studies, log-data analyses) offer weaker causal control. No single method resolves this; the field advances by triangulating across methods, and by being explicit about what kind of claim each design can support.

Experimental and quasi-experimental designs

Experiments randomly assign learners to conditions (e.g., AI tutor vs. human tutor, or AI-scaffolded vs. unassisted) to estimate causal effects on outcomes like learning gains, engagement, or motivation. Randomized controlled trials are the gold standard for internal validity. A randomized field study of human support plus AI tutoring and an RCT on generative AI in teaching use assignment to isolate causal effects. Quasi-experimental designs (pre/post, between-subjects, or matched groups without randomization) are more feasible in intact classrooms but weaker on causal claims.

  • Strengths: strongest causal inference; clean outcome measurement; supports effect-size estimation and efficacy claims.
  • Limitations: costly and slow; artificial conditions can reduce ecological validity; fast-changing AI tools make long experiments date quickly; small samples often underpower detection of meaningful effects; ethical constraints on withholding potentially helpful tools.
  • Exemplars: Access Not Enough AI Tutoring 2026, GenAI Can Harm Teaching RCT 2026, Adaptive Pretesting Retention, Agent Voice Accents K12 Group Learning, AI Use Critical Thinking Medical Students 2026.
  • Survey and structural-equation-modeling studies

    Cross-sectional surveys measure self-reported attitudes, perceptions, motivation, self-efficacy, and technology acceptance, often modeled with regression or structural equation modeling (SEM/PLS-SEM) to test hypothesized relationships and mediators. These dominate the wiki's corpus, particularly for acceptance, motivation, and psychological-mechanism questions.

  • Strengths: large samples; broad, low-cost coverage; can test complex mediational models of psychological mechanisms; feasible for studying attitudes that are hard to observe.
  • Limitations: cross-sectional data cannot establish causation; common-method/self-report bias; convenience sampling limits generalizability; mediators inferred from covariance, not manipulation.
  • Exemplars: Acceptance AI English Tools 2026, GenAI Motivation Engagement 2026, AI Autonomous Learning Accomplishment 2026, GenAI Over Reliance Learning 2026, AI Use Critical Thinking Medical Students 2026.
  • Qualitative methods

    Interviews, focus groups, and thematic analysis produce rich, contextual accounts of how students and teachers experience AI tools, the meanings they attach to them, and the tensions and harms that standardized measures miss. Research on AI tutor safety and how AI changes teaching workflows rely heavily on qualitative evidence.

  • Strengths: deep ecological and conceptual insight; surfaces unexpected phenomena, risks, and mechanisms; essential for theory-building and for studying contested constructs like trust, autonomy, and authorship.
  • Limitations: limited generalizability; interpretive and researcher-dependent; small samples; weaker support for causal claims; findings can be hard to synthesize across studies.
  • Exemplars: AI Tutor Safety Harms, AI Changing Teaching Workflows, AI Education Global Capacity, Scaffolding Critical Engagement GenAI Minority Students.
  • Mixed-methods designs

    Mixed-methods studies combine quantitative and qualitative strands — often sequentially (e.g., QUAL→QUAN→qual) — so that qualitative data explains or contextualizes quantitative findings. A mixed-method study of GenAI and sustainable learning pairs three-wave surveys with educator interviews; the competence-paradox study uses instructor focus groups, a student survey, and follow-up interviews.

  • Strengths: triangulation increases confidence; quantitative breadth plus qualitative depth; can explain unexpected results and bridge mechanism and magnitude.
  • Limitations: complex, resource-intensive, and methodologically demanding; integration can be shallow if not carefully designed; still inherits the weaknesses of each strand (e.g., self-report).
  • Exemplars: GenAI Over Reliance Learning 2026, T2i Competence Paradox 2026, Same AI Different Pathways, Fouad Bentley Trust Utility Gap Physics 2026.
  • Design-based research (DBR)

    DBR iteratively designs, implements, and refines an educational intervention in authentic contexts, cycling between theory, design, and real-world practice. It is prominent in the wiki for developing AI learning environments and pedagogical models.

  • Strengths: high ecological validity and practical relevance; produces both usable artifacts and theory; responsive to the complexity of real classrooms and evolving AI tools.
  • Limitations: weak internal validity; findings are context-bound and hard to generalize; long timelines; difficult to isolate which design element caused an outcome.
  • Exemplars: AI Assisted Collaborative Learning Model Dbr, GenAI Literacy Training Teacher Education Dbr 2026, Critical Thinking GenAI Scaffolding.
  • Systematic reviews and meta-analyses

    Reviews synthesize the evidence base. Systematic and scoping reviews map and appraise the literature; meta-analyses pool effect sizes across studies. A comprehensive ITS review and a systematic review of GenAI in higher education exemplify the approach.

  • Strengths: efficient synthesis of a large, fragmented literature; meta-analysis yields pooled effect estimates and detects moderators; essential for evidence-based practice and identifying gaps.
  • Limitations: depend on the quality of included studies (garbage-in/garbage-out); publication bias; heterogeneous methods and outcome measures make synthesis hard; rapidly aging given the speed of AI change.
  • Exemplars: Zerkouk Comprehensive Review ITS 2025, GenAI Higher Education Systematic Review 2026, Chatgpt Critical Creative Thinking Review, AI Tutor Effectiveness Review, Agentic AI Education Scoping Review.
  • Computational and benchmark evaluation

    Computational evaluation assesses AI systems directly β€” against benchmarks, ground-truth labels, or human judgments β€” rather than studying human learners. This includes benchmarks, grading accuracy, teaching-ability evaluation, and LLM-as-judge approaches. This is the closest method to AI Ed Evaluation (see the distinction below).

  • Strengths: fast, scalable, reproducible; enables head-to-head comparison of models and system versions; essential for system development and quality assurance.
  • Limitations: measures system output, not learning β€” high benchmark accuracy does not entail educational effectiveness; ground-truth and rubric quality are themselves contested; can miss pedagogical quality that humans perceive.
  • Exemplars: Teachbench LLM Teaching Evaluation, Jeon Isd Agent Bench 2026, Ground Truth Reliability AIED, Automatic Short Answer Grading, Educational Vlm Evaluation.
  • Other designs: longitudinal, case, and simulation studies

    Beyond the major families, the wiki uses longitudinal designs that track learners over time (a longitudinal LMS study), case and in-the-wild studies of authentic usage (large-scale analysis of real student interactions), and simulation studies in which LLMs stand in for students or patients (LLMs as simulated learners, Simulation). These trade breadth or control for realism and for access to phenomena that are otherwise hard to observe.

    Expert-consensus methods: the Delphi technique

    The Delphi method is a structured technique for establishing expert consensus on a question where the answer is not yet known empirically β€” most often used in the wiki to develop frameworks, competency lists, and definitions that practitioners and researchers can agree on. In a Delphi study, a panel of experts responds to successive rounds of questionnaires; after each round, an anonymized summary of the group's responses is fed back, and experts revise their answers until the group converges on agreement (typically defined by a pre-set threshold, e.g., 75%). It is a way to build construct validity and professional consensus through iterative, anonymized consultation rather than a single survey or vote.

  • Strengths: produces consensus from a diverse expert panel without in-person group pressures (anonymity reduces dominance effects); well-suited to defining constructs, competencies, and frameworks when no validated measure exists; iterative rounds let experts refine and converge; feasible where full experiments or large samples are impractical.
  • Limitations: consensus reflects expert judgment, not empirical evidence β€” it establishes agreement, not effect; results depend on panel composition and the (subjective) consensus threshold; can be slow across multiple rounds; a single panel's judgement may not generalize.
  • Exemplars: the SAIL framework study (three rounds, 17 experts, refining AI-literacy competency levels), the HCAP framework study (three rounds, 30 teachers, defining 25 AI-teacher competencies), the AI Literacy Heptagon (which used expert input/consensus alongside a PRISMA-guided review), and the Brookings students-and-AI report.
  • Delphi is often combined with other methods β€” for example, expert consensus can be used to validate a framework (as in SAIL and HCAP) that is then tested or implemented via design-based research or survey studies. It sits alongside qualitative and expert-judgment approaches and contributes to the validity of framework-based instruments.

    Research vs. evaluation: connections and distinctions

    Research and evaluation are closely related but distinct. Research asks generalizable questions about how AI affects learning β€” "does scaffolding improve learning outcomes?" β€” and aims to build theory and evidence that transfers beyond the specific study. Evaluation (see AI Ed Evaluation) assesses whether a specific AI tool or system works β€” is accurate, reliable, pedagogically sound, and fit for purpose β€” against benchmarks, rubrics, or stakeholder-defined criteria. Research emphasizes internal validity and generalization; evaluation emphasizes system quality and local decision-making.

    The boundaries blur: benchmark studies are evaluation that can feed research, and evaluation instruments (rubrics, ground-truth sets, validity frameworks) depend on the Educational Measurement and Assessment Validity concerns that research clarifies. Conversely, research findings on what supports learning should inform how AI tools are evaluated. The wiki treats them as complementary: computational and benchmark evaluation (Benchmark, AI Ed Evaluation) tells us whether an AI system is technically sound, while efficacy and survey research (Efficacy Study, RCT) tells us whether it helps people learn.

    Choosing among methods

    Method choice follows the research question. Causal-effect questions favor experiments (RCT); mechanism and perception questions favor surveys and qualitative work; system-quality questions favor computational evaluation (Benchmark, AI Ed Evaluation); synthesis questions favor reviews and meta-analyses; design questions favor DBR; and questions about what experts agree a construct, competency, or framework should contain favor expert-consensus methods like the Delphi technique. Given the field's heterogeneity and the speed of AI change, the wiki's corpus reflects a deliberate move toward triangulation β€” combining computational evaluation with efficacy, qualitative, and expert-consensus evidence to judge both whether a tool works and whether it helps learning.

    Connected Concepts

  • AI Ed Evaluation
  • Efficacy Study
  • RCT
  • Benchmark
  • Educational Measurement
  • Assessment Validity
  • Simulation
  • AI Education
  • Higher Ed
  • Connected Articles

  • Access Not Enough AI Tutoring 2026 β€” Access is Not Enough: Human Support Improves Engagement with AI Tutoring
  • GenAI Can Harm Teaching RCT 2026 β€” Generative AI Can Harm Teaching
  • GenAI Over Reliance Learning 2026 β€” From Enhancement to Over-Reliance: A Mixed-Method Study
  • Acceptance AI English Tools 2026 β€” Acceptance of AI-Assisted English Language Learning Tools
  • AI Tutor Safety Harms β€” AI Tutor Safety and Pedagogical Harms
  • Zerkouk Comprehensive Review ITS 2025 β€” Comprehensive Review of Intelligent Tutoring Systems
  • AI Assisted Collaborative Learning Model Dbr β€” Design-Based Research for an AI-Assisted Collaborative Learning Model
  • Teachbench LLM Teaching Evaluation β€” TeachBench: Evaluating LLM Teaching Ability
  • Ground Truth Reliability AIED β€” Modernizing Ground Truth: Four Shifts Toward Reliability and Validity
  • LLM Student Simulation Teacher Insights β€” Can LLMs Effectively Simulate Human Learners?
  • AI LMS Middle School Longitudinal β€” AI-Integrated Learning Management System: A Longitudinal Study
  • AI In The Wild College β€” AI in the Wild: Large Scale Analysis of Authentic Interactions
  • Same AI Different Pathways β€” Same AI, Different Pathways: Unpacking Mechanisms
  • T2i Competence Paradox 2026 β€” The Competence Paradox: Text-to-Image GenAI in Art and Design