Concept
Research Methods in AIED
Research methods in AIED — the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study AI in education: whether and how AI tools support (or harm) learning, and under what conditions. The knowledge base's corpus spans experimental, survey, qualitative, design-based, computational-benchmark, and review methods. Each has distinct strengths and limitations, and choosing among them involves trade-offs among internal validity (confidence in causal claims), external validity (generalizability), ecological validity (real-world authenticity), and the feasibility of studying fast-moving AI tools.
Questions to Consider
- The page's central tension: the strongest designs for causal inference (randomized experiments) are the hardest to run in real classrooms, while the most authentic settings offer weaker causal control. If you had to decide whether an AI tutor helps learning, which of these two failures would you rather live with — and why?
- Before you read, can you name the difference between internal, external, and ecological validity? The page argues every design trades these off. How might a study that's rigorously causal still tell you almost nothing useful about a real classroom?
- A benchmark shows an AI scores high on accuracy, but the page insists high benchmark accuracy does not entail educational effectiveness. Why might a system that 'passes the test' still fail to help students learn — and what kind of evidence is missing?
- Design-based research iterates on a real intervention but can't attribute gains to a specific mechanism, while an RCT isolates causes but runs in artificial conditions. Given the fast pace of AI change, how long do you think a rigorous RCT remains relevant before the tool it tested is obsolete?
- Delphi expert consensus establishes agreement among experts, not empirical effect. When is it legitimate to build a competency framework from what experts believe, versus from data about what works — and how would you tell the difference in practice?
- The page advocates triangulation — combining benchmark evaluation, experiments, measurement, and qualitative work to judge both whether a tool works and how. Before you read, where in a claim like 'this AI improves learning' would each method be needed to make you confident?
Introduction
The central tension in AIED research is that the strongest designs for causal inference — randomized experiments — are often the hardest to run with authentic AI tools in real classrooms, while the most authentic settings (field deployments, case studies, log-data analyses) offer weaker causal control. No single method resolves this; the field advances by triangulating across methods, and by being explicit about what kind of claim each design can support. Every method also carries cross-cutting limitations — generalizability, measurement validity, the fast pace of AI change, reproducibility, and weak theory use — that readers must weigh; see Limitations in AIEd Research.
The page's subject is method rather than findings. Learning Sciences is the substantive field these methods serve: where this page covers how a study should be designed, measured and reported, that page covers what the field has established about how people learn and how learning environments should be designed, and it treats design-based and mixed methods as the learning sciences' signature approaches rather than two options among many.
Reporting rigor and the TEP-AIED model
The reporting quality of AI-in-education studies is itself a research concern. The TEP-AIED model (Hwang, Xie, Wah & Gasevic, 2026) offers a structured framework for presenting AI-in-education research with rigor, organizing the essential components of a study report — theory/technology/educational problem framing, design, data, analysis, and results — so that readers and reviewers can assess whether claims are supported and whether the work is reproducible. It responds to the field's chronic weaknesses in reporting (vague tool descriptions, unstated model versions, omitted evaluation details) that the limitations page documents. Reporting frameworks like TEP-AIED sit alongside established reporting checklists (e.g., CONSORT-style guidance for trials, PRISMA-style guidance for reviews) as part of the field's broader move toward methodological transparency and reproducibility.
RAISE (Allison, 2026) approaches the same problem from the opposite direction — as a checklist rather than a narrative structure. It sets out 30 items across ten thematic domains (educational justification and theoretical grounding, AI system specification, AI role and interaction, Accessibility and cultural fit, setting and participants, human involvement, study design and evaluation, ethics and trustworthiness, transparency and reproducibility, and limitations and implications), with an editable version and a companion Ethics and Risk Matrix covering learner agency, equity of access, data AI Governance and algorithmic transparency. The two frameworks address each other directly: TEP-AIED characterizes RAISE as comprehensive but faults its breadth, arguing that "its breadth and granularity may make it complex and less accessible for routine empirical applications," while RAISE's own framing is that it mandates no method or model and only requires that choices be made visible. Read together they mark the trade-off in this literature — the fuller audit checklist versus the leaner three-dimension narrative — and they converge on the same non-negotiables: name and version the AI system, disclose prompts and interaction design, define treatment and comparison conditions, report ethical review and risk mitigation, and state whether outcomes measure performance, retention, or transfer. For a study to be adjudicated on those terms, the reporting instrument has to be adopted at design time rather than assembled at manuscript stage, which is the point both frameworks insist on. Corpus-level historical analysis is itself a methodological choice with transparency obligations: Rismanchian & Doroudi locate each paper in their AI×Ed framework based on author judgment of abstracts and full texts, explicitly acknowledge that this is not a systematic or scalable data-driven categorization, and make their full dataset publicly available as supplementary material for replication.
Experimental and quasi-experimental designs
An efficacy study tests whether an intervention produces its intended learning effect, typically using experimental or quasi-experimental designs that compare outcomes with and without the intervention. Experiments randomly assign learners to conditions (e.g., AI tutor vs. human tutor, or AI-scaffolded vs. unassisted) to estimate causal effects on outcomes like learning gains, engagement, or motivation. Randomized controlled trials are the gold standard for internal validity. A randomized field study of human support plus AI tutoring and an RCT on generative AI in teaching use assignment to isolate causal effects. Quasi-experimental designs (pre/post, between-subjects, or matched groups without randomization) are more feasible in intact classrooms but weaker on causal claims.
- Strengths: strongest causal inference; clean outcome measurement; supports effect-size estimation and efficacy claims.
- Limitations: costly and slow; artificial conditions can reduce ecological validity; fast-changing AI tools make long experiments date quickly; small samples often underpower detection of meaningful effects; ethical constraints on withholding potentially helpful tools.
- Exemplars: Access is Not Enough: Human Support Improves Engagement with AI Tutoring, Generative AI Can Harm Teaching, Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention Study, Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning, From AI Use to Critical Thinking Among Medical Students: A Moderated Mediation Perspective on Cognitive Load and Self-Regulated Learning.
Survey and structural-equation-modeling studies
Cross-sectional surveys measure self-reported attitudes, perceptions, motivation, Self-Efficacy, and technology acceptance, often modeled with regression or structural equation modeling (SEM/PLS-SEM) to test hypothesized relationships and mediators. These dominate the knowledge base's corpus, particularly for acceptance, motivation, and psychological-mechanism questions.
- Strengths: large samples; broad, low-cost coverage; can test complex mediational models of psychological mechanisms; feasible for studying attitudes that are hard to observe.
- Limitations: cross-sectional data cannot establish causation; common-method/self-report bias; convenience sampling limits generalizability; mediators inferred from covariance, not manipulation.
The instrument itself deserves separate scrutiny. What a questionnaire, interview, or diary can and cannot establish — and the documented gap between what people report and what they do — is gathered on Self-Report Measures.
- Exemplars: Acceptance of AI-Assisted English Language Learning Tools in Higher Education: Psychological Correlates Across Disciplinary and Proficiency Groups, Examining the Impact of Generative AI on Student Motivation and Engagement: The Mediating Role of Autonomy-Support and Autonomous Motivation in Education, AI-Assisted Autonomous Learning and Reduced Academic Accomplishment in Vocational Higher Education: The Mediating Role of Hardiness, From Enhancement to Over-Reliance: A Mixed-Method Study of Generative AI and Sustainable Learning Performance, From AI Use to Critical Thinking Among Medical Students: A Moderated Mediation Perspective on Cognitive Load and Self-Regulated Learning.
Qualitative methods
Interviews, focus groups, and thematic analysis produce rich, contextual accounts of how students and teachers experience AI tools, the meanings they attach to them, and the tensions and harms that standardized measures miss. Research on AI tutor safety and how AI changes teaching workflows rely heavily on qualitative evidence. See the dedicated Qualitative Research concept page for the full treatment of qualitative approaches — thematic analysis, grounded theory, phenomenology/phenomenography, discourse analysis, observations and ethnography, case studies, and interviews/focus groups — each with knowledge base exemplars.
- Strengths: deep ecological and conceptual insight; surfaces unexpected phenomena, risks, and mechanisms; essential for theory-building and for studying contested constructs like trust, autonomy, and authorship.
- Limitations: limited generalizability; interpretive and researcher-dependent; small samples; weaker support for causal claims; findings can be hard to synthesize across studies.
- Exemplars: SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems, How AI Is Changing Teaching Workflows, Scaffolding critical engagement with GenAI: Transforming ethnic minority preparatory students' collaborative discourse.
Mixed-methods designs
Mixed-methods studies combine quantitative and qualitative strands — often sequentially (e.g., QUAL→QUAN→qual) — so that qualitative data explains or contextualizes quantitative findings. A mixed-method study of GenAI and sustainable learning pairs three-wave surveys with educator interviews; the competence-paradox study uses instructor focus groups, a student survey, and follow-up interviews.
- Strengths: triangulation increases confidence; quantitative breadth plus qualitative depth; can explain unexpected results and bridge mechanism and magnitude.
- Limitations: complex, resource-intensive, and methodologically demanding; integration can be shallow if not carefully designed; still inherits the weaknesses of each strand (e.g., self-report).
- Exemplars: From Enhancement to Over-Reliance: A Mixed-Method Study of Generative AI and Sustainable Learning Performance, The Competence Paradox: Negotiating Ease, Risk, and Creative Identity in Text-to-Image Generative AI Use Among Art and Design Students, Same AI, different pathways: Unpacking mechanisms of AI-mediated learning across discipline-institution contexts, Trust-utility gap in introductory physics education: Students' adoption, domain-specific skepticism, and preferences for AI integration.
Design-based research (DBR)
DBR iteratively designs, implements, and refines an educational intervention in authentic contexts, cycling between theory, design, and real-world practice. It is prominent in the knowledge base for developing AI learning environments and pedagogical models. See the dedicated Design-Based Research concept page for the full DBR cycle, exemplars, and its strengths/limitations. A canonical AIEd example is the AI-Assisted Collaborative Learning model study (Putra et al.), which ran a four-phase DBR cycle — needs analysis, model design, eight-week classroom implementation, and model refinement — iterating on a four-stage learning cycle (problem identification → AI-assisted collaborative inquiry → collaborative Problem Solving → reflection and presentation). Other exemplars develop AI Literacy teacher training (Development and evaluation of artificial intelligence literacy training for teacher education students) and GenAI Scaffolding for critical thinking (Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher).
- Strengths: high ecological validity and practical relevance; produces both usable artifacts and theory; responsive to the complexity of real classrooms and evolving AI tools; well-suited to developing a model and refining it based on authentic implementation evidence.
- Limitations: weak internal validity (few/no control groups); findings are context-bound and hard to generalize; long timelines; difficult to isolate which design element caused an outcome — DBR demonstrates feasibility and improvement but cannot attribute learning gains to a specific mechanism.
- Exemplars: Design-Based Research for Developing an AI-Assisted Collaborative Learning Model to Enhance Critical Thinking and Problem-Solving Skills in Higher Education, Development and evaluation of artificial intelligence literacy training for teacher education students, Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher, Human-centered AI for teacher educators: Designing professional learning for critical AI literacy.
DBR trades the causal control of experiments for ecological authenticity and iterative refinement: it is the right tool for "how do we design this AI learning environment to work in practice?" questions, and its evidence is strongest as proof-of-concept and design guidance rather than causal efficacy. Reading DBR learning gains requires the same caution as other designs — without an unassisted, controlled outcome measure, gains can reflect the same AI-inflated-performance confound documented under learning gains.
Systematic reviews and meta-analyses
Reviews synthesize the evidence base rather than running a new experiment. Systematic and scoping reviews apply a transparent protocol to search, screen, appraise, and synthesize a body of studies; meta-analyses additionally pool effect sizes across studies to produce a weighted summary estimate and test moderators. A comprehensive ITS review and a systematic review of GenAI in higher education exemplify the approach.
- Strengths: efficient synthesis of a large, fragmented literature; meta-analysis yields pooled effect estimates and detects moderators; essential for evidence-based practice and identifying gaps.
- Limitations: depend on the quality of included studies (garbage-in/garbage-out); publication bias; heterogeneous methods and outcome measures make synthesis hard; rapidly aging given the speed of AI change.
- Meta-research caveat (2026): critiques of the AIED synthesis base show that many early meta-analyses are undermined by construct incoherence, unresolved heterogeneity, unaddressed dependence among effect sizes, and invalid publication-bias assessment — inflating headline AI effect sizes (see Effect of Artificial Intelligence on Learning: A Meta-Meta-Analysis, Presumed Effective: The Manufacturing of an Evidence Base for AI-in-Education Through Flawed Meta-Analysis, and ChatGPT in Education: An Effect in Search of a Cause). Treat pooled AIED effect sizes as upper bounds.
- Exemplars: Comprehensive Review of Intelligent Tutoring Systems, Generative AI in Higher Education: A Systematic Review of Opportunities, Challenges, and Pedagogical Innovations (2022–2025), The cognitive impact of ChatGPT in higher education: A systematic review of critical and creative thinking outcomes, Comprehensive Review of Intelligent Tutoring Systems, A scoping review of generative AI-powered agentic AI in education: Research landscape, agentic capabilities.
See the dedicated Meta-Analysis and Systematic Review concept page for a fuller treatment of systematic review and meta-analysis in AI in education — including their relationship to primary designs, PRISMA reporting, and their strengths and limitations.
Computational and benchmark evaluation
Computational evaluation assesses AI systems directly — against benchmarks, ground-truth labels, or human judgments — rather than studying human learners. This includes benchmarks, Large Language Models (LLMs)-as-judge approaches. This is the closest method to AI Ed Evaluation (see the distinction below).
- Strengths: fast, scalable, reproducible; enables head-to-head comparison of models and system versions; essential for system development and quality assurance.
- Limitations: measures system output, not learning — high benchmark accuracy does not entail educational effectiveness; ground-truth and rubric quality are themselves contested; can miss pedagogical quality that humans perceive. Rismanchian & Doroudi argue that LLMs' natural-language flexibility makes purely technical metrics insufficient, requiring human-inspired evaluation approaches — simulated students, AI-teacher tests, and behavioral-science analyses previously reserved for human subjects — to judge learning-relevant quality, and that studying LLMs cautiously can generate insight into human learning.
- Exemplars: TeachBench - Evaluating LLM Teaching Ability, ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents, Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education, Confidence Estimation in Automatic Short Answer Grading with LLMs, The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors.
- Reporting standards for automated pipelines, and audited benchmarks. Two 2026 papers extend methodological accountability beyond the study itself. PRISMA-LLM maps 888 review-automation papers and 14,726 annotations, finding that 38.0% of software or product papers reported no evaluation against 9.3% of LLM papers and that 52% of positive-only LLM evaluations left a high-bar concern unmet, and proposes reporting that identifies where in the review workflow automation acted (PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews). An expert re-grading audit of six physics benchmarks shows the same problem at the instrument level: of 250 audited rejections only 12 (4.80%) were genuine model errors, while 143 were item defects and 95 grader errors (How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks). Both argue that computational evaluations need an audited error budget before their results are read as findings about learners or models.
Other designs: longitudinal, case, and simulation studies
Beyond the major families, the knowledge base uses longitudinal designs that track learners over time (a longitudinal LMS study), case and in-the-wild studies of authentic usage (large-scale analysis of real student interactions), and simulation studies in which LLMs stand in for students or patients (LLMs as simulated learners, Simulation). These trade breadth or control for realism and for access to phenomena that are otherwise hard to observe.
Expert-consensus methods: the Delphi technique
The Delphi method is a structured technique for establishing expert consensus on a question where the answer is not yet known empirically — most often used in the knowledge base to develop frameworks, competency lists, and definitions that practitioners and researchers can agree on. In a Delphi study, a panel of experts responds to successive rounds of questionnaires; after each round, an anonymized summary of the group's responses is fed back, and experts revise their answers until the group converges on agreement (typically defined by a pre-set threshold, e.g., 75%). It is a way to build construct validity and professional consensus through iterative, anonymized consultation rather than a single survey or vote.
- Strengths: produces consensus from a diverse expert panel without in-person group pressures (anonymity reduces dominance effects); well-suited to defining constructs, competencies, and frameworks when no validated measure exists; iterative rounds let experts refine and converge; feasible where full experiments or large samples are impractical.
- Limitations: consensus reflects expert judgment, not empirical evidence — it establishes agreement, not effect; results depend on panel composition and the (subjective) consensus threshold; can be slow across multiple rounds; a single panel's judgment may not generalize.
- Exemplars: the SAIL framework study (three rounds, 17 experts, refining AI-literacy competency levels), the HCAP framework study (three rounds, 30 teachers, defining 25 AI-teacher competencies), the AI Literacy Heptagon (which used expert input/consensus alongside a PRISMA-guided review), and.
Delphi is often combined with other methods — for example, expert consensus can be used to validate a framework (as in SAIL and HCAP) that is then tested or implemented via design-based research or survey studies. It sits alongside qualitative and expert-judgment approaches and contributes to the validity of framework-based instruments.
Research vs. evaluation: connections and distinctions
Research and evaluation are closely related but distinct. Research asks generalizable questions about how AI affects learning — "does scaffolding improve learning outcomes?" — and aims to build theory and evidence that transfers beyond the specific study. Evaluation (see AI Ed Evaluation) assesses whether a specific AI tool or system works — is accurate, reliable, pedagogically sound, and fit for purpose — against benchmarks, rubrics, or stakeholder-defined criteria. Research emphasizes internal validity and generalization; evaluation emphasizes system quality and local decision-making.
The boundaries blur: benchmark studies are evaluation that can feed research, and evaluation instruments (rubrics, ground-truth sets, validity frameworks) depend on the Educational Measurement and Assessment Validity concerns that research clarifies. Conversely, research findings on what supports learning should inform how AI tools are evaluated. The knowledge base treats them as complementary: computational and benchmark evaluation (Benchmark, AI Ed Evaluation) tells us whether an AI system is technically sound, while efficacy and survey research (RCT) tells us whether it helps people learn.
Choosing among methods
Method choice follows the research question. Causal-effect questions favor experiments (RCT); mechanism and perception questions favor surveys and qualitative work; system-quality questions favor computational evaluation (Benchmark, AI Ed Evaluation); synthesis questions favor reviews and meta-analyses; design questions favor DBR; and questions about what experts agree a construct, competency, or framework should contain favor expert-consensus methods like the Delphi technique. Given the field's heterogeneity and the speed of AI change, the knowledge base's corpus reflects a deliberate move toward triangulation — combining computational evaluation with efficacy, qualitative, and expert-consensus evidence to judge both whether a tool works and whether it helps learning.
Equally important is reading any single study with awareness of the cross-cutting limitations that affect AIED research as a whole — methodological constraints, the fast pace of AI change versus slow publication, reproducibility and FAIR-practice gaps, reliance on proprietary tools, and weak or uncritical theory use. See Limitations in AIEd Research.
Contrasting the major research traditions
The three major research traditions — quantitative, qualitative, and experimental — differ fundamentally in what they can claim, what they sacrifice, and when each is appropriate. Understanding these contrasts is essential for both designing and reading AI-in-education research.
What each tradition establishes
| Dimension | Quantitative / survey | Qualitative | Experimental |
|---|---|---|---|
| Core question | How much? How related? | What does it mean? How is it experienced? | Does X cause Y? |
| Primary data | Numbers, scales, self-report | Words, observations, artifacts | Outcome measures across assigned conditions |
| Inference target | Patterns, correlations, mediation | Meaning, mechanisms, categories | Causal effects |
| Internal validity | Weak (correlational) | Weak (no control) | Strong (random assignment) |
| External validity | Strong (large samples) | Limited (small, context-bound) | Moderate (controlled conditions) |
| Ecological validity | Moderate | High | Lower (artificial conditions) |
- Quantitative research measures and models relationships among variables — surveys, SEM/PLS-SEM, measurement, longitudinal tracking. It provides breadth, precision, and generalizability but cannot establish causation from cross-sectional data and inherits measurement limitations (including self-report bias).
- Qualitative research interprets meaning and experience — interviews, focus groups, thematic analysis, grounded theory, phenomenography, discourse analysis, observation/ethnography, case studies. It provides depth, mechanism, and theory-building (see Theory Development in AI in Education) but limited generalizability and weak causal support.
- Experimental and quasi-experimental designs (see RCT) estimate causal effects via random assignment or matched comparison — the gold standard for internal validity, at the cost of cost, speed, and ecological validity.
The measurement and mixed-methods links
Quantitative work depends on Educational Measurement — reliable, valid instruments for the constructs being studied. Qualitative work reveals the mechanisms and meanings those instruments may miss. Experimental work estimates whether an intervention causes the outcomes the instruments measure. The three are complementary layers: instruments quantify constructs, experiments establish causality, and qualitative work explains the how and why behind the numbers.
Mixed-methods designs intentionally combine quantitative and qualitative strands so their strengths offset each other's weaknesses — quantitative breadth plus qualitative depth, with triangulation increasing confidence.
Usability and HCI research
A distinct methodological strand — usability and HCI research — evaluates how users interact with an AI system: its usability, usefulness, learnability, and user experience, using think-aloud protocols, structured user studies, interviews, and observation. It is the closest to AI Ed Evaluation and answers a prerequisite question: even a pedagogically sound tool fails if it is unusable. Usability research shares data-collection methods with qualitative research but aims at evaluating an artifact rather than interpreting meaning.
Benefits and limitations across traditions
- Quantitative/survey: benefits — large samples, broad coverage, tests complex mediators, efficient. Limitations — no causation, self-report bias, convenience sampling, instruments may measure the wrong construct.
- Qualitative: benefits — deep insight, surfaces unexpected phenomena and harms, essential for theory-building, centers under-represented voices. Limitations — limited generalizability, researcher dependence, small samples, weak causal support, hard to synthesize.
- Experimental: benefits — strongest causal inference, clean outcome measurement, effect-size estimation. Limitations — costly/slow, artificial conditions, fast-changing AI dates results, underpowered small samples, ethical constraints.
- Mixed-methods: benefits — triangulation, breadth + depth, explains unexpected results. Limitations — complex, resource-intensive, integration can be shallow, inherits each strand's weaknesses.
- Usability/HCI: benefits — identifies adoption barriers, actionable design guidance, fast and cheap. Limitations — does not establish learning effects, small samples, self-report satisfaction can mislead.
In practice, AI-in-education research rarely falls cleanly into one tradition. The strongest evidence triangulates: a computational or usability evaluation establishes that a system works, an experiment establishes that it causes learning, quantitative instruments measure the constructs, and qualitative work reveals the mechanisms and meanings — together answering both whether a tool helps learning and how and why.
Connected Concepts
- Interpreting and Applying AIEd Research
- AI Ed Evaluation
- RCT
- Benchmark
- Meta-Analysis and Systematic Review
- Educational Measurement
- Assessment Validity
- Simulation
- AI in Education
- Higher Education
- Limitations in AIEd Research
- Learning Gains
- Theory Development in AI in Education — Theory Development in AI in Education
- Qualitative Research — Qualitative Research
- Quantitative Research — Quantitative Research
- Mixed-Methods Research — Mixed-Methods Research
- Design-Based Research — Design-Based Research
- Usability Research — Usability Research
- Self-Report Measures
- Learning Sciences
Connected Articles
- Access is Not Enough: Human Support Improves Engagement with AI Tutoring — Access is Not Enough: Human Support Improves Engagement with AI Tutoring
- Generative AI Can Harm Teaching — Generative AI Can Harm Teaching
- From Enhancement to Over-Reliance: A Mixed-Method Study of Generative AI and Sustainable Learning Performance — From Enhancement to Over-Reliance: A Mixed-Method Study
- Acceptance of AI-Assisted English Language Learning Tools in Higher Education: Psychological Correlates Across Disciplinary and Proficiency Groups — Acceptance of AI-Assisted English Language Learning Tools
- SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems — AI Tutor Safety and Pedagogical Harms
- Comprehensive Review of Intelligent Tutoring Systems — Comprehensive Review of Intelligent Tutoring Systems
- Design-Based Research for Developing an AI-Assisted Collaborative Learning Model to Enhance Critical Thinking and Problem-Solving Skills in Higher Education — Design-Based Research for an AI-Assisted Collaborative Learning Model
- TeachBench - Evaluating LLM Teaching Ability — TeachBench: Evaluating LLM Teaching Ability
- Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education — Modernizing Ground Truth: Four Shifts Toward Reliability and Validity
- Can LLMs Effectively Simulate Human Learners? Teachers' Insights from Tutoring LLM Students — Can LLMs Effectively Simulate Human Learners?
- RAISE the Standard: A Framework for Transparent Reporting of Artificial Intelligence Studies in Education — RAISE: 30 items in ten domains for transparent reporting of AI-in-education studies (Allison 2026)
- AI-Integrated Learning Management System for Middle School: A Longitudinal Study of Learning Outcomes Through High — AI-Integrated Learning Management System: A Longitudinal Study
- AI in the Wild: A Large Scale Analysis of Authentic Interactions of College Students with Generative AI — AI in the Wild: Large Scale Analysis of Authentic Interactions
- Same AI, different pathways: Unpacking mechanisms of AI-mediated learning across discipline-institution contexts — Same AI, Different Pathways: Unpacking Mechanisms
- Presenting Your AI in Education Research with Rigor: The TEP-AIED Model — The TEP-AIED model for reporting AI-in-education research with rigor (Hwang, Xie, Wah & Gasevic 2026)
- The Competence Paradox: Negotiating Ease, Risk, and Creative Identity in Text-to-Image Generative AI Use Among Art and Design Students — The Competence Paradox: Text-to-Image GenAI in Art and Design
- The Evolution of Research on AI and Education Across Four Decades: Insights from the AIxEd Framework
- ChatGPT in Education: An Effect in Search of a Cause — ChatGPT in Education: An Effect in Search of a Cause
- Effect of Artificial Intelligence on Learning: A Meta-Meta-Analysis — Meta-meta-analysis of AI effect on learning
- Presumed Effective: The Manufacturing of an Evidence Base for AI-in-Education Through Flawed Meta-Analysis — Presumed Effective: flawed AIED meta-analysis audit
- What Fidelity Metrics Miss: A Structural Check on Synthetic Educational Data — What Fidelity Metrics Miss: A Structural Check on Synthetic Educational Data
Connected FAQs
- What Are Notable Gaps in the Research Literature on AI in Education?
- What Measures and Research Methods Can an Instructor Use to Evaluate AI-Related Interventions?
- How Should AI in Education Research Incorporate Equity, Accessibility, Privacy, Ethics, and Pedagogical Safety?
- What Are Best Practices for Reporting and Interpreting AI in Education Research?