FAQ
What Measures and Research Methods Can an Instructor Use to Evaluate AI-Related Interventions?
Match the method to the claim. If you want to know whether students liked an AI activity, a survey can help. If you want to know whether they learned, use performance measures. If you want to know whether the AI caused an improvement, you need a credible comparison condition and preferably random assignment.
Method options
The Research Methods in AIED page distinguishes several useful options:
- Randomized experiments provide the strongest causal inference.
- Quasi-experimental pre/post or matched-group designs are often more practical in intact classes but support weaker causal claims.
- Qualitative interviews, focus groups, observations, and artifact analysis reveal mechanisms and unexpected experiences.
- Mixed methods combine outcome evidence with explanations of why effects occurred.
- Design-based research is useful when instructors are iteratively developing and refining an intervention in an authentic course.
A manageable classroom evaluation
For a manageable classroom evaluation, a useful minimum is a baseline measure, the intervention, an immediate post-measure, and a later unassisted measure. Wherever possible, include a comparison condition such as existing practice, no AI, unrestricted AI versus scaffolded AI, or two alternative designs. Measure assisted performance and independent learning separately.
The AI Ed Evaluation synthesis recommends outcomes such as unassisted learning gain, delayed retention, transfer to a new task, quality of reasoning, misconceptions, feedback uptake, and subgroup performance. Engagement, satisfaction, AI-use logs, self-efficacy, and perceived usefulness can be valuable secondary measures but should not be treated as substitutes for learning. Instructor workload and time savings are also legitimate implementation outcomes.