Research Article
Effects of an AI-supported inquiry model on AI literacy and authentic performance: A quasi-experimental study with preservice teachers
Synthesis: Effects of an AI-supported inquiry model on AI literacy and authentic performance: A quasi-experimental study with preservice teachers
Key Findings
- A 10-week quasi-experimental study with 95 preservice teachers in an educational research methods course compared two intact classes (experimental n = 52; comparison n = 43). The groups were equivalent at baseline, with no statistically significant pre-intervention differences in age (M = 20.38 vs 20.05), gender (12:40 vs 5:38 male:female), grade level, major, or AI-literacy pretest scores (M = 79.46 vs 80.33).
- The QUEST+AI model structures AI-supported inquiry in five phases: Question, Understand, Engage, Solve, and Teach.
- The experimental group showed higher overall AI Literacy, with small but meaningful gains concentrated in applying AI, AI-supported problem solving, and emotion AI Regulation in Education during AI use (partial η² ≈ .05–.09). Within-group gains ran from M = 79.46 (SD = 12.22) at pretest to M = 84.94 (SD = 8.27) at posttest (t(51) = 3.66, p < .001), whereas the comparison group showed no statistically reliable change.
- No clear group differences were found for more concept-focused or evaluative dimensions of AI literacy (concepts, detection, Ethics, creation, and persuasion).
- The experimental group earned higher scores on the final research proposal (partial η² = .137), indicating a moderate advantage in authentic performance: after covariate adjustment, adjusted means were 84.14 vs 80.72 (F(1, 90) = 14.24, p < .001), a difference of 3.42 points (95% CI [1.62, 5.22]).
Study Design & Method
Grounded in Deweyan Inquiry and the Practical Inquiry model, the study examined the effects of QUEST+AI, an AI-supported inquiry model built around five phases: Question, Understand, Engage, Solve, and Teach. Both groups received the same in-class instruction, but the experimental group completed two QUEST+AI cycles with coached generative AI use, whereas the comparison group completed conventional homework. Outcomes included a multidimensional AI literacy measure and a capstone research proposal scored with a common rubric. Within-group change was assessed with paired-samples t-tests, and group differences with ANCOVA controlling for pretest AI literacy, gender, and grade. The coached cycles embedded prompt logs and verification routines that made AI-supported decisions visible, and the same inquiry routines mapped directly onto the rubric dimensions used to score the proposals (problem framing, synthesis quality, methodological coherence, and argumentation).
What this means for practice
- Faculty developers. Embed coached GenAI use inside authentic coursework instead of adding stand-alone AI-literacy lectures: the experimental group improved overall AI literacy with no additional direct instruction on the topic (M = 79.46 to 84.94, t(51) = 3.66, p < .001).
- Structure the AI work around the five QUEST+AI phases — Question, Understand, Engage, Solve, and Teach — and run at least two coached cycles so prompt logs and verification routines make AI-supported decisions visible.
- Judge the intervention by disciplinary output rather than tool comfort: after covariate adjustment the experimental group's research proposals scored 84.14 against 80.72, a difference of 3.42 points (95% CI [1.62, 5.22]).
- Add brief targeted activities for the dimensions the cycles did not move — concepts, detection, Ethics, creation, and persuasion — such as concept refreshers with retrieval checks and rubric-scored ethics cases.
- Keep claims proportionate to the effect sizes when making the case for Educational Development investment: gains were small for AI literacy (partial η² ≈ .05–.09) and moderate for authentic performance (partial η² = .137).
Limitations
- Assignment was nonrandom, using two intact classes at a single institution (experimental n = 52; comparison n = 43), so causal inference is limited even though the groups were equivalent at baseline.
- Exposure was brief — two QUEST+AI cycles inside a 10-week study — with no follow-up measure to test whether the gains persisted.
- AI literacy was measured with self-report subscales, several of them short with modest reliability, which may reduce sensitivity to change.
- Authentic performance was judged with a common rubric on capstone research proposals rather than objective performance-based measures triangulated with protocol-adherence data.
Citation
Cao, D., Yan, Y., Xiong, A., & Wicks, D. (2026). Effects of an AI-supported inquiry model on AI literacy and authentic performance: A quasi-experimental study with preservice teachers.