On this page

Synthesis: This systematic review of 24 empirical studies maps ChatGPT-supported inquiry-based learning in STEAM education across the five inquiry phases. ChatGPT was primarily used during conceptualization, investigation, and discussion — as learning tool, tutor, learning peer, domain expert, and teaching assistant — improving performance, critical thinking, engagement, and motivation. But learner-level challenges (over-reliance, hallucination, superficial conclusions) and educator/institutional challenges (curricular misalignment, reduced instructional depth, Academic Integrity) persist. An integrated framework synthesizes ChatGPT's roles, advantages, and challenges across the inquiry phases.

Key Findings

  1. 24 empirical studies, uneven disciplinary coverage. Science received the most attention; technology, arts, engineering, and integrated STEAM were underrepresented. Studies split evenly between K-12 and higher education, with undergraduates the largest group and no study focused on elementary learners.
  2. Roles cluster in three of five inquiry phases. ChatGPT was used mainly in conceptualization, investigation, and discussion — supporting question formulation, inquiry design, Problem Solving, and reflection. Its use in the orientation and conclusion phases remains largely unexplored.
  3. Multiple roles. ChatGPT functioned as a learning tool, tutor, learning peer, domain expert, and teaching assistant. For students it enhanced academic performance, critical thinking, engagement, and motivation while supporting individualized learning; for processes it improved interaction, reduced educator workload, and enabled differentiated learning.
  4. Multifaceted challenges. Learners may over-rely on ChatGPT-generated answers, accept inaccurate feedback, experience conceptual confusion, or produce superficial conclusions when AI outputs are treated as authoritative. Hallucinations undermine evidence synthesis in the conclusion phase. Institutional challenges include curricular misalignment, reduced instructional depth, unequal access, privacy concerns, academic-integrity risks, and algorithmic bias.
  5. An integrated framework and four research agendas. The review proposes a framework synthesizing ChatGPT's roles, advantages, and challenges across the five inquiry phases, plus four agendas for future research on responsible, pedagogically aligned integration.

What this means for practice

  • Instructors. Deploy ChatGPT where the reviewed evidence is strongest — conceptualization, investigation, and discussion — for question formulation, inquiry design, Problem Solving, and reflection, rather than assuming the "AI-powered co-inquirer" fits every phase of inquiry.
  • Instructors. Put critical-evaluation checkpoints into the conclusion phase and keep mediating the interaction throughout: the 24 studies report students accepting inaccurate or irrelevant feedback, experiencing conceptual confusion, and reaching superficial conclusions when AI output is treated as authoritative, with hallucinations undermining evidence synthesis.
  • Instructors. Provide the pedagogical Scaffolding the reported advantages depend on — the gains in academic performance, critical thinking, engagement, and motivation were not achieved by unaided access to the tool, and uncritical acceptance risks the Cognitive Offloading and AI Literacy problems the knowledge base documents.
  • Designers. Align ChatGPT use with curricular goals and design sustained teacher mediation into the activity, since misalignment with curricular goals and reduced instructional depth were among the educator-level challenges the review identified.
  • Researchers. Extend the evidence into the orientation and conclusion phases, which remain largely unexplored, and into the technology, arts, engineering, and integrated STEAM settings the review found underrepresented, complementing the empirical problem-posing study at the primary level (Inquiry-Based Learning in STEM Education: The Impact of Generative AI-Based Chatbots on Primary School Students' Problem Posing Ability in Science).

Limitations

  • The review rests on 24 empirical articles with uneven disciplinary coverage: science received the most attention while technology, the arts, engineering, and integrated STEAM remained underrepresented.
  • No reviewed study focused on elementary learners; studies split evenly between K–12 and higher education, with undergraduates the largest single participant group.
  • Reporting clusters in three of the five inquiry phases, so the framework's guidance for the orientation and conclusion phases is extrapolation from little direct evidence.
  • The review is a secondary synthesis of heterogeneous designs and does not model effect sizes, so its advantages and challenges are descriptive patterns rather than measured magnitudes.

Citation

Jiang, H., Li, Y., Chugh, R., Xin, Y., & Cheng, H. (2026). The AI-powered co-inquirer: a systematic review of ChatGPT for inquiry-based learning in STEAM education. International Journal of STEM Education, 13, 42.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.