On this page

Science education — the study and practice of how students learn science and how to teach it, now being reshaped by generative AI (LLMs, simulations, virtual labs, and AI grading) across physics, chemistry, and biology. The science-education articles in this knowledge base reveal a field negotiating a core tension: AI demonstrably supports inquiry, misconception correction, and assessment at scale, yet its value depends on instructional design, and it carries real risks of over-reliance, hallucination, and dehumanized, profit-driven learning.

Questions to Consider

  • The page opens with a tension: AI demonstrably supports inquiry and misconception correction at scale, yet its value depends entirely on instructional design. Before you read, when has a tool you used in science teaching seemed powerful yet pedagogically hollow — and what made the difference?
  • An AI-facilitated physics inquiry 'invented' mass values that no one had measured when students pressured it. What does this incident reveal about treating AI outputs as authoritative in science, and how should students be trained to treat a confident-sounding answer?
  • Research suggests AI may currently add more value as a scalable content generator than as an interactive tutor — in one study, well-crafted conceptual-change texts (expert or AI-made) beat a prompted AI dialogue. Does that surprise you, and what does it suggest about where the real pedagogical payoff of AI in science lies right now?
  • AI graded 10,364 handwritten physics assessments with high agreement with human graders, yet multimodal models still show gaps on complex reasoning and diagram generation. What is it about 'authentic, representation-heavy' science work that might resist automated assessment?
  • Physics students show a striking trust–utility gap — high use, low trust — and split into pragmatic users and skeptical nonusers. If students themselves are domain-calibrated skeptics, what does that imply for one-size-fits-all policies about AI in science classrooms?
  • A critical voice warns of 'AI colonization' of science education — extraction without consent, algorithmic monoculture, dehumanized reform. Before you read, where do you think the benefits of AI in science education are being oversold, and what would a human-centered, justice-oriented alternative look like?

Introduction

Science education is where AI's promise and its limits collide most visibly, because the disciplines demand rigorous Multimodal reasoning — visual-spatial thinking in physics, laboratory skills in chemistry, and organism-level systems in biology — alongside well-structured, verifiable content that LLMs handle well. The eighteen articles synthesized here span all three disciplines and levels, from K 12 secondary classrooms to Higher Ed university courses and pre-service teacher programs. Across them, a consistent picture emerges: AI functions best not as an answer generator but as an embedded partner — a co-inquirer, a content generator, a virtual lab assistant — whose contribution is decided by instructional design and the surrounding pedagogical structure.

Virtual labs and simulations

AI is dramatically lowering the barrier to creating customized, embodied science instrumentation. Levy et al. show that a four-element natural-language prompt can generate a hand-controlled augmented-reality physics simulation (a pinch-and-spread gesture tunes a virtual lamp's wavelength), with 86% of pilot students reporting greater engagement. Similarly, Suñer et al. generated a browser-based rotation laboratory entirely through prompting, validated against video analysis to better than 1%. These works reframe generative AI as a programming tool that lets teachers design software around pedagogical intent rather than adapting activities to fixed apps — advancing embodied and Simulation-based learning. In Chemistry Education, Abdikayumova & Madybekova embedded PhET simulations and ChatGPT tutoring within a context-based 7E inquiry cycle for Grade 10 students, achieving significantly higher achievement and engagement than either component alone — evidence that contextualization, structured inquiry, and adaptive AI act synergistically.

AI tutors and inquiry-based learning

Inquiry-based learning is a central theme. Jiang et al.'s systematic review of 24 studies positions ChatGPT as an "AI-powered co-inquirer" used mainly in the conceptualization, investigation, and discussion phases of STEAM inquiry, improving performance, critical thinking, and engagement — yet risks over-reliance, hallucination, and superficial conclusions when outputs are treated as authoritative. Aydın found that eight weeks of AI-supported guided inquiry produced large gains in pre-service teachers' conceptual understanding of photosynthesis and respiration, while AI literacy and computational thinking showed no significant change — suggesting disciplinary learning can improve even when broader competencies need longer or more explicit instruction. Tufino & Damiani show AI can scaffold the epistemic core of ISLE inquiry while remaining confined to the verbal channel, but found facilitation fragile: under student pressure, the AI "invented" mass values that no one had measured, underscoring Hallucination Risk and the need to verify rather than assume AI fidelity. Kuhn et al. propose the AIRIS framework (Activate–Inquire–Reflect) to keep prediction, interpretation, and evaluation non-delegable human tasks, and call for "withdrawal condition" experiments testing whether learning survives AI's removal — a direct response to the "boiling frog problem" of eroding epistemic practice.

Misconceptions and conceptual change

On conceptual change, the evidence is nuanced and partly counterintuitive. Akdoğan found in a 413-student Solomon Four-Group design that both expert-written and AI-generated conceptual change texts were equally and significantly more effective at reducing heat-and-temperature Misconceptions than interactive AI dialogue, which offered no advantage over control — positioning AI's current pedagogical value as a scalable content generator rather than an interactive tutor, with gains concentrated among high-achieving students (an equity concern). This is reinforced by Borse et al., who show that prompt specificity shapes physics-solution quality and that MAPS-guided critique better prepares students to evaluate AI output than independent problem solving, treating AI fallibility as a Critical Thinking learning resource.

AI grading and assessment

AI grading is advancing rapidly on high-stakes work. Pathak et al. graded 10,364 scanned pages of handwritten physics assessments (including a national Olympiad and team-selection camp), achieving high score correlations (0.91–0.97) and recovering the same five-student team as human grading — arguing AI works as a valid second reader and audit tool under examiner control when given detailed, physics-specific rubrics. Yet Chen et al.'s OmniPhys Benchmark reveals that multimodal LLMs still show significant gaps on complex reasoning and diagram generation, cautioning against over-trusting automated physics assessment on the authentic, representation-heavy tasks that define deep disciplinary competence.

Teacher perceptions and the critical view

Teacher readiness is decisive. Amponsah et al. found Ghanaian pre-service science teachers hold positive attitudes and strong intentions toward AI but only moderate actual classroom use — an intention–use gap pointing to Teacher AI Competency and institutional support as the real levers. Becker et al. and Fouad & Bentley document that physics students are domain-calibrated skeptics, not uncritical adopters: a 50-point trust-utility gap (91% use, 41% trust) and two distinct user profiles (Pragmatic Users vs. Skeptical Non-Users) that challenge one-size-fits-all policies. Counterbalancing the techno-optimism, Avraamidou warns of an "AI colonization" of science education — extraction without consent, algorithmic monoculture, and dehumanized, profit-centered reform — and calls for a feminist, human-centered AI prioritizing justice over profit, while Erümit & Özdemir Sarıalioğlu's systematic review of 18 studies emphasizes ethical risks (bias, hallucination, plagiarism, erosion of independent thinking) and the need for teacher training and conscious use. The collective lesson across all eighteen articles is that AI in science education delivers gains when embedded in sound inquiry, Scaffolding, and Assessment design — and undercuts learning when it displaces the epistemic work students must do themselves.

Connected Concepts

Connected Articles