Research Article
Exploring the Capacity of Large Language Models to Simulate Students' Scientific Thinking: Insights for Responsive Teaching
Synthesis: Can large language models stand in for a teacher's students? This two-phase study generated 8,820 simulated ideas with six models from 49 NGSS-aligned lesson plans and compared them with real-lesson ideas and with what six teachers saw. Overlap was moderate and mostly appropriate, but the models drifted most for younger students, producing reasoning that was too broad, too technical, and too certain. Teachers found them useful when the ideas triggered their own sensemaking, positioning Simulating Students as a starting point for instructional planning rather than a finished product. The authors frame the work as evidence for Generative AI-supported responsive science teaching and teacher sensemaking, not a replacement.
Key Findings
- Six LLMs generated 30 ideas per lesson from 49 lesson plans across three domains and grade bands: 1,470 per model, 8,820 total (lesson plans averaged 28, SD = 10).
- Simulated ideas matched lesson ideas at mean cosine similarity 0.52 (SD = 0.07), inside the 0.42–0.70 band prior work accepts; every model beat Mistral-7B (0.47), Gemini-2.5-Flash highest (0.55).
- Overall 69% of ideas sat at or below the target grade's reading level; Llama-4 best (over 80%), GPT-5-mini worst, exceeding it 0.41 more often (OR = 153, 95% CI [21.2, 1103], p < .001).
- Overall 91% of ideas stayed within expected knowledge scope, GPT-4o and Llama-4 strongest; GPT-5-mini exceeded grade-level boundaries 0.27 more often than either (Δ proportion = −0.27).
- Grade level, not subject, drove most gaps: elementary and middle school ideas matched lesson ideas more closely than high school ideas (β = 0.05, p = .01 both), though within-scope proportions were lower for middle school (OR = 0.14) and elementary (OR = 0.07).
- Divergence coding found 60.54% of ideas reasoned broadly from lesson contexts, 35.71% used more technical vocabulary, 14.63% carried fewer uncertainty markers, 7.48% used analogies.
- Six teachers found the ideas realistic and aligned with lesson objectives, most useful when they sparked sensemaking, but flagged advanced vocabulary and leading prompts.
How the study tested six models against real lesson ideas
Phase one measured overlap, grade-level language, and knowledge scope; phase two asked how teachers judge them. The 49 lesson plans came from nine OpenSciEd units; each prompt supplied lesson context and asked for partial or incorrect student reasoning. Models: GPT-5-mini, GPT-4o, Claude Sonnet 4, Gemini 2.5 Flash, Mistral 7B, Llama-4-17B. Ideas were clustered (threshold = 0.30, similarity ≥ 0.70), scored on Flesch–Kincaid Grade Level, and scope-judged by an LLM-as-a-judge (GPT-5), human-checked on 20% (κ = 0.69). Six teachers (45-minute interviews, 1–15 years' experience) assessed GPT-4o ideas for their own lessons, coded qualitatively.
Similarity, readability, and scope: the three benchmarks
Similarity varied by model (χ²(5) = 218.69, p < .001), four models clustered at 0.54–0.55. Domain did not predict it (χ²(2) = 3.83, p = .15); grade did (χ²(2) = 10.94, p = .004), elementary (0.54) and middle school (0.55) above high school (0.50), β = 0.05, p = .01 both.
Language level was weakest: 69% fell at or below target grade (χ²(5) = 57.44, p < .001), Llama-4 beating Mistral-7B by 0.33 (OR = 67.57, 95% CI [10.33, 443], p < .001). Domain was flat (χ²(2) = 0.51, p = .77), grade was not (χ²(2) = 58.54, p < .001): high school OR = 160 (95% CI [31.1, 819]), middle school OR = 3.16 (95% CI [1.29, 7.74]). A grade-level revision prompt lifted most models to 83.67–95.92%, GPT-5-mini only 55.10%.
Knowledge scope was strongest, with GPT-4o and Llama-4 best; GPT-5-mini exceeded scope 0.27 more often than either and 0.22 more than Gemini-2.5-Flash. Domain did not matter (χ² = 0.06, p = .97), grade did (χ² = 17.18, p < .001): against high school, within-scope proportions fell 0.09 for middle school (OR = 0.14, 95% CI [0.03, 0.68], p = .02) and 0.16 for elementary (OR = 0.07, 95% CI [0.02, 0.33], p < .001); only GPT-5-mini showed a within-model grade effect (χ² = 9.18, p = .01). Models could not represent younger students' reasoning: a fourth-grade lesson drew "atoms move but we couldn't see shape change."
Where simulated ideas diverged from classroom ideas
Ideas overlapped lesson concepts but reasoned less specifically from lesson contexts, evidence, and mechanisms (60.54%, by model 51.02–75.51%). Misconception substance diverged in 17.69% of ideas, technical vocabulary in 35.71%, uncertainty markers fewer than in lesson ideas (14.63%; Llama-4 24.49%, Claude 4 22.45%), analogies more common (7.48%). Table 2 reports n = 294, 49 per model.
What teachers saw in the simulated ideas
Teachers judged outputs on authenticity and usefulness: realistic, aligned with lesson objectives, most useful when they prompted sensemaking; they valued instructional scaffolds for surfacing unexpected student ideas. Two found vocabulary too advanced, Sam and Alyssa found prompts overly leading, and Jan asked for "how/why, not yes/no questions."
What this means for practice
- Instructors. Treat simulated ideas as a rehearsal aid, not a script; correct vocabulary and certainty before use.
- Curriculum and instructional designers. Validate lower-grade output against NGSS grade-band expectations: readability and scope slipped for elementary and middle school.
- Model selection is not one-dimensional. Models strong on similarity can still overshoot grade level; Llama-4, Gemini-2.5-Flash, and GPT-4o were most balanced, GPT-5-mini weak on level and scope.
- Teacher educators. Pair deployment with professional development on refining AI-generated ideas through iterative re-prompting.
- Administrators. Set data-privacy rules for student information fed to LLMs; over-reliance on inaccurate simulations can narrow responsiveness.
Limitations
- Simulated ideas were never tested in real classrooms, so instructional impact is unmeasured.
- The comparison corpus was not student discourse: suggested ideas from pilot data and teacher predictions may not represent students.
- Scope relied on an LLM-as-a-judge (GPT-5) scoring other models, mitigated with a human rater (κ = 0.69) and second judge (κ = 0.64); generation was zero-shot only, and one elementary teacher was interviewed.
- The lineup is superseded: submission 19 November 2025, acceptance 13 May 2026, no data-collection window, so results describe models current then.
Citation
Nguyen, H., & Cao, J. (2026). Exploring the Capacity of Large Language Models to Simulate Students' Scientific Thinking: Insights for Responsive Teaching. Journal of Science Education and Technology.