Research Article
Using LLMs to Detect Growth in Computational Thinking in Introductory Physics
Synthesis: Savage, Shanker, Michlitsch & Rebello (2026) investigate using LLMs to evaluate students' written explanations of computational physics problems at scale. Establishing a human-coded baseline grounded in CT literature, they found significant growth in Data Practices and Computational Problem Solving Practices. The LLM successfully mirrored human evaluations for these constructs, but both human raters and the LLM struggled with more complex constructs like Systems Thinking. This work demonstrates that LLMs offer a viable, scalable method for assessing computational thinking in large-enrollment physics courses.
As computation becomes more central to physics education, scalable methods to assess authentic computational thinking (CT) are critically needed. This study establishes a human-coded baseline grounded in CT literature, identifies significant pre/post growth in Data Practices and Computational Problem-Solving Practices, and demonstrates that an LLM can mirror human evaluations — scaling CT assessment across large datasets. Notably, both human raters and the LLM struggled with more complex constructs like Systems Thinking, revealing the limits of current automated approaches.
Key Findings
- Students in a computationally intensive introductory physics course showed large, statistically significant growth in Data Practices (p < 0.001, d = 1.03) and moderate growth in Computational Problem-Solving (d = 0.52) and Systems Thinking (d = 0.45), as coded by human raters grounded in the Weintrop CT taxonomy.
- An LLM mirrored human evaluations for well-defined constructs (Data Practices κ = 0.90, Modeling & Simulation κ = 0.78) and reproduced the same growth trends when scaled across the full 936-student dataset (over 2,800 responses).
- Both human raters and the LLM struggled with multi-component constructs like Systems Thinking (κ ≈ 0.48–0.51), indicating the difficulty stems from construct complexity rather than a model deficiency.
- A ceiling effect on Modeling & Simulation Practices (pre-instruction mean 1.62 out of 2) masked growth on the simulation-design prompt, showing some items lack sensitivity to instructional growth.
Background
The American Association of Physics Teachers frames computation as the "third pillar" of physics alongside experiment and theory. Yet integrating computation into introductory courses means fostering computational thinking (CT) — moving students from superficial programming toward authentic sensemaking — which is difficult to measure. Historically, physics education research relied on multiple-choice instruments that excel at identifying broad performance trends but cannot capture students' mental models and reasoning. Written explanations offer a richer window into cognition, but evaluating open-ended responses at scale is resource-intensive and requires many hours of qualitative coding to reach acceptable inter-rater reliability.
This assessment challenge has intensified as generative AI lets students bypass the cognitive effort of algorithmic design. To evaluate complex written work at scale, the authors turn to recent work showing AI can rate written responses, identify Misconceptions about AI, and assist in generating Feedback. Rather than evaluating code correctness, they deploy a custom-prompted LLM to categorize how participants frame and reason through computational physics problems.
Methods
The assessment was anchored in the taxonomy developed by Weintrop et al., which categorizes CT in math and science into four domains: Data Practices, Computational Problem-Solving Practices, Modeling and Simulation Practices, and Systems Thinking Practices. Systems Thinking — the ability to define system boundaries and synthesize micro-level interactions into macro-level behavior — was integrated into all three open-ended prompts alongside other practices. The instrument was iteratively refined by experts to ensure construct validity, with questions requiring simultaneous application of physics concepts and specific CT practices.
Data came from an introductory calculus-based engineering physics course at a large public university that integrated computation heavily into its curriculum, requiring weekly labs in Python via Jupyter notebooks. The Assessment was administered online via the Brightspace learning management system, proctored to ensure integrity, as a pre-test in Week 1 and post-test in Week 15. A total of 936 students completed both surveys.
Three researchers independently coded an initial subset of 10 students' responses (60 responses), achieving a Fleiss' κ of 0.53 for CT practices, then resolved discrepancies through iterative discussion. Following calibration they coded 50 students (300 responses), establishing human ground truth via majority vote. For LLM validation, the authors used the Multimodal AI GPT-5.4-mini API with a temperature of 0.2, feeding the model the survey image, prompt, and response. A structured prompt instructed the model to act as an expert physics education researcher and to generate structured justifications before assigning a 0–2 score. Every response was evaluated across three independent runs with a final score by majority vote, mirroring the human methodology.
Findings
Human consensus demonstrated substantial reliability on explicit computational tasks: Modeling and Simulation (κ = 0.80), Data Practices (κ = 0.80), Physics Correctness (κ = 0.73), and Computational Problem-Solving (κ = 0.60), while Systems Thinking yielded the lowest agreement (κ = 0.51). The LLM achieved substantial agreement on well-defined practices — Data Practices (κ = 0.90, 93% agreement), Modeling & Simulation (κ = 0.78), and Computational Problem-Solving (κ = 0.69) — but lower agreement on multi-component constructs like Physics Correctness (κ = 0.53) and Systems Thinking (κ = 0.48).
When applied to the full dataset, the LLM confirmed the macroscopic trends, detecting highly significant growth (p < 0.001) on the graph-interpretation and code-tracing items, but no measurable growth on the simulation-design item. Paired t-tests on the human-graded sample showed that students grew most in Data Practices (d = 1.03), aligning with the course's lab design, with moderate gains in Computational Problem-Solving (d = 0.52) and Systems Thinking (d = 0.45). Modeling and Simulation showed no growth (d = 0.00) due to an assessment ceiling: students entered the course already able to list simulation parameters, reflecting high prior knowledge.
What this means for practice
- Instructors. Use a custom-prompted LLM to track CT growth in large-enrollment courses: the model reproduced the human-coded trends when scaled across the full 936-student dataset of over 2,800 responses.
- Assessment professionals. Validate automated scoring against human ground truth construct by construct before scaling it: agreement reached κ = 0.90 for Data Practices but fell to κ = 0.48 for Systems Thinking.
- Assessment professionals. Do not read low automated agreement on an integrated construct as model failure — human raters were no more consistent on Systems Thinking (κ ≈ 0.48–0.51) — and instead rewrite the rubric to separate identifying system components from explaining their interactions.
- Instructors. Check new items for ceiling effects before interpreting flat growth: Modeling & Simulation showed d = 0.00 because pre-instruction means were already 1.62 out of 2, so the simulation-design prompt measured baseline knowledge rather than growth.
- Researchers. Keep humans in the loop on complex constructs and retain periodically re-coded human subsamples (calibration reached only Fleiss' κ = 0.53) as a check on automated drift.
Limitations
Both human raters and the LLM showed lower agreement on multi-component constructs, highlighting the difficulty of assessing complex reasoning in brief written responses. The near-ceiling pre-instruction scores on Modeling and Simulation Practices meant the simulation-design prompt served mainly as an indicator of baseline knowledge rather than of new cognitive growth. The authors note that future rubrics should define Systems Thinking more explicitly, distinguishing identifying system components from explaining their interactions and consequences.
Citation
Savage, S., Shanker, A., Michlitsch, G., & Rebello, N. S. (2026). Using LLMs to Detect Growth in Computational Thinking in Introductory Physics.