🧠 AI Ed Wiki

Synthesis: Savage, Shanker, Michlitsch & Rebello (2026) investigate using LLMs to evaluate students' written explanations of computational physics problems at scale. Establishing a human-coded baseline grounded in CT literature, they found significant growth in Data Practices and Computational Problem-Solving Practices. The LLM successfully mirrored human evaluations for these constructs, but both human raters and the LLM struggled with more complex constructs like Systems Thinking. This work demonstrates that LLMs offer a viable, scalable method for assessing computational thinking in large-enrollment physics courses.

As computation becomes more central to physics education, scalable methods to assess authentic computational thinking (CT) are critically needed. This study establishes a human-coded baseline grounded in CT literature, identifies significant pre/post growth in Data Practices and Computational Problem-Solving Practices, and demonstrates that an LLM can mirror human evaluations — scaling CT assessment across large datasets. Notably, both human raters and the LLM struggled with more complex constructs like Systems Thinking, revealing the limits of current automated approaches.

  • LLMs successfully mirrored human coding for Data Practices and Computational Problem-Solving Practices
  • Both human raters and LLMs struggled with Systems Thinking — revealing construct complexity
  • The approach scales CT assessment to large-enrollment physics courses where manual coding is infeasible
  • Submitted to Physics Education Research Conference (PERC) 2026
  • Connected Concepts

  • Physics Education
  • Computational Thinking
  • STEM Education
  • Automated Grading
  • Educational Measurement
  • Higher Ed
  • Connected Articles

  • Hashmi Socratic Physics Chatbot 2025
  • AI Scoring Language Bias Physics
  • Citation

    Savage, S., Shanker, A., Michlitsch, G., & Rebello, N. S. (2026). Using LLMs to Detect Growth in Computational Thinking in Introductory Physics.