AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics

Created: 2026-07-31 | Tags: assessment-validityautomated-gradingbias-mitigationequitymultilingual-learningphysics-education

Authors: Markus S. Feser, Paul L. Tschisgale (Leibniz Institute for Science and Mathematics Education, Kiel, Germany)

Source: arXiv:2607.28210 (physics.ed-ph, July 2026)

Key Findings

This study examined whether AI-based scoring can assess students' conceptual understanding independently of the linguistic quality of their text-based explanations in physics. The researchers compared scores from 9 machine learning (ML) approaches and 2 large language model (LLM) approaches against human expert scores for 116 secondary-school students' physics explanations.

The Language Bias Problem

Disproportionate Impact

The stakes fall hardest on multilingual learners, whose language proficiency may be misread as weaker conceptual understanding. This is especially concerning as AI-based scoring takes on higher-stakes assessment decisions.

Relevance to AI in Education

This paper makes a critical contribution to the automated-assessment and automated-essay-scoring literature by demonstrating that the bias-mitigation problem in AI scoring is not merely a technical artifact of specific models but appears to be fundamental to the task itself. Key connections:

Implications

1. Benchmarking AI scoring: AI-based scoring systems should be explicitly evaluated for language bias, not just overall agreement with human scores. 2. High-stakes caution: As AI scoring moves toward higher-stakes decisions, the asymmetric language bias becomes increasingly consequential. 3. Multimodal assessment: The findings support calls for assessment approaches that reduce dependence on linguistic production, particularly for multilingual-learning populations. 4. Teacher-AI collaboration: Rather than replacing teacher assessment, AI scoring may be most useful when teachers remain in the loop to calibrate for language effects.

Related Pages