On this page

Synthesis: Wang, Chen, Huang, and Lai (2026) systematically compare the scoring behavior of three LLMs (Qwen, GPT, and Gemini) with human raters on English essays written by non-native learners. Analyzing sixteen textual features, they find strong overall alignment but distinct feature weighting patterns: the LLMs placed greater emphasis on grammatical accuracy, lexical sophistication, and syntactic complexity, while human raters prioritized content completeness and visual presentation with greater tolerance for minor linguistic errors. Across proficiency levels, human raters exhibited a more stable scoring framework, while LLMs showed larger cross-group shifts — placing more weight on language errors for low-proficiency students and increasingly rewarding linguistic sophistication for high-proficiency students.

Key Findings

  • Three LLMs (Qwen, GPT, Gemini) showed strong overall score alignment with human raters but distinct feature weighting patterns.
  • LLMs emphasized grammatical accuracy, lexical sophistication, and syntactic complexity; human raters prioritized content completeness and visual presentation.
  • Human raters exhibited a more stable scoring framework across proficiency levels; LLMs showed larger cross-group shifts.
  • LLMs placed more weight on language errors for low-proficiency students and increasingly rewarded linguistic sophistication for high-proficiency students.
  • LLMs integrated multiple features when scoring, with integration patterns varying by proficiency level.

Implications for AI in Education

The study highlights both the potential and limitations of LLM-based scoring and underscores the importance of interpretability and transparency to enhance scoring validity. The finding that LLMs weight formal linguistic features more heavily than human raters — and shift their weighting by proficiency level — raises validity and fairness concerns, particularly for non-native writers. For assessment designers, the results suggest that LLM scoring should be calibrated against human judgment and audited for feature-weighting patterns. The study connects to Educational NLP, Writing Education, and Equity In AI Education research on automated assessment.

Connected Concepts

Connected Articles

  • [llm-essay-assessment-framework-reliability-2026] — framework for evaluating LLMs in essay assessment
  • [llms-do-not-grade-essays-like-humans-2026] — LLMs do not grade essays like humans
  • [ai-scoring-language-bias-physics] — language bias in AI-based scoring
  • [choi-anchor-aes-prompting-2025] — anchor-paper prompting for AES

Citation

Wang, M., Chen, Y., Huang, X., & Lai, Y. (2026). Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity. Computers and Education: Artificial Intelligence, 10, 100568.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.