Research Article
Opening the Blackbox of LLM-Based Automated Essay Scoring: Insights into Feature Weighting Patterns and Score Validity
Synthesis: Wang, Chen, Huang, and Lai (2026) systematically compare the scoring behavior of three LLMs (Qwen, GPT, and Gemini) with human raters on English essays written by non-native learners. Analyzing sixteen textual features, they find strong overall alignment but distinct feature weighting patterns: the LLMs placed greater emphasis on grammatical accuracy, lexical sophistication, and syntactic complexity, while human raters prioritized content completeness and visual presentation with greater tolerance for minor linguistic errors. Across proficiency levels, human raters exhibited a more stable scoring framework, while LLMs showed larger cross-group shifts — placing more weight on language errors for low-proficiency students and increasingly rewarding linguistic sophistication for high-proficiency students.
Key Findings
- Three LLMs (Qwen, GPT, Gemini) showed strong overall score alignment with human raters but distinct feature weighting patterns.
- LLMs emphasized grammatical accuracy, lexical sophistication, and syntactic complexity; human raters prioritized content completeness and visual presentation.
- Human raters exhibited a more stable scoring framework across proficiency levels; LLMs showed larger cross-group shifts.
- LLMs placed more weight on language errors for low-proficiency students and increasingly rewarded linguistic sophistication for high-proficiency students.
- LLMs integrated multiple features when scoring, with integration patterns varying by proficiency level.
Implications for AI in Education
The study highlights both the potential and limitations of LLM-based scoring and underscores the importance of interpretability and transparency to enhance scoring validity. The finding that LLMs weight formal linguistic features more heavily than human raters — and shift their weighting by proficiency level — raises validity and fairness concerns, particularly for non-native writers. For assessment designers, the results suggest that LLM scoring should be calibrated against human judgment and audited for feature-weighting patterns. The study connects to Educational NLP, Writing Education, and Equity In AI Education research on automated assessment.
Connected Concepts
- Automated Essay Scoring
- LLM
- Educational NLP
- Assessment Validity
- Writing Education
- Bias Mitigation
- Equity In AI Education
- Automated Assessment
Connected Articles
- [llm-essay-assessment-framework-reliability-2026] — framework for evaluating LLMs in essay assessment
- [llms-do-not-grade-essays-like-humans-2026] — LLMs do not grade essays like humans
- [ai-scoring-language-bias-physics] — language bias in AI-based scoring
- [choi-anchor-aes-prompting-2025] — anchor-paper prompting for AES
Citation
Wang, M., Chen, Y., Huang, X., & Lai, Y. (2026). Opening the blackbox of LLM-based automated essay scoring: Insights into feature weighting patterns and score validity. Computers and Education: Artificial Intelligence, 10, 100568.