Hadi et al. (2026) โ EDM 2026 (Educational Data Mining Conference).
๐ Full text (arXiv)
LLMs systematically underestimate the difficulty of misconception-driven items ('The Easy Trap'). While LLM ratings show moderate rank correlation with empirical student difficulty (rho=0.52-0.70), they misclassify several fraction items as easy that are among the hardest for students (e.g., 34% correct). LLMs approximate curricular rather than cognitive difficulty.
Relevance to AI in Education: This paper contributes to the understanding of llm-assessment, personalized-learning, and student-experience. The findings have implications for adaptive-learning systems, formative-assessment design, and the broader edtech-platform landscape. Future work should explore how these results generalize across stem-education and higher-ed contexts.
This research connects to the growing body of work on ai-literacy and teacher-role, highlighting both the promise and limitations of AI tools in educational settings.
Related Pages
- llm-assessment โ related concept
- formative-assessment โ related concept
- adaptive-learning โ related concept
- student-experience โ related concept
- stem-education โ related concept
- feedback-loop โ related concept