On this page

Synthesis: LLMs systematically underestimate the difficulty of misconception-driven items ('The Easy Trap'). While Large Language Models (LLMs) ratings show moderate rank correlation with empirical student difficulty (rho=0.52-0.70), they misclassify several fraction items as easy that are among the hardest for students (e.g., 34% correct). LLMs approximate curricular rather than cognitive difficulty.

Relevance to AI in Education: This paper contributes to the understanding of Automated Assessment, Personalized Learning, and Student Experience. The findings have implications for Adaptive Learning systems, Formative Assessment design, and the broader Edtech Platform landscape. Future work should explore how these results generalize across STEM Education and Higher Education contexts.

This research connects to the growing body of work on AI Literacy and Teaching, highlighting both the promise and limitations of AI tools in educational settings.

What this means for practice

  • Instructors. Do not accept LLM-estimated difficulty ratings for misconception-heavy topics: model ratings correlated with empirical student difficulty at Spearman's rho of only 0.52–0.77 and marked items easy on which 34.16% of students were correct.
  • Instructors. Validate item difficulty against your own students' p-values before using model ratings to sequence instruction or assemble practice sets.
  • Designers. Add a misconception-informed signal or a human review step to any adaptive system that selects items or estimates learner ability from model-generated difficulty.
  • Students. Spend deliberate time on fraction operations even when a tool rates them easy; the clearest underestimation appeared in 11 fraction items where conceptual understanding, not procedure count, drives difficulty.
  • Designers. Expect difficulty estimates to track curricular position rather than cognitive demand, and calibrate per topic rather than trusting a model's overall ranking ability.

Limitations

  • Participants were 770 second-year undergraduates at a single Indonesian institution, all with the Indonesia-K13 curriculum, so generalization to other systems and cultures is untested.
  • The analysis rested on only 32 arithmetic items — 5 number-operation, 10 integer, 11 fraction, 4 decimal, and 2 exponent/root items — and the clearest evidence of systematic underestimation came from 11 fraction items.
  • Model difficulty came from 640 ratings (4 models × 5 repetitions × 32 items), a snapshot tied to specific model versions and prompts.
  • Prompting deliberately withheld misconception information to mirror everyday educator use, so the study cannot show whether richer prompts would remove the bias.

Citation

Amanda La Hadi, Muhammad Johan Alibasa, Guanliang Chen, A. Taufiq Asyhari (2026). The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty. EDM 2026 (Educational Data Mining Conference).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.