🧠 AI Ed Wiki

Bannò, Knill & Gales (2026) propose a paradigm shift in automated essay scoring: from inter-learner ranking to intra-learner profiling. Instead of asking "how does this essay rank against others?", their self-referential framework asks "what are this specific learner's strengths and weaknesses?"

Self-Referential L2 Writing Assessment with LLMs

Core Contribution

Bannò, Knill & Gales (2026) propose a paradigm shift in automated essay scoring: from inter-learner ranking to intra-learner profiling. Instead of asking "how does this essay rank against others?", their self-referential framework asks "what are this specific learner's strengths and weaknesses?"

Key Findings

Using the ICNALE GRA dataset annotated by up to 80 trained raters and calibrated with two-facet Rasch modeling:

  • LLMs outperform single human raters at identifying relative weaknesses (negative feedback) across proficiency aspects
  • Human raters remain stronger at identifying relative strengths (positive feedback)
  • Traditional rank-based correlation metrics mask diagnostic behavior — high correlations can hide poor intra-learner discrimination
  • Implications for AIED

    This connects to Automated Grading but challenges its dominant evaluation paradigm. The finding that LLMs are strong at weakness detection but weaker at strength identification has practical implications for Formative Assessment design — AI might best serve as a complementary weakness detector while teachers focus on strengths.

    The self-referential approach aligns with Personalized Learning goals and the AI Learning Companions Framework emphasis on prioritizing learning over performance. It extends Writing Education research on AI in composition and connects to Automated Question Generation work on AI-generated assessment. The use of Rasch modeling for calibration connects to Ground Truth Reliability AIED calls for more rigorous measurement in AIED.

    Connections to Wiki

  • Paradigm shift from Automated Grading ranking to profiling
  • Aligns with Sequenced AI Feedback Learning emphasis on feedback quality over quantity
  • Extends LLM Student Modeling Memory to assessment contexts — profiling over time
  • Complements Human In The Loop AI by identifying where humans vs. AI add value
  • Connected Concepts

  • Automated Grading
  • Formative Assessment
  • Personalized Learning
  • Writing Education
  • Automated Question Generation
  • Human In The Loop AI
  • Connected Articles

  • AI Learning Companions Framework
  • Ground Truth Reliability AIED
  • Sequenced AI Feedback Learning
  • LLM Student Modeling Memory
  • Icle Plus Plus Essay Scoring
  • Citation

    Gales, A.S.B.K.K.M., Approach, T.S.A.A.A.P., LLMs, T.L.W.E.W., Gales, S.B.K.K.M., prac-, A.I.R.W.C.D.T.A.C.E., & (PCC), G.A.U.D.S.W.S.A.P.C.C. (2026). Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs