🧠 AI Ed Wiki

Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun — arXiv preprint (2026).

Synthesis

This study challenges the standard practice of evaluating LLM qualitative coding by agreement with human coders, using data from a K-12 AI platform: five LLM systems and three trained human coders applied a 72-item hierarchical codebook to 2,560 educator messages.

An independent domain expert judged 855 pairwise code-set comparisons blind to source, treating human and machine outputs symmetrically. Human-LLM agreement (mean Jaccard 0.30) fell well below human-human agreement (0.52), yet the blind verifier preferred human and LLM coding at indistinguishable rates (51.5% vs 48.5%, p = 0.537).

A Bradley-Terry ranking placed two LLMs above two of three human coders, and for several substantive codes human consensus encoded shared bias that the verifier rejected in favor of the LLM interpretation — evidence that agreement metrics can mislead automation decisions.

The study contributes a transferable blind-verification protocol for evaluating qualitative coding quality in AIED research, with implications for how LLM-assisted analysis of educator and student data should be validated.

Connected Concepts

  • Teacher AI Competency
  • Bias Mitigation
  • K 12 AI Education
  • Student Experience
  • Equity In AI Education
  • Culturally Relevant Pedagogy
  • AI Education
  • Formative Assessment
  • Connected Articles

  • Human LLM Collaborative Coding K12 Educator AI — Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
  • Agent Voice Accents K12 Group Learning — Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning
  • AI Changing Teaching Workflows — How AI Is Changing Teaching Workflows
  • AI Education Global Capacity — What AI in Education Needs Next: Lessons from Youth Leaders Across Five Countries
  • Civic Education AI Lesson Plans — AI-Generated Lesson Plans in Civic Education
  • Lodge Loble Cognitive Offloading 2026 — Artificial intelligence, cognitive offloading and implications for education
  • Citation

    Liu, A., Esbenshade, L., Xiao, M., Tian, V., Zhang, Z., He, K., & Sun, M. (2026). Agreement is not quality: Blind expert verification of human and LLM qualitative coding. arXiv:2607.28890.