Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun โ arXiv preprint (2026). ๐ Full text (arXiv) Synthesis This study challenges the standard practice of evaluating LLM qualitative coding by agreement with human coders, using data from a K-12 AI...
Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
Alex Liu, Min Sun, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He โ arXiv preprint (2026). ๐ Full text (arXiv) Synthesis A multi-phase human-LLM collaborative pipeline adapted open, axial, and selective coding to build a hierarchical codebook from 45,000 mess...
Randomized Controlled Trials in AI Education Research
Randomized controlled trials are the wiki's gold-standard causal evidence: genai-can-harm-teaching-rct-2026 , access-not-enough-ai-tutoring-2026 , lets-chat-chatbot-outreach-2026 , and genai-policy-prompting-rct are pre-registered field experiments showing AI's educational effect...