Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

Created: 2026-08-03 | Tags: llmqualitative-researchk-12teacher-roleai-ed-evaluationequityresearch-methods

Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun โ€” arXiv preprint (2026). ๐Ÿ“„ Full text (arXiv)

Synthesis

This study challenges the standard practice of evaluating LLM qualitative coding by agreement with human coders, using data from a K-12 AI platform: five LLM systems and three trained human coders applied a 72-item hierarchical codebook to 2,560 educator messages.

An independent domain expert judged 855 pairwise code-set comparisons blind to source, treating human and machine outputs symmetrically. Human-LLM agreement (mean Jaccard 0.30) fell well below human-human agreement (0.52), yet the blind verifier preferred human and LLM coding at indistinguishable rates (51.5% vs 48.5%, p = 0.537).

A Bradley-Terry ranking placed two LLMs above two of three human coders, and for several substantive codes human consensus encoded shared bias that the verifier rejected in favor of the LLM interpretation โ€” evidence that agreement metrics can mislead automation decisions.

The study contributes a transferable blind-verification protocol for evaluating qualitative coding quality in AIED research, with implications for how LLM-assisted analysis of educator and student data should be validated.

Related Pages

Citation

APA: Liu, A., Esbenshade, L., Xiao, M., Tian, V., Zhang, Z., He, K., & Sun, M. (2026). Agreement is not quality: Blind expert verification of human and LLM qualitative coding. arXiv:2607.28890.