📄 Research Article
Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation
Tseng et al. (2026) investigate human-machine alignment in LLM-based pretest question evaluation — a critical bottleneck for scalable AI-assisted assessment. Their AI-assisted workflow combines automated generation, rubric-based evaluation, and iterative selection. Through a 2×2 experimental design varying rubric operationalization and evaluation mode, they find that human-machine disagreements are systematic rather than random, rubric revision has a larger effect on alignment than rationale-first evaluation, and the two interventions are complementary. The core insight is that scalable AI-assisted pretesting depends not only on generation capability but crucially on how pedagogical quality is operationalized for machine interpretation. This work contributes to AI Ed Evaluation by providing empirical evidence for aligning LLM judgment with human pedagogical standards in Formative Assessment contexts, and has direct implications for Automated Grading and Assessment system design.
Connected Concepts
Connected Articles
Citation
Pei-Yu Tseng, Mahir Akgun, Peng Liu (2026). Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation. arXiv:2606.23629. arXiv:2606.23629 (cs.HC)