🧠 AI Ed Wiki

Tseng et al. (2026) investigate human-machine alignment in LLM-based pretest question evaluation — a critical bottleneck for scalable AI-assisted assessment. Their AI-assisted workflow combines automated generation, rubric-based evaluation, and iterative selection. Through a 2×2 experimental design varying rubric operationalization and evaluation mode, they find that human-machine disagreements are systematic rather than random, rubric revision has a larger effect on alignment than rationale-first evaluation, and the two interventions are complementary. The core insight is that scalable AI-assisted pretesting depends not only on generation capability but crucially on how pedagogical quality is operationalized for machine interpretation. This work contributes to AI Ed Evaluation by providing empirical evidence for aligning LLM judgment with human pedagogical standards in Formative Assessment contexts, and has direct implications for Automated Grading and Assessment system design.

Connected Concepts

  • AI Ed Evaluation
  • LLM
  • Formative Assessment
  • Automated Grading
  • Assessment
  • Connected Articles

  • Responsible Assessment AI Era Stanford 2026 — Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
  • Cotal Formative Assessment Scoring 2026 — CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
  • Authentic Products Authenticated Processes 2026 — From authentic products to authenticated processes: authentic assessment in AI-rich higher education
  • Automated Formative Assessments A Level Sciences — The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences
  • Hybrid E Assessment Semi Automated Grading — Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations
  • Cross Dataset Bloom Question Classification — Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
  • Citation

    Pei-Yu Tseng, Mahir Akgun, Peng Liu (2026). Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation. arXiv:2606.23629. arXiv:2606.23629 (cs.HC)