On this page

Core Finding

Can large language models (ChatGPT-5 and Microsoft Copilot Pro) conduct meta-assessment — evaluating the quality of assessment reports? Comparing AI ratings to a human expert across three report versions (strong/moderate/weak), an evaluation format (checklist vs. rubric), three assessment elements, and five replications, AI aligned with human ratings at 87% (checklist) and 44–50% (rubric) agreement overall — but the number alone obscures deep limitations: AI struggled most with measurement quality and weak/high-error reports, and even when scores agreed, AI's rationale often conflicted with expert reasoning. AI is a valuable supplemental tool, not a replacement for human expertise.

Key Findings

  • Format matters greatly. The binary checklist produced high overall agreement (ChatGPT-5 = 87%, Copilot Pro = 89%) because its simple decision structure limits disagreement; the more nuanced rubric produced far lower agreement (44% and 50%) and revealed much larger divergences. The only major checklist failure was Use of Results on the weak report, where both models credited intent rather than evidence.
  • AI is better on high-quality reports. Agreement was strongest for the strong report across both formats, and declined as quality fell. On weaker reports, AI frequently inflated scores and failed to penalize missing or misclassified evidence (e.g., Copilot consistently scored the weak report's Improvement element a perfect 1 in the checklist).
  • Measurement is the hardest element. Use of Results had the highest agreement (68%), followed by SLOs (61%) and Measures (54%) — but this understates the problem: AI had consistent difficulty judging outcome–measure alignment and distinguishing direct from indirect evidence, errors that cascaded into other ratings.
  • Four recurring error types: (1) hallucinations (fabricated SLOs — two true instances), (2) misses (failing to detect missing/low-quality evidence on weaker reports), (3) misplaced attention (e.g., Copilot basing an Interpretation of Results rating on the Use of Results section), and (4) misapplication of criteria (faulty judgment applied to correctly identified content). These often intersected.
  • Secure agreement ≠ sound rationale. Even where human and AI scores matched, the underlying reasoning often differed. AI accepted report labels at face value (e.g., taking "exit interview" as a direct measure) and treated any mention of "change" as proof of data-informed improvement — suggesting credible-sounding but misleading labels could fool untrained LLMs. Human expertise adds contextual, nonverbal-perceptual reasoning AI currently lacks.
  • Model differences: ChatGPT-5 was stricter, more conservative, and more variable (occasionally returning fractional scores to signal uncertainty and adopting an unnecessarily strict linguistic standard); Copilot Pro was highly stable but consistently lenient/inflated.

Practical Implications

  • Use AI as a supplement, not a replacement. For strong reports or a simple checklist, AI can streamline initial review and efficiently flag vague language and non-student-centered phrasing — useful especially for faculty/staff new to assessment.
  • Keep human judgment on measurement and weak reports. AI is least reliable at evaluating measurement quality/alignment and low-quality reports — exactly where faculty need developmental support. A hybrid model (AI for initial screens, humans for nuanced elements) can offset staffing limits in assessment offices.
  • Expect and guard against label-based overcredulity. LLMs accepted surface labels at face value; institutions should verify AI verdicts against evidence rather than trust labels, and mirror this training for new human raters too.
  • Establish review and verification guidelines. Institutions adopting AI for meta-assessment should set clear policies for human oversight, disclosure of AI use in feedback, and data-privacy handling (what is stored or used for training) — especially with real institutional data.
  • Use AI meta-assessment as a reflective training byproduct. The process of prompting and asking AI to justify ratings encourages assessment professionals to critically reflect on their own criteria and rationale — itself a valuable professional-development outcome.

Connected Concepts

Connected Articles

Citation

Green, K., Bao, Y., LeRoy, S., & Good, M. (2026). Can AI Evaluate Assessment? A Study of Large Language Model Meta-Assessment Performance. Research & Practice in Assessment, 21(2).