AI-ed evaluation covers how AI education tools are benchmarked and assessed โ see cotal-formative-assessment-scoring-2026, benchmark, ground-truth-reliability-aied, and teachbench-llm-teaching-evaluation โ with recurring findings that evaluation must be domain-specific and reliability-aware, not headline-accuracy-driven.
Related Pages
- cotal-formative-assessment-scoring-2026 โ Cross-domain grading evaluation
- authentic-products-authenticated-processes-2026 โ Six-dimension framework for assessment design
- llm-cognitive-diagnosis-handwritten-math โ MathCog benchmark: 18 LLMs evaluated on cognitive skill diagnosis from handwritten math; all F1 < 0.5; systematic over-attribution and hallucination of evidence (2025)
- machines-misread-pedagogical-quality -- Human-machine disagreements in AI pretest evaluation are systematic; rubric revision has a larger alignment effect than rationale-first evaluation, and the two are complementary.
๐ 13 other pages tagged ai-ed-evaluation
- Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
- AICoFE: AI-Powered Feedback System
- Authentic Assessment
- Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
- Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
- CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
- Formative Assessment in AI Education
- From authentic products to authenticated processes: authentic assessment in AI-rich higher education
- Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
- ISD Agent Benchmark
- Randomized Controlled Trials in AI Education Research
- Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
- Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation