AI-ed evaluation — the body of methods, benchmarks, and criteria used to assess whether AI education tools (LLM-based tutors, automated graders, feedback systems, agents) actually work — not just on headline accuracy, but on reliability, pedagogical quality, validity, and real learning impact. A recurring theme across the wiki's research is that evaluation must be domain-specific, reliability-aware, and anchored in human judgment and educational outcomes rather than single aggregate accuracy numbers.
AI-ed evaluation spans several distinct objects of assessment. It can evaluate the output (is the AI's answer, grade, or feedback correct and reliable?), the process (does the tool support valid, defensible assessment and learning?), and the agent (does an AI tutor or agent teach effectively and behave appropriately?). Each requires different methods and raises different validity questions.
How AI-ed evaluation appears in the research
Output reliability and ground truth: Modernizing ground truth argues that reliability problems in AI-ed evaluation often trace back to the reference data itself — the "ground truth" labels systems are judged against — and proposes four shifts toward improving reliability and validity. Calibrating trustworthiness co-designs evaluation metrics and visualizations with stakeholders so that trust in an AI tool rests on demonstrated, interpretable evidence.Automated grading and scoring: LLM short-answer grading, confidence-aware ASAG, CoTAL human-in-the-loop prompt engineering, and cognitive-diagnosis of handwritten math show that LLMs can grade and diagnose, but that reliability depends on human oversight, domain-specific grounding, and confidence calibration rather than raw model size.Benchmarking and domain specificity: TeachBench evaluates LLM teaching ability, ISD Agent Benchmark evaluates agentic instructional-design agents, educational VLM evaluation assesses multimodal models, and a tool-invariant framework assesses computational-method competency. These share a warning: generic benchmarks mislead, and evaluation must be tailored to the specific educational task and context.Pedagogical quality and alignment: Why machines misread pedagogical quality documents human–machine misalignment in judging what makes instruction good, and the Tutoring Effectiveness Index predicts tutor quality from teaching behavior. Responsible assessment in the AI era and authenticated processes argue that evaluation must reach beyond correct answers to whether assessment remains authentic, valid, and defensible when AI can produce the "products" of learning.Evaluating learning outcomes and agents: the ITS systematic review, LLM-difficulty calibration, Socratic conversational tests, and valid student simulation broaden evaluation to learning gains, test validity, and whether simulated students are a valid proxy for real learners.Why evaluation is hard in AI-ed
AI-ed evaluation is difficult for several reasons. First, reliability is not enough — a system can agree with a rubric yet misjudge pedagogy, as human–machine alignment research shows. Second, ground truth is contested — what counts as a "correct" answer, grade, or teaching move is itself a judgment that varies across disciplines and experts, per ground-truth modernization. Third, educational validity is multidimensional — Assessment Validity, Formative Assessment, and Authentic Assessment each impose different criteria that a single accuracy metric cannot capture. Finally, the target keeps moving — agentic AI and multimodal models demand evaluation frameworks (Agentic AI, tool-invariant assessment) rather than reuse of text-model benchmarks.
Connections to related concepts
AI-ed evaluation sits at the center of the wiki's methods and risks. It operationalizes Assessment Validity, Educational Measurement, and Benchmark within Assessment and Automated Assessment. Its call for human oversight connects to Human In The Loop AI and Teacher Role, while its focus on reliability connects to Hallucination Risk, Confidence Aware AI Assessment, and Trust Calibration. The distinction between evaluating performance and evaluating learning links to performance vs. learning and to Student Modeling; and evaluation of pedagogical agents connects to Intelligent Tutoring, Pedagogical LLM Training, and Pedagogical Safety.
Connected Concepts
Assessment ValidityEducational MeasurementBenchmarkAutomated AssessmentFormative AssessmentAuthentic AssessmentHuman In The Loop AIConfidence Aware AI AssessmentHallucination RiskTrust CalibrationIntelligent TutoringAgentic AITeacher RoleAssessmentLLMConnected Articles
Ground Truth Reliability AIED — Modernizing Ground Truth: Four Shifts Toward Improving Reliability and ValidityCalibrating Trustworthiness LLM Education 2026 — Calibrating Trustworthiness: Co-Designing Metrics and VisualizationsTeachbench LLM Teaching Evaluation — TeachBench: Evaluating LLM Teaching AbilityMachines Misread Pedagogical Quality — Why Machines Misread Pedagogical Quality: Human-Machine AlignmentAutomatic Short Answer Grading — Automatic Short Answer Grading With LLMsCotal Formative Assessment Scoring 2026 — CoTAL: Human-in-the-Loop Prompt Engineering for Formative AssessmentCong Confidence ASAG 2026 — Confidence-Aware Automatic Short Answer GradingLLM Cognitive Diagnosis Handwritten Math — Benchmarking LLMs for Diagnosing Students' Cognitive SkillsTutoring Effectiveness Index — The Tutoring Effectiveness Index: Predicting LLM Math Tutor QualityJeon Isd Agent Bench 2026 — ISD Agent BenchmarkEducational Vlm Evaluation — Educational VLM EvaluationTool Invariant Framework Agentic AI — A Tool-Invariant Framework for Teaching and Assessing Computational MethodsValid Student Simulation LLM 2026 — Towards Valid Student Simulation With Large Language ModelsLLM Difficulty Calibration Programming Exams 2026 — From Evaluated Models to Evaluation AidsSocratic Tests Conversational Assessment — The Theoretical Foundation of Socratic TestsResponsible Assessment AI Era Stanford 2026 — Responsible Assessment in the AI EraAuthentic Products Authenticated Processes 2026 — From Authentic Products to Authenticated ProcessesZerkouk Comprehensive Review ITS 2025 — Comprehensive Review of Intelligent Tutoring SystemsAI Peer Feedback Systems — AI Peer Feedback SystemsBecerra Aicofe Feedback 2026 — AICoFE: AI-Powered Feedback SystemElbench Education LLM Benchmark 2026Teaching Monster Pck Benchmark 2026