🧠 AI Ed Wiki

🏷️ ai-ed-evaluation

22 pages tagged with ai-ed-evaluation(17 articles, 5 concepts)

📄 ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
> **Synthesis:** Jiang et al. (2026) introduce **ELBench**, the first benchmark to evaluate education-facing LLMs on all four required dimensions — General Capability, Safety and Trustworthiness, Basi…
2026-08-13 · benchmark, llm-evaluation, pedagogical-safety, llm, safety
📄 Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
> **Synthesis:** Lin et al. (2026) present the **Teaching Monster Challenge**, the first instructional-video generation benchmark that treats the learner persona as an explicit evaluation criterion, m…
2026-08-13 · benchmark, llm-evaluation, agentic-ai, pedagogical-agent, content-generation
🏷️ Benchmark
> **Benchmark** — standardized test suites and evaluation frameworks used to measure AI model performance on educational tasks. Benchmarks enable reproducible comparison across models and approaches, …
2026-08-09 · assessment, llm, generative-ai, benchmark
🏷️ Hallucination Risk
> **Hallucination Risk** — the danger that AI systems generate plausible but factually incorrect or fabricated content in educational contexts, where such errors can mislead learners, undermine trust,…
2026-08-09 · hallucination-risk, generative-ai, llm, pedagogical-safety, human-in-the-loop-ai
🏷️ Learning Analytics
> **Learning analytics** — the measurement, collection, analysis, and reporting of data about learners and their contexts for the purpose of understanding and optimizing learning. AI has transformed l…
2026-08-09 · knowledge-tracing, student-modeling, formative-assessment, privacy, edtech-platform
🏷️ Learning Gains
> **Learning gains** — measurable improvements in student knowledge, skills, or competencies resulting from educational interventions, including AI-assisted instruction. In AI in education research, l…
2026-08-09 · assessment, student-experience, higher-ed, k-12
📄 What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries
> **Synthesis:** This paper reports a blind Turing Test evaluating leading LLMs on three Italian professional legal examinations: the Bar exam, Judges exam, and Notary exam. LLMs generated full writte…
2026-08-07 · llm, assessment, professional-training, benchmark, automated-grading
📄 Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
> **Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education** — Longitudinal co-design with learning engineers building an LLM-powered digital textbook. C…
2026-08-05 · llm, trust-calibration, human-in-the-loop, instructional-design, edtech-platform
📄 EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
> **EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners** — Introduces a 30-day long-horizon benchmark for pedagogical LLM agents using simulated learners ground…
2026-08-05 · intelligent-tutoring, llm, agentic-ai, benchmark, knowledge-tracing
📄 Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
> **GPT-4o-mini can produce stable rubric-based scores for open-ended music analysis responses, with few-shot chain-of-thought prompting agreeing most strongly with teacher means while RAG systematica…
2026-08-04 · automated-grading, llm, assessment-validity, higher-ed, rag
📄 Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
> **Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun** — arXiv preprint (2026).…
2026-08-03 · llm, qualitative-research, k-12, teacher-role, equity
📄 From authentic products to authenticated processes: authentic assessment in AI-rich higher education
> Generative AI has not created the need for authentic assessment — it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by to…
2026-08-03 · authentic-assessment, assessment, assessment-validity, generative-ai, academic-integrity
📄 CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
> 1. **Evidence-Centered Design (ECD)** — assessments and rubrics aligned to curriculum goals from the start 2. **Human-in-the-loop prompt engineering** — labelled examples and prompts refined iterati…
2026-08-03 · formative-assessment, automated-grading, human-in-the-loop, prompt-engineering, benchmark
📄 Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
> **Alex Liu, Min Sun, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He** — arXiv preprint (2026).…
2026-08-03 · llm, qualitative-research, k-12, teacher-role, generative-ai
📄 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
> **Responsible assessment in the AI era** — assessment grounded in learners' sociocultural contexts and designed to generate valid, trustworthy, context-specific inferences from accumulated evidence,…
2026-08-03 · assessment, assessment-validity, formative-assessment, generative-ai, educational-theory
📄 AICoFE: AI-Powered Feedback System
> **AICoFE** (AI-based Collaborative Feedback) is a multi-LLM feedback generation system for higher education that combines independently fine-tuned language models with **teacher-in-the-loop mediatio…
2026-07-29 · feedback-loop, student-experience, human-in-the-loop-ai, higher-ed, learning-analytics
📄 Confidence-Aware Automatic Short Answer Grading
> **Confidence-Aware ASAG** — A hybrid confidence estimation framework for Automatic Short Answer Grading with LLMs that fuses model-based confidence signals (verbalized, latent, consistency-based) wi…
2026-07-29 · assessment, automated-grading, confidence, psychometrically-aware-ai, hybrid-e-assessment-semi-automated-grading
📄 ISD Agent Benchmark
> **ISD-Agent-Bench** is a comprehensive benchmark for evaluating LLM-based instructional design agents, comprising **25,795 scenarios** generated via a Context Matrix framework that combines 51 conte…
2026-07-29 · agentic-ai, benchmark, rag, llm, agentic-workflows
📄 Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation
Tseng et al. (2026) investigate human-machine alignment in LLM-based pretest question evaluation — a critical bottleneck for scalable AI-assisted assessment. Their AI-assisted workflow combines automa…
2026-06-23 · llm, formative-assessment, teacher-role, assessment, automated-grading
📄 Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
> **MathCog** benchmark (3,036 teacher-annotated diagnostic verdicts, 639 handwritten responses, 18 LLMs): all models severely underperform (macro F1 < 0.5) — over-attributing evidence, overthinking m…
2026-05-31 · knowledge-tracing, multimodal, benchmark, human-in-the-loop, critical-thinking
📄 Authentic Assessment
> Wiggins (1990) proposed AA as a counterbalance to standardised tests: direct examination of "student performance on worthy intellectual tasks." > Authentic assessment (AA) has evolved from workplace…
2026-05-07 · ai-education, assessment, formative-assessment, higher-ed, metacognition
🏷️ Formative Assessment in AI Education
Assessment designed to inform ongoing instruction and learning, as opposed to summative evaluation. AI systems can generate, validate, and adapt formative assessment items at scale, though quality var…
2026-05-07 · agentic-ai, ai-education, assessment, pedagogical-safety, llm

Related Tags

llm (15)generative-ai (9)assessment (9)benchmark (8)formative-assessment (7)higher-ed (6)automated-grading (6)k-12 (5)agentic-ai (4)human-in-the-loop (4)assessment-validity (4)pedagogical-safety (3)human-in-the-loop-ai (3)knowledge-tracing (3)teacher-role (3)