🧠 AI Ed Wiki

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education — Longitudinal co-design with learning engineers building an LLM-powered digital textbook. Co-constructed five trustworthiness metrics with 20 measures tailored to pedagogical use. Designed visualizations mapping trustworthiness violations onto LLM res... LLM AI Ed Evaluation Over Reliance Human In The Loop AI Instructional Design Edtech Platform

Longitudinal co-design with learning engineers building an LLM-powered digital textbook. Co-constructed five trustworthiness metrics with 20 measures tailored to pedagogical use. Designed visualizations mapping trustworthiness violations onto LLM responses. Making trustworthiness explicit increased inter-rater reliability and helped learning engineers resolve conflicting objectives and produce more consistent judgments. Proposes design guidelines for future LLM evaluation tools that enable pedagogically-aligned learning tools.

Abstract

LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; designed visualizations that map trustworthiness violations onto LLM responses; and evaluated how these tools help learning engineers make A/B comparisons of LLM responses.

Connected Concepts

  • LLM
  • AI Ed Evaluation
  • Over Reliance
  • Human In The Loop AI
  • Instructional Design
  • Edtech Platform
  • Connected Articles

  • LLM Cognitive Diagnosis Handwritten Math — Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
  • Cotal Formative Assessment Scoring 2026 — CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
  • Veriforge Narrative Drafting Scaffolding 2026 — VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding
  • LLM Intervention Design CS Review — A review of intervention designs of LLM Integration in Undergraduate Computer Science Education
  • Cong Confidence ASAG 2026 — Confidence-Aware Automatic Short Answer Grading
  • Jeon Isd Agent Bench 2026 — ISD Agent Benchmark
  • Citation

    Adam Coscia, Sujata Duwal, Langdon Holmes, Scott Crossley, & Alex Endert (2026). Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education. arXiv:2608.04006. arXiv:2608.04006 [cs.HC] (under review).