📄 Research Article
Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education — Longitudinal co-design with learning engineers building an LLM-powered digital textbook. Co-constructed five trustworthiness metrics with 20 measures tailored to pedagogical use. Designed visualizations mapping trustworthiness violations onto LLM res... LLM AI Ed Evaluation Over Reliance Human In The Loop AI Instructional Design Edtech Platform
Longitudinal co-design with learning engineers building an LLM-powered digital textbook. Co-constructed five trustworthiness metrics with 20 measures tailored to pedagogical use. Designed visualizations mapping trustworthiness violations onto LLM responses. Making trustworthiness explicit increased inter-rater reliability and helped learning engineers resolve conflicting objectives and produce more consistent judgments. Proposes design guidelines for future LLM evaluation tools that enable pedagogically-aligned learning tools.
Abstract
LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; designed visualizations that map trustworthiness violations onto LLM responses; and evaluated how these tools help learning engineers make A/B comparisons of LLM responses.
Connected Concepts
Connected Articles
Citation
Adam Coscia, Sujata Duwal, Langdon Holmes, Scott Crossley, & Alex Endert (2026). Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education. arXiv:2608.04006. arXiv:2608.04006 [cs.HC] (under review).