🏷️ hallucination-risk
18 pages tagged with hallucination-risk(12 articles, 6 concepts)
📄 ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills
> Pardos & Bhandari (2024) report a randomized efficacy study (N=274) comparing ChatGPT-generated hints to human tutor-authored hints and a no-help control across four mathematics subject areas. Only …
🏷️ Trust Calibration
> **Trust calibration** — the metacognitive capacity to align one's confidence in an AI system with its actual reliability in a given context, knowing when to trust and when to question its output. Tr…
🏷️ Generative AI
> **Generative AI** — AI systems capable of producing text, code, images, and other content, most prominently large language models like GPT-4 and Claude. Generative AI is the technology driving the c…
🏷️ Hallucination Risk
> **Hallucination Risk** — the danger that AI systems generate plausible but factually incorrect or fabricated content in educational contexts, where such errors can mislead learners, undermine trust,…
🏷️ Large Language Models (LLMs)
> **Large Language Models (LLMs)** — neural network models trained on vast text corpora that generate human-like text, powering most modern AI in education applications. LLMs are the computational bac…
🏷️ Pedagogical Safety
> **Pedagogical safety** — the design principle that AI education systems must protect learners from harm, including inappropriate content, unsafe advice, biased treatment, and manipulative interactio…
🏷️ RAG (Retrieval-Augmented Generation)
> **RAG (Retrieval-Augmented Generation)** — an AI architecture that combines information retrieval with text generation, allowing LLMs to ground responses in external knowledge sources rather than re…
📄 EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
EduGuard is a retrieval-augmented generation (RAG) tutoring framework that directly confronts the safety and pedagogical failures of unrestricted LLM tutors in introductory programming. Unrestricted t…
📄 AI as a Partner in Learning about, Doing, and Engaging with Science: Vigilance as the Key to Productive Augmentation
Argues that epistemic vigilance — the human evaluation of AI output calibrated to how far a fallible source can be trusted — is the binding constraint on productive augmentation. AI's fluent, confiden…
📄 Stuck in a Spiral": Shame and Guilt as Social Regulators of AI Use in Computing Education
> An interview study with 19 computing students through a functionalist perspective of shame and guilt. Findings show these emotions regulate when and how students make their AI use visible, engaging …
📄 Warning About AI Fallibility Increases Help-Seeking in an Intelligent Tutoring System
> **Synthesis:** Recent work in Technology-Enhanced Learning and HumanComputer Interaction highlights the importance of transparency and trust calibration in AI-supported learning environments as they…
📄 Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
> **MathCog** benchmark (3,036 teacher-annotated diagnostic verdicts, 639 handwritten responses, 18 LLMs): all models severely underperform (macro F1 < 0.5) — over-attributing evidence, overthinking m…
📄 The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration
LLM sycophancy creates a feedback loop where user errors propagate into AI advice, degrading outcomes; AI literacy training reduces but doesn't eliminate this contextual sycophantic dependence. This A…
📄 Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
LLM tutors achieve near-ceiling on correct steps but systematically over-reject valid-suboptimal reasoning and over-validate incorrect solutions — precisely where adaptive tutoring matters most. This …
📄 Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
This paper exposes a critical failure mode in using LLMs as simulated students for [[intelligent-tutoring]] development and evaluation. The authors introduce **misconception faithfulness** — the prope…
📄 Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
> Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks **Kasneci & Kasneci (2026)** — Position paper. arXiv cs.AI/cs.HC.…
📄 Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
> Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs **Maiorano (2026)** — arXiv cs.CR/cs.AI.…
📄 Pedagogical Promise and Peril of AI: A Text Mining Analysis of ChatGPT Research Discussions in Programming Education
This book chapter presents a **text mining analysis** of how scholarly literature frames ChatGPT's role in programming education. Using term frequency analysis, phrase pattern extraction, and topic mo…