🧠 AI Ed Wiki

🏷️ efficacy-study

47 pages tagged with efficacy-study(46 articles, 1 concepts)

📄 Methodologies for Improving the Quality of AI Tutoring in K-12 Education
> **Synthesis:** Udeshi et al. (2026), the team behind **Khanmigo** (Khan Academy's K-12 AI tutor, launched 2023), describe the metrics they use to measure AI tutoring quality and student engagement, …
2026-08-13 · ai-tutoring, intelligent-tutoring, k-12, llm, personalized-learning
🏷️ Research Methods in AIED
> **Research methods in AIED** — the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study AI in education: whether and how AI tools support (or harm) …
2026-08-13 · ai-education, educational-measurement, rct, benchmark, methodology
📄 Efficacy of an Intensive Generative AI Professional Development Program on Pedagogical Content Knowledge (AI-PCK) and the Comparative Analysis of Learning Gain between Experienced and Pre-service Teachers
> **Synthesis:** This quasi-experimental study of an intensive 8-hour generative-AI professional development program with 163 teachers and pre-service teachers found significant gains across all five …
2026-08-11 · teacher-professional-development, teacher-ai-competency, generative-ai, professional-training, teacher-training
📄 A framework for characterising and capturing the quality of digital interactions and experiences in early childhood education
> **Synthesis:** This study introduces a Digital Interactions Quality (DigIQ) framework and scale as a protocol to observe and index the quality of interactions and experiences involving digital techn…
2026-08-10 · ai-education, ai-tutoring, educational-technology, edtech-platform, evaluation
📄 Generative AI and the Productivity Divide: Human-AI Complementarities in Education
> **Synthesis:** Idan & Anand (2026) conduct an RCT showing that GenAI access significantly increases task performance on average — but the gains are highly uneven, NOT predicted by GPA or prior knowl…
2026-08-08 · generative-ai, ai-literacy, equity, productivity, higher-ed
📄 A systematic review of generative AI in education: Empirical insights from a human–AI interaction perspective
> **Synthesis:** Liang, Yang, Sha, Gašević, Yan & Chen (2026) systematically review 56 empirical studies on GenAI in education through the AIED-HCD framework, analyzing three human–AI interaction mode…
2026-08-08 · systematic-review, generative-ai, human-ai-interaction, ai-literacy, higher-ed
📄 Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results
> **Shuang Geng, Helen Lallos-Harrell, Jiya Ashar, Thomas J. McKenna, Annwesa Dasgupta, Caleb Farny, Emma Lejeune** — arXiv preprint (2026).…
2026-08-03 · llm, stem-education, higher-ed, student-experience, teacher-role
📄 Artificial intelligence in vocational education and training: A systematic review of educational purposes, theoretical conceptualizations, and empirical effectiveness
> **Authors:** Viola Deutscher, Herbert Thomann, Olga Zlatkin-Troitschanskaia, Ulrike Weyland, Stephan Abele, Amory H. Danek, Samuel Greiff, Andreas Rausch, Susan Seeber, Jürgen Seifried, Esther Winth…
2026-08-01 · professional-training, intelligent-tutoring, generative-ai, adaptive-learning, simulation-based-learning
📄 A review of intervention designs of LLM Integration in Undergraduate Computer Science Education
This scoping review analyzed **13 experimental studies** on LLM integration in undergraduate [[cs-education]], examining how intervention design choices shape learning outcomes. The central finding: *…
2026-07-31 · cs-education, generative-ai, llm, scaffolding, instructional-design
📄 Is Solving Better Than Evaluating GenAI Solutions?
Randomized A/B crossover study (N=220) in a junior-level algorithms course comparing solution evaluation/critique tasks against traditional solution generation. Finds that evaluation-centered tasks pr…
2026-07-31 · generative-ai, stem-education, higher-ed, student-experience
📄 Designing a mobile chatbot-based learning journaling system for intrinsic motivation and engagement
A **randomized 2×2 full-factorial field experiment** (N = 179 German university students, 22 days of app use, 12-week follow-up) testing two design principles for a **mobile chatbot-based learning jou…
2026-07-29 · self-regulated-learning, generative-ai, higher-ed, student-experience, engagement-metrics
📄 Generative AI Availability, Grades, and Student Satisfaction at a Large University
This large-scale observational study tests the "GenAI substitution hypothesis" — the concern that students offload cognitive effort to [[generative-ai]] and earn inflated grades without learning. Usin…
2026-07-24 · generative-ai, higher-ed, learning-gains, student-experience, llm
📄 Experiential Versus Instructional Approaches for Eliciting Metacognitive Awareness in AI-Assisted Learning
A quasi-experimental, short-term longitudinal study with 126 first-year engineering students comparing two ways of teaching students how to learn with generative AI: an experiential, hands-on session …
2026-07-23 · generative-ai, higher-ed, student-experience, scaffolding, self-regulated-learning
📄 Cross-Subject Predictive Validity for Learning Outcomes of Delayed Start Behavior
This study examines the [[student-modeling]] validity of **delayed start behavior** — when students begin assignments or practice sessions past a recommended start time — as a predictor of learning-ga…
2026-06-25 · learning-analytics, student-modeling, higher-ed, engagement-metrics, self-regulated-learning
📄 Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
> **Luyang Fang, Yingchuan Zhang, Jongchan Park, Zhaoji Wang, Ping Ma, Xiaoming Zhai** (2026). arXiv cs.AI preprint…
2026-06-19 · automated-grading, stem-education, formative-assessment, k-12, multi-representational-tools
📄 AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
AI-driven assessment of human tutor training performance correlates with real-life tutoring quality; bridges the gap between training metrics and classroom practice. AI-Driven Assessment of Human Tuto…
2026-06-18 · intelligent-tutoring, automated-grading, feedback-loop, teacher-role, adaptive-virtual-patient-psychotherapy-training
📄 Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
> **Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson** (2026). Pluralistic Alignment Workshop @ ICML 2026…
2026-06-17 · scaffolding, intelligent-tutoring, llm, benchmark, student-experience
📄 Self-Efficacy and Favorability Shape Learning from Tutoring Systems and Paper Practice
> **Xinfei Cen, Vincent Aleven, Kenneth R. Koedinger, Conrad Borchers, Paulo F. Carvalho** (2026). EC-TEL 2026…
2026-06-17 · intelligent-tutoring, personalized-learning, higher-ed, student-experience, self-regulated-learning
📄 Simulating Students' Java Programming Errors with Large Language Models
This paper investigates whether [[llm|large language models]] can serve as scalable proxies for students by simulating realistic logical errors in code submissions. Using the CodeWorkout dataset of 74…
2026-06-15 · llm, stem-education, student-experience, intelligent-tutoring, learning-analytics
📄 The Environmental Cost of LLMs in AIED: Reporting and Practices
> **Sabrina C. Eimler, Lukas Erle, Daniel Flood, Aditi Haiman, Luca Häckert, André Helgert, Lachlan McGinness, Büsra Yapici**…
2026-06-11 · llm, generative-ai, policy-maker, privacy, ethics
📄 AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education
This field experiment shows that AI-generated feedback drafts can measurably increase the rate and length of feedback that teaching assistants actually deliver to students, without sacrificing perceiv…
2026-06-04 · automated-grading, feedback-loop, higher-ed, llm, teacher-role
📄 Effects of an AI-supported inquiry model on AI literacy and authentic performance: A quasi-experimental study with preservice teachers
> **Synthesis:** Effects of an AI-supported inquiry model on AI literacy and authentic performance: A quasi-experimental study with preservice teachers…
2026-06-03 · ai-literacy, faculty-development, generative-ai, higher-ed
📄 The Main Barrier to AI Adoption in the Public Sector is Lack of Training
Through Brazilian government case studies, demonstrates that a four-layer pedagogical methodology (Literacy, Protocol, Prompt Engineering, Audit) is the key to productivity gains (up to 50%), rather t…
2026-06-02 · ai-literacy, public-sector, training-methodology, prompt-engineering, scaffolding
📄 REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
**REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models** advances the [[automated-grading]] frontier by solving a fundamental trust problem: even accurate AI graders are unusable if educat…
2026-05-28 · automated-grading, llm, formative-assessment, higher-ed, scaffolding
📄 Position: Adopting AI in Practice Does Not Guarantee the Productivity Boost
This ICML 2026 position paper argues that adopting AI in organizational practice does not automatically yield productivity gains — human and environmental factors critically moderate the relationship.…
2026-05-26 · generative-ai, teacher-role, higher-ed, policy-maker, persistent-ai-agents-academic-research
📄 The Illusion of Competence: Self-Perceived Digital Literacy and AI Readiness Among European Secondary Students
This multicenter study (N=243 European secondary students) systematically challenges the 'Digital Native' paradigm by demonstrating a severe confidence-competence gap in digital and AI literacy. Stude…
2026-05-26 · k-12, ai-literacy, student-experience, equity, genai-minoritized-knowledges-disability
📄 Explaining Too Much? Understanding How Large Language Model Reasoning Traces Influence Performance and Metacognition
This preregistered between-subjects study (N=559) provides the first rigorous evidence that [[llm]] reasoning traces — increasingly common in AI interfaces — do not improve performance and can activel…
2026-05-26 · llm, metacognition, student-experience, over-reliance, self-regulated-learning
📄 Cognitive offloading and the speedup illusion in human-AI interaction
This preregistered large-scale study (N = 1,237) investigates whether people are well-calibrated in estimating the time savings from AI assistance on simple cognitive tasks. The key finding is a **spe…
2026-05-25 · over-reliance, metacognition, student-experience, cognitive-offloading, ai-assistance-reduces-persistence
📄 Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs
Automated grading systems have enabled scalable assessment for many response types, but handwritten mathematics remains a barrier due to the complexity of multi-step solutions. Vision-capable large la…
2026-05-22 · automated-grading, llm, stem-education, higher-ed, multimodal
📄 How AI Is Changing Teaching Workflows
📄 [Full article](https://edtechinsiders.substack.com/p/how-ai-is-changing-teaching-workflows) AI saves teachers roughly 30% of lesson preparation time with no measurable quality loss — but whether th…
2026-05-21 · generative-ai, teacher-role, faculty-development, rct, k-12
📄 Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom
Oral code reviews paired with a flipped classroom represent a pragmatic harm-reduction approach to generative AI in CS education. Rather than banning LLMs, Fowles et al. (2026) designed weekly formati…
2026-05-21 · generative-ai, higher-ed, cs-education, over-reliance, academic-integrity
📄 Evidence of a Cognitive Shift in AI Education: How Students Are Rethinking Human Intelligence?
This paper presents a striking longitudinal finding: as AI becomes a routine educational tool, students systematically revalue **human intelligence (HI) over artificial intelligence (AI)**. Drawing on…
2026-05-20 · ai-literacy, student-experience, higher-ed, stem-education, over-reliance
📄 From Heuristics to Analytics: Forecasting Effort and Progress in Online Learning
This paper tackles a core ITS challenge: predicting when students will disengage so tutors can intervene before it's too late. It introduces **engagement forecasting** as a supervised prediction task …
2026-05-20 · intelligent-tutoring, learning-analytics, engagement-metrics, k-12, benchmark
📄 An Interpretable Closed-Loop Intelligent Tutoring System for Multimodal Affective Feedback in Asynchronous Presentation Training
Closed-loop ITS with multimodal affective scoring (facial, vocal, textual, oculomotor) produced significant presentation skill gains (Cohen's d = 0.39-0.90, N=204) over 30 days. This paper presents on…
2026-05-19 · intelligent-tutoring, affective-computing, multimodal, higher-ed, professional-training
📄 The Effects of Structured LLM-Generated Feedback on Programming Assignment Performance
LLM-generated feedback produces faster time-to-solution than compiler-only baseline; counterintuitively, less guided feedback showed stronger effects than more guided variants. This study provides emp…
2026-05-19 · llm, generative-ai, feedback-loop, higher-ed, scaffolding
📄 LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
> LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework **Rodríguez (2026)** — Oregon State University. Submitted to Computers & Education.…
2026-05-15 · automated-grading, higher-ed, stem-education, llm, generative-ai
📄 Little Impact of ChatGPT Availability on High School Student Test Score Performance
This paper uses a clever identification strategy: measure the **seasonal drop in ChatGPT activity during non-school summer months** (2023 and 2024). Areas with larger summer dropoffs have heavier scho…
2026-05-14 · generative-ai, k-12, over-reliance, academic-integrity, k-12-ai-education
📄 Improving Hybrid Human-AI Tutoring by Differentiating Human Tutor Roles Based on Student Needs
In a large-scale quasi-experiment with 635 students (grades 5-8), hybrid human-AI tutoring produced substantial gains over AI-only tutoring: +25% time on task, +36% skill proficiency, and +61% standar…
2026-05-14 · intelligent-tutoring, k-12, personalized-learning, learning-gains, equity-in-ai-education
📄 Assessing the Impact and Underlying Pathways of Sequenced AI Feedback on Student Learning
**Sequenced AI feedback harms learning despite boosting engagement and positive perceptions.** In a randomized experiment with 199 participants, the authors compared two types of AI-generated feedback…
2026-05-11 · feedback-loop, formative-assessment, scaffolding, generative-ai, student-experience
📄 The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
> A framework for evaluating AI tutoring systems that extends beyond pedagogical quality of feedback to measure what students actually *do* with that feedback — whether they act on it and whether they…
2026-05-09 · intelligent-tutoring, higher-ed, benchmark, engagement-metrics, llm
📄 NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
> Boateng et al. (2026) introduce **NSMQ Riddles**, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's **National Science and Maths Quiz** — a live TV competition f…
2026-05-08 · benchmark, stem-education, k-12, llm, pedagogical-llm-training
📄 Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
> Schleifer, Ariely & Klebanov (2026) investigate a critical gap in [[automated-grading]]: **how scoring quality degrades for mid-range student responses**. Most ASAS evaluations focus on clearly corr…
2026-05-08 · automated-grading, formative-assessment, llm, benchmark, human-in-the-loop-ai
📄 AI Tutor Effectiveness Review
> Zerkouk, Mihoubi & Chikhaoui (2025) systematically analyzed qualified studies from 2010–2025 across: > A comprehensive systematic review of AI-based Intelligent Tutoring Systems (2010–2025) reveals …
2026-05-07 · intelligent-tutoring, benchmark, higher-ed, k-12, pedagogical-llm-training
📄 Educational LLM Alignment
> Hardy & Kim (2026) identify a **cascading proxy** problem in AI-for-education evaluation: > The gap between what LLMs are *capable* of and what actually *benefits learners* — benchmark performance, …
2026-05-07 · llm, benchmark, bias-mitigation, teacher-role, pedagogical-llm-training
📄 A meta-analysis of the effect of generative AI on productivity and learning in programming
> Maier, Gunzenhäuser & Schweisthal (2026) conduct a **meta-analysis synthesizing evidence** on how generative AI tools affect both programming productivity and learning outcomes. This is a **confiden…
2026-05-06 · rct, generative-ai, higher-ed, learning-gains, meta-analysis

Related Tags

higher-ed (25)llm (24)generative-ai (22)student-experience (16)scaffolding (14)intelligent-tutoring (12)benchmark (10)ai-literacy (10)formative-assessment (10)automated-grading (10)k-12 (9)stem-education (9)metacognition (8)ai-education (6)teacher-role (6)