🏷️ multimodal
19 pages tagged with multimodal(19 articles, 0 concepts)
📄 Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform for High Dose Rate (HDR) Brachytherapy
> **Synthesis:** Xu et al. (2026) present an agentic AI-driven immersive simulation for training in **High Dose Rate (HDR) brachytherapy**, integrating VR and mobile computing to create a high-fidelit…
📄 AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques
> **Synthesis:** This dissertation develops an AI-guided learning framework that supports three interconnected stages — Consume, Understand, and Imitate — with three deep-learning systems for audio/vi…
📄 Multimodal Item Parameter Estimation using Simulated Response Probabilities
> **Synthesis:** This paper fine-tunes a multimodal large language model (Qwen3.5-based) to reconstruct multiple-choice model (MCM) and three-parameter logistic (3PL) item characteristic curves. By le…
2026-08-12 · item-response-theory, educational-measurement, student-modeling, llm, automated-assessment
📄 Enhancing online learning outcomes through virtual companion AI: The role of identity anthropomorphism
> **Synthesis:** Grounded in social presence theory, this study introduces the concept of identity anthropomorphism and adopts multimodal learning analytics (MMLA) combining questionnaires, EEG and ey…
📄 Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
> **Synthesis:** This paper introduces an evidence-grounded multimodal pipeline that constructs provenance-rich [[knowledge-tracing|knowledge graphs]] from lecture videos by integrating speech transcr…
📄 NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts
> **Synthesis:** Systematic study of domain-adapted text-to-image models for nuclear engineering education. Fine-tunes Stable Diffusion on nuclear domain images; fine-tuned model achieves 78% domain a…
📄 Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
> **Synthesis:** Pilot study on privacy-aware computer vision for classroom incident detection. Introduces a hybrid benchmark combining generative CCTV-style videos with real classroom pose data. Prop…
📄 SAVVY: Student Attention Visualization for Video-based Learning Analysis
> **Shixian Zhou, Minghuan Shen, Xiaolin Wen, Zijun Qiu, Yongliang Jiang, Xiangyang Wu, Fei Wu, Yong Wang, Zhiguang Zhou** — arXiv preprint (2026).…
📄 Multimodal Dialogue in STEM Education
> **The Multimodal Interference Effect** describes a systemic accuracy drop when LLMs encounter image-rich STEM problems: from ~96% on text-only physics problems to ~74% on multimodal ones. A simple t…
📄 Students' multimodal prompting practices as epistemic work in AI literacy development
> **Synthesis:** Students' multimodal prompting practices as epistemic work in AI literacy development…
📄 TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics
> **Synthesis:** Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focuses on visual programming for pro…
📄 Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
> **MathCog** benchmark (3,036 teacher-annotated diagnostic verdicts, 639 handwritten responses, 18 LLMs): all models severely underperform (macro F1 < 0.5) — over-attributing evidence, overthinking m…
📄 Slide Deck Q&A Quality Assurance App: A Multi-Stage Pipeline for Pedagogical Question Generation
SlidesQAQA is a Flask-based system that extracts text and rendered images from PDF lecture slides and processes them through a four-stage [[llm]] pipeline: **window planning** (segment extraction), **…
📄 Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs
Automated grading systems have enabled scalable assessment for many response types, but handwritten mathematics remains a barrier due to the complexity of multi-step solutions. Vision-capable large la…
📄 ANVIL: Analogies and Videos for Lecturers
Noviello, Birillo, and Migut (2026) present ANVIL, an end-to-end multimodal generation pipeline for educational content — one of the first systems to automate the full journey from concept definition …
📄 An Interpretable Closed-Loop Intelligent Tutoring System for Multimodal Affective Feedback in Asynchronous Presentation Training
Closed-loop ITS with multimodal affective scoring (facial, vocal, textual, oculomotor) produced significant presentation skill gains (Cohen's d = 0.39-0.90, N=204) over 30 days. This paper presents on…
2026-05-19 · intelligent-tutoring, affective-computing, higher-ed, professional-training, efficacy-study
📄 LLM-based Multimodal AI Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback
**AI multimodal feedback matches educator feedback for learning while significantly outperforming it on student perceptions.** The authors built a real-time AI-facilitated multimodal feedback system i…
📄 Educational VLM Evaluation
> Benchmarking vision-language models (VLMs) not on their ability to solve problems, but on their ability to *support learners* — particularly struggling learners and those making errors. Traditional …
📄 Multimodal Learning with Generative AI
> The guide adopts a middle way between "techno-fixing" and rejecting AI as an existential threat. It argues that: > A comprehensive educator's guide to integrating Generative AI into multimodal teach…