On this page

Multimodal AI — AI systems that process, understand, or generate content across multiple modalities — text, images, audio, video, and structured data — and the educational questions these systems raise. In AI in education, multimodal AI appears in three distinct roles: as the learning content learners create and engage with (multimodal learning), as the capability boundary of tutoring systems that must interpret diagrams and graphs (multimodal tutoring), and as the assessment signal used to evaluate understanding (multimodal measurement).

Questions to Consider

  • Think of a graph, force diagram, or schematic you've ever struggled to explain in words. What does that experience suggest about the limits of a text-only AI tutor trying to help with image-rich problems?
  • A physics tutor answers text-based problems ~96% of the time but drops to ~74% on problems that embed meaning in diagrams. Before reading, what do you think causes this 'multimodal interference'—and can you think of a fix that doesn't involve retraining the model?
  • You've likely generated both text and images with AI tools. Have you found that 'prompting for pictures' differs from prompting for text? What skills might students need to translate an abstract idea into a precise visual prompt?
  • Multimodal AI can grade essays, generate feedback with audio narration, and even reconstruct exam item statistics from image-and-text items. What does the shift from text-only to multimodal assessment signal (or risk) for fairness and validity?
  • How could the fact that AI support is less reliable on the very diagram-heavy problems that build deep STEM understanding create an equity gap between learners? Who is most affected?
  • Multimodal systems can translate text to audio or visuals to support inclusive learning, but they also enable fine-grained classroom sensing. Where is the line between helpful multimodal access and surveillance?

Introduction

Multimodality in AI refers to the capacity to work across different representational forms rather than text alone. Modern Generative AI and Large Language Models (LLMs) systems increasingly accept and produce images, audio, and video in addition to text, opening new possibilities and new risks for education. Grounded in social semiotic theory, which holds that meaning is made across modes — not just words — multimodal AI changes how teaching, learning, and assessment are designed and evaluated.(Multimodal Learning with Generative AI)

Three faces of multimodal AI in education

1. Multimodal learning and content creation

Multimodal AI enables learners to produce and engage with content across text, image, audio, and video. An educator's guide to multimodal learning with generative AI positions these tools as a "cyber-social" partner: they complement — but cannot replace — human meaning-making.(Multimodal Learning with Generative AI)

  • AI literacy in multimodal contexts is layered: basic awareness of multimodal platforms, intermediate co-creation and critical evaluation of outputs, and advanced design of multimodal activities and assessments.(Multimodal Learning with Generative AI)
  • Multimodal prompting is itself a demanding epistemic practice. Students who prompt for images as well as text discover that "prompt literacy is different between prompting for text than it is for pictures" — translating abstract meaning into machine-readable multimodal prompts requires a precise visual vocabulary and exposes system limitations and bias.(Students' multimodal prompting practices as epistemic work in AI literacy development)
  • Multimodal assessment shifts from essays to artifacts combining text, image, audio, and video, with educators using AI to scaffold creation and feedback rather than replace the learner's own production.(Multimodal Learning with Generative AI)
  • Learner multimodal composing as a critical-thinking scaffold carries a trade-off. Lu et al. (2027) show that having upper-primary students turn written narratives into AI-generated images and short videos supported sustained gains in interpretation, analysis, evaluation, and explanation — but not inference. Because the visuals made story meaning explicit, students reported less need to infer implicit meaning from text alone; peer collaboration, not the multimodal tool, restored occasions for inference. Multimodal AI's value as a meaning-making partner is thus dimension-specific and depends on instructional design that deliberately re-introduces the inferential and self-regulatory work the externalization can short-circuit.
  • Learner multimodal composition as critical AI literacy. Burriss et al. (2026) analyze 22 eleventh graders' 90-second to 3-minute video public service announcements on self-chosen AI Ethics issues — surveillance through school-regulated laptops and electronic "hall passes," informed consent, and punitive algorithmic accusation — as critical AI literacy enacted through composing across moving image, sound, text, and students' own bodies. Across all seven films harm was portrayed as emerging from human–machine entanglement rather than from the tool alone (an anthropomorphized "AI stalker" was played by a human actor in three of seven), and 15 of 18 end-of-unit responses said composing changed their understanding of AI ethics. The authors argue multimodal products both demonstrate and communicate critical competence — productive artifacts, reflections, and civic discourse can serve as Assessment evidence that text-only literacy scales structurally miss.

2. Multimodal tutoring and the capability boundary

When LLM-based tutors must solve problems that embed meaning in graphs, force diagrams, schematics, or tables, their accuracy degrades sharply — the Multimodal Interference Effect.(Multimodal Dialogue in STEM Education)(Multimodal Dialogue in STEM Education)

  • On OpenStax physics problems, text-only accuracy of ~96% drops to ~74% on image-rich problems, consistently across model families.(Multimodal Dialogue in STEM Education)
  • Visual Processing Errors — failures to extract information from graphs or diagrams — dominate the error taxonomy and are the most correctable failure mode.
  • A simple structured-dialogue intervention (have the model describe what it sees, correct only observable misreadings without giving away physics, then re-prompt) restores accuracy to ~95% with zero retraining.(Multimodal Dialogue in STEM Education)
  • This is an equity concern: students working on image-rich problems — precisely the problems that build deep conceptual understanding in STEM — currently receive less reliable AI support than those on text-only exercises.
  • The boundary is a profile, not a level — and artistic imagery sits outside the region models handle well. MUSE (Zhu et al., 2026) evaluates 30 open- and proprietary VLMs on 12 tasks over 1,174 commissioned artworks, and the capability spread across dimensions is wider than any aggregate score suggests: scene classification is near-mature (23 of 30 models above 75.0, median 81.0) while emotion detection tops out at 39.5 and the open-ended tasks that require models to articulate their evidence score at 50.90 (visual clue identification) and 49.18 (emotion cause inference) on semantic similarity. Compositional and viewpoint-dependent reasoning fail hardest — where the ground truth specifies no definite lateral or vertical relation, 90.0% and 73.3% of models assert one anyway, only 43.3% place the girl correctly in depth, and no model resolves all three dimensions of a single item. Failures also cascade: a mis-grounded character is then justified with a fluent rationale built from nearby visual semantics (butterflies, birds), which is the outcome most dangerous in tutoring because the explanation reads as competent. For image-based language learning this argues for dimension-level validation on the imagery a course actually uses, rather than importing a general multimodal score, and for extending the grounding checkpoint described below — describe what is seen, and where, before reasoning from it — to situated artistic content (MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education).

The practical design implication is a visual grounding checkpoint in multimodal tutoring: a deliberate step where the system describes what it sees before attempting a solution, giving the student or a human supervisor a chance to correct perceptual errors.(Multimodal Dialogue in STEM Education)

3. Multimodal assessment and measurement

Multimodal AI broadens both the content of assessment and the signal used to score it.

Multimodal AI for language and accessible learning

Multimodal systems also expand access and personalization. AI-guided audio-video learning tools adapt playback speed, produce multimodal video summaries, and support pronunciation practice.(AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques) Multimodal knowledge graphs reason across images and text for educational tasks,(Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning) and multimodal representations improve Inclusive Learning by translating information across modes (e.g., text to audio or visual). Domain applications include handwritten-math grading and diagnosis,(Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work) affective tutoring with multimodal signals,(An Interpretable Closed-Loop Intelligent Tutoring System for Multimodal Affective Feedback in Asynchronous Presentation Training)(MathBuddy: Affective Math Tutoring) text-to-image learning in specialized fields,(NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts) and privacy-aware multimodal classroom sensing.(Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition) Bird (2026) demonstrates a text-internal form of multimodality: fusing a fine-tuned ELECTRA transformer with computational-linguistics feature analysis to classify English literature by UK Key Stage, where the fused model (F1 0.996) far surpassed every unimodal baseline — evidence that combining representational forms, even within text, can outperform single-model approaches.

Challenges and design implications

  1. Close the multimodal gap. Multimodal tutoring systems should include visual grounding and structured-dialogue scaffolds rather than assuming vision capabilities are robust.(Multimodal Dialogue in STEM Education)
  2. Treat multimodal prompting as a teachable skill. AI literacy curricula must address modality-specific prompting, coherence across modes, and critical evaluation of multimodal outputs.(Students' multimodal prompting practices as epistemic work in AI literacy development)
  3. Preserve human meaning-making. Multimodal AI should augment, not replace, the learner's own construction and evaluation of meaning across modes.(Multimodal Learning with Generative AI)
  4. Extend evaluation to multimodal validity. Assessment validity, bias, and reliability must be examined when AI scores or generates multimodal artifacts.(Multimodal Item Parameter Estimation using Simulated Response Probabilities)(AI Ed Evaluation)
  5. Watch equity and privacy. Unreliable support on image-rich problems and the data demands of multimodal sensing both carry equity and privacy implications.(Multimodal Dialogue in STEM Education)(Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition)

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.