Research Article
Multimodal Learning with Generative AI
Synthesis: The guide adopts a middle way between "techno-fixing" and rejecting AI as an existential threat. It argues that:
A comprehensive educator's guide to integrating Generative AI into Multimodal AI teaching, learning, and assessment across higher education. Built on Kress's social semiotic theory of multimodality, the guide positions GenAI as a 'cyber-social' partner that complements—but cannot replace—human meaning-making. It proposes the MMLD-AI unifying model (UDL + ABC Learning Design) and the Dual-Track Cyber-Social Learning Model for designing effective human-AI collaboration.^varga-atkins-educators-guide-multimodal-learning-genai-2025
Core Position: Pragmatic, Not Uncritical
The guide adopts a middle way between "techno-fixing" and rejecting AI as an existential threat. It argues that:
- GenAI is already embedded in daily life; ignoring it does students a disservice
- GenAI is not intelligent (no consciousness, understanding, or ethical judgment)
- Human educators and learners bring vision, purpose, nuanced critique, and meaning-making that AI cannot replicate
- Effective use requires cyber-social partnership: humans and machines with complementary strengths
The Four Costs of GenAI
| Cost Domain | Key Concern | Educational Response |
|---|---|---|
| Individual | Privacy, data protection, equity of access, mental health, over-reliance | Transparent authorship; approved tools list; scaffolded critical engagement |
| Environment | Image/video/audio generation uses significantly more energy than text | Mindful use; limit iterations; group demonstrations; digital decluttering |
| Knowledge | Removes sourcing process integral to retention; short-term gains may displace deep learning | Students clarify own understanding before consulting AI; metacognitive Scaffolding |
| Future Jobs | Entry-level white-collar roles vulnerable to automation | Focus on human strengths: contextual reasoning, ethical judgment, craftsmanship |
AI Literacy in Multimodal Contexts
Three Levels
- Basic literacy — Awareness of multimodal GenAI platforms, capabilities, and appropriate uses (creating prompts, generating visual outputs)
- Intermediate literacy — Co-create multimodal content, critically evaluate AI outputs, scaffold uses (transform lecture notes into visuals or podcasts)
- Advanced literacy — Design activities/assessments incorporating multimodal GenAI; lead ethical and philosophical discussions
Four Implementation Scales
| Scale | Strategies |
|---|---|
| Individual | Workshops on creative multimodal tasks; prompt crafting practice; reflective assignments documenting AI use |
| Module | Embed GenAI literacy into learning outcomes; optional multimodal tasks with clear rubrics; creative/reflective critique components |
| Program | Cross-module policies; consistency and transparency via workshops and discussion; alignment with graduate attributes (criticality, Creativity, digital fluency) |
| Institutional | Clear policies with checklists; vetted tools; data privacy protocols enforced; avoid rigid mandates in favor of flexible guidance |
The MMLD-AI Unifying Model
The Multimodal Learning Design with GenAI model merges:
- Universal Design for Learning (UDL): multiple means of engagement, representation, and action/expression
- ABC Learning Design: storyboarding the student journey through learning types
Six Multimodal Engagement Types (adapted from ABC)
- Acquisition of information
- Investigation and/or research
- Collaboration with others
- Production of artifacts (learning, teaching, or assessment)
- Practice of approaches/theories/principles/skills
- Discussion/discourse, including critique/evaluation
For each engagement type, educators decide on multimodal affordances of GenAI and explore respective cyber-social strengths.
The Dual-Track Cyber-Social Learning Model (Galla et al., 2025)
This complementary model maps human vs. AI strengths across Bloom's taxonomy processes:
| Process | Human Strengths | AI Strengths | Cyber-Social Approach |
|---|---|---|---|
| Knowledge & framing | Contextual understanding; embodied knowledge; critical verification | Rapid data retrieval; pattern recognition; broad topical coverage | Humans define purpose and frame problems; AI generates background data and inspiration; humans filter and verify |
| Interpretation & analysis | Causal reasoning; cultural/ethical awareness; implicit meaning | Correlation analysis; theme identification; feature extraction at scale | AI identifies statistical patterns; humans determine causality, relevance, deeper significance |
| Application & prototyping | Situated judgment; adaptive Problem Solving; ethical decision-making; craftsmanship | Rapid simulation; consistent rule application; code/digital artifact generation | AI generates digital prototypes; humans adapt for real-world complexity, apply physical craft, ensure ethics |
| Synthesis & creation | Novel conceptual blending; purpose-driven integration | Cross-domain pattern integration; combinatorial exploration | AI explores possible combinations; humans evaluate, refine, and integrate meaningfully |
Practical Integration: Three Strands
Teaching (Educator-Created Content)
- Generating visuals, diagrams, and infographics from text prompts
- Creating podcast scripts and video summaries
- Building interactive simulations and virtual scenarios
- Using AI to get feedback on marking rubrics and assessment briefs
Learning (Student-Created Content)
- Students transform lecture notes into multimodal artifacts (visuals, podcasts, videos)
- Collaborative group projects using GenAI for brainstorming and prototyping
- Critical evaluation: students annotate AI-generated outputs for accuracy, bias, coherence
- Ethical protocols establishing clear boundaries (e.g., "do not use AI to write reflections; do use it for brainstorming visuals")
Assessment and Feedback
- Multimodal assessment: students submit artifacts combining text, image, audio, video
- AI-assisted peer and self-assessment with structured rubrics
- Educators use GenAI to generate formative feedback at scale, then verify and personalize
- Transparent: assessment briefs explicitly state when and how GenAI may be used
Relationship to Existing Research
| Guide Principle | Knowledge Base Connection |
|---|---|
| Cyber-social partnership (complementary strengths) | A principled way to think about AI in education: guidance for educators and policy makers based on goals, models — "AI must augment, not displace" aligns perfectly |
| Four costs framework (individual, environment, knowledge, jobs) | SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems — Costs to knowledge overlap with cognitive offloading; environmental costs are a new dimension |
| AI literacy levels and scales | AI Literacy — ICAP framework; collaborative learning; this guide adds institutional scaling and multimodal specificity |
| MMLD-AI model (UDL + ABC + six engagement types) | Adaptive Learning — Multi-resolution personalization; Evolution of AI in Education: Agentic Workflows — Planning and reflection paradigms |
| Dual-Track Cyber-Social Model | Training Pedagogical LLMs for Tutoring — Reward "guiding" over "answering"; Human-in-the-Loop — Human verification of AI outputs |
| Multimodal assessment redesign | Authentic Assessment — Six-dimensional framework; Formative Assessment — AI-generated feedback with human validation |
| Scaffolding and metacognition | Self-Regulated Learning — UDL's emphasis on student agency; Metacognition — Cyber-social metacognitive awareness |
| Faculty development across four scales | Educational Development — CTL pragmatic transition model; this guide adds module-level and program-level strategies |
Case Study Themes from the Guide
The guide includes 15+ educator case studies spanning:
- Healthcare: AI avatars for patient communication training (H5P interactive scenarios)
- Bioscience: Multimodal groupwork designing organisms for future Earth scenarios
- Business/HR: Peer conflict resolution with AI-generated scenarios and video avatars
- Education/Teacher training: AI visual metaphors for reflective practice
- Chemistry: AI-generated molecular visualizations and 3D models
- Languages: Text-to-speech and avatar creation for pronunciation practice
- General: Explainer videos, digital posters, podcast scripts, interactive quizzes
Open Questions
- Environmental cost awareness: How can educators and students make informed trade-offs between the pedagogical value of multimodal GenAI artifacts and their energy costs?
- Transfer across modalities: Does competence in AI-assisted multimodal creation in one domain (e.g., visual design) transfer to another (e.g., audio production)?
- Assessment validity: When students use GenAI to create multimodal assessment artifacts, how can assessors distinguish genuine human meaning-making from AI-generated polish?
- Scaling the MMLD-AI model: Can the six engagement types be operationalized as automatic learning design recommendations, or does human pedagogical judgment remain essential?
What this means for practice
- Instructors. Require critique before adoption: build tasks in which students must extend, adapt, critique, or even abandon genAI output rather than submit it wholesale. The guide names uncritical adoption as its first challenge to creative thinking.
- Instructors. Protect unmediated work. The guide warns that the convenience of these tools can foster dependence that erodes independent research, critical analysis, and self-regulation, so state explicitly when GenAI may support a task and when the task is to be done without it.
- Instructional designers. Insert deliberate pause points for reflection into AI-mediated tasks: automated summaries and visualizations can supply answers too quickly and compress the stages of the learning cycle where reflection happens.
- Faculty developers. Choose platforms by purpose rather than novelty. Work through the guide's selection checks — the intended learning outcome, whether visuals, audio or video are genuinely needed, students' digital skills and device access, and bias, privacy and representation — and keep the vetted tools list current, because free access and premium tiers change constantly.
- Faculty developers. Design for the tasks students themselves valued: in the project's focus groups, students responded positively to real-world tasks such as building websites or designing exhibitions, which points to assessment artifacts worth the multimodal effort.
Limitations
- The guide is a synthesis, not an empirical study: it reports data from a literature review, a case-study collection exercise, a survey, and focus groups with educational developers, educators, and students from a 2024/25 SEDA Small Grants project, and measures no learning outcome of its own.
- Its case studies are practitioner submissions reproduced in full in an appendix rather than controlled comparisons — one describes 198 students in a single marketing module across three stages — and none is tested against a no-GenAI condition.
- The environmental argument cannot be quantified: the guide notes precise energy costs for different GenAI platforms are very difficult to extract from their producers and suppliers, so it demonstrates only that multimodal generation uses substantially more energy than text, not how much.
- The evidence base ages quickly and the guide says so: GenAI advanced even during final editing (it reports GPT-5's release in that window), and its cases were gathered from self-selected practitioners already using GenAI, so they document early adopters rather than typical practice.
Citation
Varga-Atkins, T., Saunders, S., Beckingham, S., Hartley, P., Keshishi, N., Lacković, N., et al. (2026). Multimodal Learning with Generative AI.